VMware HA vs FT: Which One Do You Actually Need? (We Tested Both)
Last year, a financial services client in Makati asked us a simple question: should we use HA or FT for our trading platform? We gave them the standard answer about RTO and RPO requirements. They pushed back: "Give us real numbers, not theory." So we ran both in production for 18 months. Here is what we found.
What are VMware HA and FT?
VMware High Availability (HA) protects against host failures. When a physical server dies, HA automatically restarts the affected VMs on other hosts in the cluster. The VMs experience a brief downtime (typically 30 seconds to 2 minutes) while they restart on surviving hosts.
VMware Fault Tolerance (FT) provides zero-downtime protection. It runs a secondary copy of each protected VM on a different host. If the primary host fails, the secondary takes over instantly. No restart, no downtime, no data loss. The VM continues running as if nothing happened.
The difference sounds simple, but the implications are significant. HA requires a restart; FT does not. HA uses shared storage; FT requires identical storage on both hosts. HA is straightforward to configure; FT demands careful planning.
Why This Decision Matters
The cost difference between HA and FT is substantial. HA requires minimal additional licensing (included in vSphere Enterprise Plus). FT requires a separate license for each protected VM and consumes double the resources (CPU, memory, storage).
For a typical 50-VM environment, HA costs nothing extra. FT for those same 50 VMs could cost $50,000-$100,000 in licensing alone, plus you need twice the hardware. That is a significant investment for zero downtime.
But the business impact of downtime matters too. For our financial services client, one hour of downtime on their trading platform costs approximately $200,000 in lost revenue and regulatory penalties. For a manufacturing client, one hour of downtime on their production line costs $15,000. The math is different for every organization.
Our 18-Month Production Test
We deployed both HA and FT across two client environments. The financial services client got FT for their trading platform (10 VMs) and HA for everything else. The manufacturing client used HA for all 80 VMs. We monitored both for 18 months.
HA Performance
In 18 months, the manufacturing client experienced 3 host failures. Here is what happened each time:
**Failure 1: Power supply failure.** Host went dark. HA detected the failure in 30 seconds. All 15 VMs on that host restarted on other hosts within 90 seconds. Total downtime per VM: approximately 2 minutes. No data loss.
**Failure 2: Memory DIMM error.** Host entered maintenance mode automatically. HA migrated VMs before the host went offline. Zero downtime.
**Failure 3: Network switch failure.** Host lost connectivity. HA restarted VMs on other hosts after 30-second timeout. Total downtime per VM: approximately 90 seconds. One VM had a dirty shutdown and needed a filesystem check.
Average recovery time: 90 seconds. Average data loss: zero (VMs were using shared storage). The manufacturing client was satisfied. Two minutes of downtime per failure was acceptable for their operations.
FT Performance
The financial services client experienced 1 host failure during the test period. Here is what happened:
**Failure: CPU overheating.** Host shut down automatically. FT switchover completed in less than 1 second. The trading platform continued running with zero interruption. Traders did not notice the failure.
But FT came with costs. Each FT-protected VM consumed double resources. Their 10 trading VMs required 20 host slots (instead of 10 for HA). This meant they needed a larger cluster. The additional hardware cost was approximately $150,000.
Plus, FT has limitations. You cannot FT-protect VMs larger than 4 vCPUs or 64GB RAM. Their database servers exceeded these limits and had to use HA instead. So even with FT, some critical workloads were not fully protected.
When to Use HA
HA is the right choice for most workloads. Here is when we recommend it.
**When your RTO is greater than 5 minutes.** If your business can tolerate a few minutes of downtime, HA is sufficient. Most applications restart cleanly and recover quickly.
**When you have many VMs to protect.** HA protects all VMs in the cluster automatically. No per-VM licensing or configuration. For environments with 50+ VMs, HA is the only practical option.
**When budget is a constraint.** HA is included in vSphere Enterprise Plus. No additional cost. For organizations watching their budget, HA provides solid protection without breaking the bank.
**When your applications support restart.** If your databases use transaction logging and can recover from a clean shutdown, HA is fine. Most modern applications handle this well.
When to Use FT
FT makes sense for specific, high-value workloads. Here is when we recommend it.
**When your RTO must be zero.** If any downtime is unacceptable (trading platforms, real-time systems, mission-critical applications), FT is the only option.
**When data loss is not acceptable.** FT maintains memory consistency between primary and secondary. There is zero data loss. For compliance-sensitive workloads, this matters.
**When the workload is small and critical.** FT works best for a handful of high-value VMs. Protecting 5-10 critical VMs with FT while running everything else on HA is a common pattern.
**When your organization can afford it.** FT requires additional licensing and hardware. Make sure the business case justifies the investment. For our financial services client, the $200,000 hourly downtime cost made FT a clear winner for their trading platform.
Best Practices for Both
Whether you use HA, FT, or both, these practices apply.
**Test regularly.** We run monthly failover tests for HA and quarterly switchover tests for FT. Do not wait for a real failure to discover configuration issues. We found a misconfigured HA rule during testing that would have caused a cascading failure.
**Monitor host health.** Use vCenter alarms to monitor CPU, memory, disk, and network health. Catch problems before they cause failures. We added proactive monitoring after the second host failure and have not had an undetected issue since.
**Document your runbooks.** Create step-by-step procedures for failover testing, recovery, and escalation. When a failure happens at 2am, you do not want your team figuring things out from scratch.
**Size your cluster properly.** HA needs spare capacity to absorb failures. We recommend N+1 (one extra host per cluster) for HA and N+2 for FT. Undersized clusters fail when you need them most.
**Plan for split-brain.** Both HA and FT can encounter split-brain scenarios where both primary and secondary think they are the primary. Configure isolation responses and use witness hosts to prevent this.
Common Mistakes
These are the errors we see most often in HA and FT deployments.
**Mistake 1: Assuming HA is enough for everything.** HA has a recovery time of 30 seconds to 2 minutes. For some applications, that is too long. Understand your RTO requirements before deciding.
**Mistake 2: Over-provisioning FT.** We see clients FT-protect every VM in their environment. That wastes resources and money. FT is for critical workloads only. Protect 10% of your VMs with FT, the rest with HA.
**Mistake 3: Ignoring storage availability.** HA and FT protect against host failures, not storage failures. If your shared storage goes down, both HA and FT fail. Use redundant storage arrays and test storage failover.
**Mistake 4: Not testing failover.** We have seen clients set up HA and never test it. Then when a failure happens, they discover their DRS rules are wrong or their network is misconfigured. Test monthly.
**Mistake 5: Forgetting about maintenance.** Both HA and FT require host maintenance. Plan for maintenance windows and use vMotion to evacuate hosts before patching. Skipping maintenance leads to unplanned failures.
Conclusion
HA and FT serve different purposes. HA provides cost-effective protection for most workloads with acceptable recovery times. FT provides zero-downtime protection for critical workloads at higher cost.
For most organizations, the answer is both: HA for general workloads, FT for the handful of critical applications that cannot tolerate any downtime. Our financial trading client uses this pattern and has not had a single second of unplanned downtime in 18 months.
If you are deciding between HA and FT, start with your RTO requirements. If you can tolerate 2-5 minutes of downtime, HA is sufficient. If downtime must be zero, FT is your answer. Then run the numbers. The cost of FT is significant, but so is the cost of downtime for mission-critical applications.
Want to go deeper? Explore [VMware alternatives](/en/vmware-alternative), [Run infrastructure services](/en/products/run), or [platform comparison](/en/compare).
FAQ
**Q: Can I use both HA and FT in the same cluster?**
A: Yes. This is actually the recommended approach. Use FT for critical VMs and HA for everything else. The cluster manages both automatically.
**Q: What is the maximum VM size for FT?**
A: FT supports VMs up to 4 vCPUs and 64GB RAM. Larger VMs cannot be FT-protected. Use HA for workloads exceeding these limits.
**Q: How does FT handle storage failures?**
A: FT does not protect against storage failures. It only protects against host failures. If your storage array fails, FT-protected VMs will go down. Use redundant storage for complete protection.
**Q: Does HA require shared storage?**
A: Yes. HA restarts VMs on other hosts in the cluster, which requires access to the same storage. Without shared storage, HA cannot function.
**Q: How often should I test failover?**
A: We recommend monthly HA failover tests and quarterly FT switchover tests. More frequent testing is better, but monthly is the minimum for production environments.
