Most businesses first meet disaster recovery strategies when they start pricing them — and a recurring pattern shows up: the strategy gets chosen by how it sounds, not by what the system needs. “Active/active” sounds like the serious, grown-up option, so it gets requested for everything, including systems that could tolerate a few hours of downtime without meaningful harm. Meanwhile, a quieter system that genuinely needs a near-instant recovery is left on basic backups, because nobody flagged it as important enough to ask about.
As covered in The RTO/RPO Gap, the fix starts with knowing your actual recovery targets before shopping for a strategy. This piece covers the other half: what the four standard AWS disaster recovery strategies look like, and how to match one to each system.
The four strategies, in increasing order of readiness (and cost)
Backup & Restore — the baseline. Data is backed up regularly, but there’s no standing infrastructure in a second location. Recovery means provisioning everything from scratch and restoring data onto it, which typically takes hours, or up to a day, depending on data volume and how automated the restore is. Cheapest by a wide margin, and often entirely appropriate for lower-priority systems.
Pilot Light — the core of a system (typically its database, kept in continuous or near-continuous replication) exists in the standby region, but the rest of the infrastructure sits switched off or minimally provisioned until needed. On failover, that infrastructure is scaled up around the already-replicated data. Recovery typically takes tens of minutes to a few hours, at a much lower ongoing cost than keeping everything running.
Warm Standby — a scaled-down but fully working copy of the environment runs continuously in the standby region: smaller instance sizes or lower capacity than production, but live and able to serve traffic. Failover means scaling it up to full capacity and redirecting traffic. Recovery drops to minutes, at a higher ongoing cost than Pilot Light, since real infrastructure runs around the clock even when nothing has gone wrong.
Multi-Site Active/Active — both regions run at full production capacity, both serving live traffic under normal conditions. If one fails, the other is already handling load and absorbs the rest. This is as close as AWS disaster recovery gets to zero recovery time, and it’s priced accordingly — effectively running, and paying for, the full environment twice.
What should actually drive the choice
Not a gut sense of “how important is this system”. The two numbers that matter are the Recovery Time Objective (how long this system can be down before the business is genuinely harmed) and the Recovery Point Objective (how much data loss, measured in time, is tolerable). Both should come from a proper Business Impact Analysis, not an assumption made in a meeting. Each step up the list raises both readiness and ongoing spend, so the honest question for every system isn’t “what’s the best option available” — it’s “what does this system’s RTO and RPO actually require, and what’s the cheapest strategy that meets it?”
The mistake we see most often
Two versions of the same error come up repeatedly. The first is paying for Multi-Site Active/Active on a system that could tolerate a few hours of downtime — permanently running duplicate infrastructure for readiness that will, in all likelihood, never be used. The second, more dangerous version is the opposite: settling for Backup & Restore on a system that needs a recovery time measured in minutes. That gap stays invisible right up until the moment it matters most — the worst possible time to discover it.
Neither error is really about technology. Both come from skipping the step of setting real RTO and RPO figures before choosing a strategy — which is exactly why that step comes first.
Matching strategy to system, not to a menu
Most businesses end up with a mix of all four across their estate, the same way a migration ends up as a mix of approaches rather than one applied everywhere. A handful of critical systems might justify Warm Standby or, more rarely, full Active/Active. Many sit comfortably on Pilot Light. Lower-priority systems are often well served by Backup & Restore. The right question is never “which strategy should we standardise on” — it’s “which strategy does each system’s impact analysis justify?”
How this maps to our resilience tiers
Each of our disaster recovery tiers is one of these four strategies, chosen per system:
| Tier | AWS strategy | Typical recovery | Typical data loss |
|---|---|---|---|
| Foundational | Backup & Restore | Hours, up to a day | Back to the last backup |
| Standard | Pilot Light | Tens of minutes to a few hours | Minutes |
| Resilient | Warm Standby | Minutes | Seconds to minutes |
| Continuous | Multi-Site Active/Active | Near-instant | Near-zero |
These ranges are typical, not guaranteed — the actual targets for each system are agreed during the assessment and then proven in a failover drill.
Where we fit
Getting from a vague sense of “this matters” to a specific, defensible strategy per system — and a tested runbook proving it works, not just a diagram claiming it would — is the core of a Yemberzal Disaster Recovery Assessment.
Frequently asked questions
What are the four disaster recovery strategies on AWS? Backup & Restore, Pilot Light, Warm Standby and Multi-Site Active/Active — in increasing order of readiness and cost.
What’s the difference between pilot light and warm standby? In Pilot Light, only the data layer runs in the second region and everything else is started on failover. In Warm Standby, a smaller but complete working copy runs all the time, so failover is a matter of scaling it up — faster, but more expensive to keep running.
Is backup and restore enough for disaster recovery? For some systems, yes. It’s the right choice when a system can tolerate hours of downtime and losing data back to the last backup. For systems that can’t, backups alone are not a disaster recovery plan — see Why “We Have Backups” Is Not a DR Plan.
Should every system use the same strategy? No. The strategy should follow each system’s RTO and RPO. Most estates use a mix.