Skip to content
Blurpx
Infrastructure 6 min read

Business continuity for enterprise email: how to build a disaster recovery plan

When email stops, business stops. Drawing on the setup we built for a client in aviation, we walk through the steps to outage-resilient email infrastructure and a tested disaster recovery plan.

Blurpx Team

Blurpx Teknoloji

Business continuity for enterprise email: how to build a disaster recovery plan
Contents
  1. Why you need a separate plan
  2. 1. Define your targets: RTO and RPO
  3. 2. Assess where you are
  4. 3. Design for redundancy
  5. High availability
  6. Geographic separation
  7. The 3-2-1 backup rule
  8. MX and DNS redundancy
  9. 4. Make security part of the plan
  10. 5. Write an incident response plan
  11. 6. Test, then test again
  12. 7. Monitor continuously
  13. Common mistakes
  14. Cloud or on-premises?
  15. Keeping the plan alive
  16. Checklist
  17. Conclusion

In critical sectors such as aviation, email is the backbone of operations: flight planning, supplier correspondence, communication with regulators and team coordination all depend on it. For Skytech Group we built email infrastructure that runs without interruption across several locations, along with a disaster recovery framework. In this post we share the steps any organisation can follow to build something similar.

Why you need a separate plan

When most organisations say "we have backups", they mean there is a copy of the data somewhere. Disaster recovery is more than a copy: it defines how quickly, in what order and under whose responsibility the service comes back. A server failure, a misconfiguration, a ransomware attack or a power outage at the data centre are all different scenarios that need different responses.

1. Define your targets: RTO and RPO

Two numbers sit at the heart of the plan:

  • RTO (Recovery Time Objective): the longest the service can be down. For example, "email must be working again within an hour".
  • RPO (Recovery Point Objective): the most data you can afford to lose. For example, "we can lose at most the last 15 minutes of email".

The smaller these numbers, the more complex and costly the infrastructure. That is why the technical team should set them together with the business, not alone. Not every team needs the same target either: an operations mailbox and an archive mailbox can have different priorities.

2. Assess where you are

A plan starts with an honest inventory. You should have clear answers to these questions:

  • Which systems are critical? (Mail servers, authentication, DNS, archive, mobile access)
  • Where are the single points of failure? (One server, one internet line, one DNS provider)
  • Where are backups kept, how, and how often? Has a restore ever been tested?
  • Is it written down who does what during an outage?

3. Design for redundancy

High availability

Running mail servers on more than one node means the service continues on another node when one fails. How mailbox data is replicated between nodes, synchronously or at short intervals, directly determines your RPO.

Geographic separation

Two servers in the same data centre do not help if the whole data centre becomes unreachable. For critical systems, at least one copy should live in a different location.

The 3-2-1 backup rule

Under this widely used rule you keep at least 3 copies of your data, on 2 different types of storage, with at least 1 copy off-site. Against ransomware, at least one backup should be immutable or offline.

MX and DNS redundancy

Email delivery relies on MX records in DNS. A secondary MX record and a queuing mechanism ensure incoming mail is not lost when the primary server is unreachable. DNS itself should not depend on a single provider, and sensible TTLs for records speed up the switch-over.

4. Make security part of the plan

Many outages come from security incidents rather than hardware failure, so the recovery plan should not be separate from security measures:

  • SPF, DKIM and DMARC records make it harder to spoof your domain and improve deliverability.
  • Two-factor authentication means a stolen password alone is not enough to access an account.
  • Encryption should apply both in transit (TLS) and at rest.
  • Admin access should follow least privilege and be logged.

5. Write an incident response plan

During an outage, time is the most valuable thing and decisions are hardest to make. So write a short, practical response plan for each scenario: who decides, who brings which system online, and how and by whom users are informed. The plan should be clear enough for a non-technical manager to follow.

6. Test, then test again

An untested recovery plan is only an assumption. Regular drills should confirm that:

  1. A single mailbox can be restored from backup within the target time.
  2. The secondary system takes over within the RTO when the primary is shut down.
  3. Incoming mail is delivered without loss during the switch-over.
  4. The people named in the plan can follow the steps using the written plan alone.

After every drill, compare the measured times with the targets and update the plan.

7. Monitor continuously

Disk usage, queue length, replication lag, certificate expiry and the success of backup jobs should be monitored continuously, with alerts when thresholds are crossed. Many outages start as a small, visible problem that was left to grow.

Common mistakes

Reviewing email infrastructure for different organisations over the years, we have seen the same mistakes again and again:

  • Keeping backups in the same environment: storing the mail server and its backups on the same disk array or under the same account means an attack such as ransomware can take out both at once.
  • Never testing a restore: a backup job showing "successful" does not mean the data can be restored. Corrupt or incomplete backups are usually discovered only when they are needed.
  • Undocumented configuration: when only one person knows how the server is configured, recovery takes far longer if that person is unavailable.
  • Not tracking certificate and domain expiry: an expired TLS certificate or an unrenewed domain can cause an outage as serious as a hardware failure.
  • Forgetting user communication: an alternative channel (SMS, corporate messaging) for telling users what happened and when service will return should be agreed in advance.

Cloud or on-premises?

Enterprise email can run on cloud services, on-premises servers or a hybrid of the two. Cloud services hand much of the infrastructure redundancy to the provider, but you still need a plan for account security, accidentally deleted data and provider-side outages. On-premises servers offer more control, but redundancy, monitoring and updates are entirely your responsibility. Regulatory requirements, data residency and cost are the main factors behind the decision.

Whichever model you choose, "the provider takes care of it" is not enough on its own. Under the shared responsibility model, protecting the data itself and securing accounts are usually the customer's job.

Keeping the plan alive

A disaster recovery plan is most accurate on the day it is written; after that, every infrastructure change, new application or staff change makes it a little more out of date. The plan therefore needs an owner, and it should be reviewed at regular intervals and after every major change. Drill results are the most valuable evidence of how well the plan really works.

Checklist

  • RTO and RPO targets agreed with the business.
  • Inventory of critical systems and single points of failure completed.
  • High availability and a geographically separate copy in place.
  • Backups follow 3-2-1, with at least one immutable copy.
  • Secondary MX and DNS redundancy configured.
  • SPF, DKIM, DMARC and two-factor authentication enabled.
  • Written response plan with named owners.
  • Regular recovery drills, with results recorded.

Conclusion

For Skytech Group, the infrastructure went live after thorough testing; operational reliability improved, the risk of downtime fell and critical data became better protected. If you want to build a continuity plan for email or another critical system, we would be glad to review your current setup with you.

Frequently asked questions

What is the difference between RTO and RPO?

RTO is the longest the service can be down; RPO is the most data you can afford to lose. One is about time, the other about data.

Are backups and disaster recovery the same thing?

No. A backup is a copy of the data. Disaster recovery is a tested plan that defines how quickly, in what order and by whom the service is restored.

How often should recovery drills be run?

It depends on how critical the system is; for critical systems, several times a year and after every major infrastructure change is a good rule.

Why do SPF, DKIM and DMARC matter?

They make it harder to spoof your domain, reduce the chance of your email landing in spam and help protect against phishing.

Share

More posts

Let's plan your next project together.

Tell us what you need, and we'll come back with a roadmap covering approach, scope and timeline.