Reliability & Disaster Recovery
A Salesforce backup is not a recovery plan
Copies are inputs to recovery
A completed backup job is comforting because it produces a timestamp, a success status and a collection of files. None of those proves that the business can resume after an administrator deletes records, an integration corrupts relationships, a deployment changes behaviour or a compromised account damages data. A backup is an input to recovery; it is not the outcome.
A recovery plan connects technical restoration to a business service. It defines what must return, how much data loss is acceptable, how long users can wait, who can authorise destructive steps and how the team will confirm that operations are safe to resume. The JSBC Labs view is that backup success should be measured by demonstrated recoverability, not by the number of green jobs on a dashboard.
Start with RPO and RTO by capability
Recovery point objective, or RPO, is the maximum acceptable data-loss window. Recovery time objective, or RTO, is the maximum acceptable duration of disruption. A daily copy implies very different exposure from continuous protection, and a manual rebuild has very different timing from an automated restore. Neither target should be invented by the platform team in isolation.
Define targets for business capabilities rather than assigning one number to the whole org. Order capture, regulated consent and payroll may require tighter protection than historical reporting or temporary workflow state. Salesforce Well-Architected guidance makes the same distinction: RPO shapes backup or replication frequency, while RTO shapes automation, redundancy and recovery investment. Aggressive promises need correspondingly mature architecture and rehearsals.
Platform availability does not reverse logical damage
Salesforce designs and operates resilient infrastructure, but infrastructure availability and application recovery solve different failures. Replication protects service continuity when underlying components fail. It can also replicate a validly committed mistake: a destructive import, faulty automation, unauthorised update or bad integration payload becomes durable everywhere the current state is copied.
Model application-layer incidents explicitly. Include accidental deletion, widespread field corruption, incorrect ownership, malicious change, broken metadata, file loss and external-system replay. For each scenario, determine the affected scope and safe recovery point. Waiting for an infrastructure incident playbook to solve a logical data incident confuses Salesforce's responsibility for the platform with the customer's responsibility for the solution built on it.
Protect data, metadata and files separately
A Salesforce service is not only rows in standard and custom objects. Its behaviour lives in metadata: objects, fields, permissions, Apex, Flow, validation, layouts and integration configuration. Its evidence may live in ContentVersion binaries, attachments and documents. A data export cannot recreate missing automation, and a source repository cannot restore customer records or file contents.
Maintain an inventory for all three layers. Back up business data at a frequency aligned to its RPO. Keep deployable metadata in version control and ensure the repository reflects authorised production state. Export file binaries with identifiers and relationships required for reattachment. Salesforce provides native export and Backup and Recover capabilities, while Metadata API and Salesforce CLI support metadata retrieval; the chosen combination must cover the actual service, not just the easiest objects.
Restore order is architecture
Salesforce relationships, ownership, sharing and automation make restoration an ordered operation. Parents may need to exist before children, users before owned records, reference data before transactions and metadata before data that depends on fields or validation. Original record identifiers may not survive every recovery approach, so lookup remapping and external identifiers require deliberate design.
Create dependency waves and test them at realistic volume. Decide when automation, duplicate rules, sharing calculations and outbound messages must be disabled or controlled. Rebuild relationships, files and derived values, then allow formulas, roll-ups and downstream processes to recompute. A CSV collection without a validated load sequence is an archive, not a recovery capability.
Choose the recovery point with evidence
The newest backup is not automatically the safest. If corruption began three days ago and spread gradually, last night's copy may contain the same damage. The team needs a timeline based on audit evidence, integration logs, deployment history, user reports and data profiling to identify when the system was last trustworthy.
Define how investigators will compare candidate restore points and how the business will approve the selected one. Preserve affected evidence before overwriting production. When partial recovery is possible, identify records and fields precisely and protect legitimate changes made after the incident began. Recovery can cause a second loss if a broad restore replaces valid work that users performed during the contamination window.
Connected systems change the plan
Restoring Salesforce to an earlier state does not rewind payment platforms, data warehouses, marketing systems or middleware. Those systems may already have consumed events and acted on corrupted records. Once Salesforce returns, replayed outbound messages can duplicate orders or communications, while inbound integrations can immediately reintroduce the bad state.
Map every critical dependency. Define how to pause traffic, drain queues, preserve correlation identifiers, reconcile both sides and resume in a controlled order. Credentials, OAuth authorisations and external identity configuration may also need re-establishment because they are not equivalent to ordinary data or metadata. Salesforce disaster-recovery guidance calls out these external dependencies; an org-only restore cannot satisfy a service-level RTO when authentication or integrations remain unavailable.
Recovery needs a clean control plane
An incident may remove access to the tools and people normally used to coordinate recovery. If the runbook, contact list and approvals exist only inside the affected Salesforce org, responders can lose them when they are most valuable. If the same compromised administrator can alter production and delete its backups, the control has an obvious common failure mode.
Keep protected recovery documentation, credentials and escalation contacts outside the primary failure boundary. Separate backup administration from ordinary production access where risk warrants it. Record who can initiate, approve and monitor a restore. Include alternate personnel for every critical role. Business continuity depends as much on reachable knowledge and authority as it does on stored bytes.
Validation determines when service is restored
A restore job completing does not mean the service is ready. Validate record counts, control totals, relationship integrity, representative files, ownership, permissions, automation and integration checkpoints. Use business scenarios such as creating an order, progressing a case or recording consent—not only technical queries—to prove that critical capabilities work end to end.
Define acceptance criteria before the incident. Name the technical and business approvers and distinguish minimum viable recovery from full restoration. Some lower-priority history or files may return later while critical transactions resume. This staged approach is safe only when the remaining gaps are understood, communicated and monitored rather than hidden behind a generic ‘system available’ message.
Rehearse until the timings are credible
The first full restore should never occur during the first real crisis. Run partial drills for metadata, selected objects and files, then conduct a broader recovery exercise in an isolated non-production environment. Measure actual extraction, preparation, loading, reconciliation and approval time. Include backup personnel rather than relying on the one expert who wrote the runbook.
Salesforce Well-Architected guidance recommends regular restore validation and timed disaster-recovery exercises. Use each rehearsal to update dependencies, automation controls, scripts, ownership and estimates. Track backup freshness, failed jobs and achieved RPO/RTO as operational measures. A credible recovery plan is executable evidence: the organisation has restored the right service, in the right order, within an agreed window—and knows how to do it again.
Official references
- Salesforce Architects: Reliability in the Well-Architected Framework
- Salesforce Architects: Disaster Recovery and Business Continuity Patterns
- Salesforce Architects: Incident Response Patterns
- Salesforce Help: Best Practices to Back Up Salesforce Data
- Salesforce Help: Back Up Metadata to Protect and Restore Customizations
- Salesforce Help: Back Up Data with Salesforce Backup and Recover
- Salesforce Developers: Backup and Recover MCP Server