Data Quality & Duplicate Management
Duplicate management is a decision system—not a cleanup switch
Detection is not resolution
Salesforce separates matching rules from duplicate rules for a reason. A matching rule defines how records are identified as possible duplicates. A duplicate rule decides what happens when a match is found as records are created or edited. That combination can alert, report or block—but it does not decide which record is authoritative, which values should survive or who is accountable for resolving the case.
The JSBC Labs position is that duplicate management is a decision system. It needs an identity policy, channel-specific behaviour, stewardship and evidence. Activating a standard rule may reduce obvious repetition, but it cannot settle whether two branches of one company, two people sharing an email address or a lead and an existing contact represent the same identity.
Define what ‘same’ means for the use case
A sales team may consider two contacts with the same work email to be one person. A service team may need separate contact records for different customer accounts. A regulated onboarding process may require stronger verified identifiers before records can be consolidated. One universal definition of duplicate can therefore create false merges in one process and missed matches in another.
Start with business scenarios, not fields. Document the entity being matched, the consequence of a false positive, the consequence of a missed match and the evidence needed for automatic versus human decisions. Separate exact identifiers from signals such as name, phone, domain or address. Matching logic should express the identity policy; it should not become the place where an undocumented policy emerges by accident.
Tune for false positives and false negatives
Custom matching rules can compare multiple fields and combine criteria with filter logic. More criteria do not automatically create better quality. Strict AND logic may miss the same customer after a surname, phone number or address changes. Broad OR logic may flag unrelated people who share a household, switchboard number or common name.
Build a labelled test set from real patterns: confirmed duplicates, confirmed distinct records and difficult borderline cases. Include abbreviations, transposed names, international phone formats, missing fields and shared contact details. Measure precision and recall in business language: how many alerts waste a user's time, and how many duplicates pass through? Tune against the cost of each error rather than pursuing an abstract match percentage.
Blocking is a business interruption
Blocking can prevent a new duplicate from entering Salesforce, but it can also stop a legitimate sale, service request or integration transaction. An alert with Allow preserves continuity yet depends on the user understanding the evidence and choosing consistently. Reporting without interruption supports later stewardship but permits bad data to affect routing, automation and analytics before it is resolved.
Choose the action by risk and channel. Block when the match is strong, the cost of duplication is high and the user has a practical path to the existing record. Alert when context is needed. Allow and report when continuity matters more than immediate certainty. Write messages that explain the next action; ‘duplicate detected’ is not useful if the operator cannot safely inspect or resolve it.
Every write channel needs a policy
Duplicate behaviour is not confined to a Lightning record page. Data loads, middleware, mobile clients, Apex and APIs can create or update the same entities. Salesforce's REST and SOAP APIs expose duplicate-rule options, including whether an alerting rule may allow the save and whether record details are returned. That flexibility is valuable—and dangerous when each integration makes a different choice.
Publish one channel matrix covering user entry, imports, migration, web forms and every integration identity. State whether the channel blocks, reports, bypasses or routes candidates for review. Treat any bypass as privileged behaviour with a named owner and monitoring. If an integration allows a duplicate for continuity, persist enough context to reconcile it later instead of turning the header into a silent permanent exemption.
Rule order and scope affect the outcome
An org can have multiple active duplicate and matching rules, but their interaction is not simply cumulative. Salesforce documents that if the first duplicate rule finds a match for a record, later duplicate rules skip that record. Conditions, object scope and cross-object comparisons therefore influence which policy is actually applied and which message or action the writer receives.
Maintain an ordered rule register showing purpose, objects, conditions, matching rule, action and owner. Use conditions to distinguish genuinely different policies, not as patches for unexplained exceptions. When introducing a rule, test it alongside the complete active set. A strong rule evaluated too late may never run, while a broad early rule can conceal more precise logic behind it.
Respect visibility without hiding risk
Duplicate detection can expose the existence or details of records a user would not normally see. Conversely, enforcing the current user's sharing can leave them unable to understand why a save was blocked. The API's DuplicateRuleHeader includes options governing record details and sharing behaviour, which makes this a deliberate security and experience decision rather than a purely technical setting.
Return only the information required to resolve the situation, and test with users who have different sharing and field access. Do not reveal sensitive names, addresses or account relationships merely to prove a match. Where the user cannot access the authoritative record, provide a controlled hand-off to a data steward or routed request. Data quality does not override least privilege.
Merge is a governed data change
Salesforce supports duplicate record sets, duplicate jobs and tools for reviewing and merging records. A merge is not clerical tidying. It can change the surviving values, parent-child relationships, ownership, activity history and downstream identifiers that integrations use. Automation may respond to the resulting updates, while external systems may still reference the losing record.
Define survivor rules by field and source authority before asking users to merge. Identify fields that must be reviewed rather than selected automatically, preserve required evidence and plan how external references are redirected. High-risk entities may require approval or a specialist stewardship queue. Never measure success only by records removed; a bad merge creates a harder data incident than an acknowledged duplicate.
Use the backlog as operational evidence
Allow-and-report rules and duplicate jobs create a queue of potential cases. That queue needs service levels, ownership and ageing measures. Track candidates created, confirmed duplicates, false positives, merge completion time, recurring source channels and records that could not be resolved. Segment results by rule so an apparently healthy total does not hide one noisy or ineffective policy.
Feed those findings back into prevention. A web form may need better verification; a migration may need stronger external IDs; an integration may be retrying without idempotency; a team may be creating contacts because account search is poor. Duplicate volume is often a symptom of process or integration design. Cleaning the records without fixing the writer guarantees another backlog.
Release matching changes like production logic
Changing a matching rule can alter who is blocked, which integrations fail and how much stewardship work is created. Validate the change in a representative environment with realistic data distributions and every material write channel. Compare old and new outcomes on the same labelled sample, assess the new backlog and prepare support guidance before activation.
Assign an owner, review date and rollback decision for each rule. Monitor the first production period closely and retain the evidence behind the change. Duplicate management works when identity policy, platform configuration and human resolution form one operating model. The goal is not a database with zero repeated-looking records; it is a trusted process that prevents, detects and resolves identity ambiguity without interrupting legitimate business.
Official references
- Salesforce Help: Duplicate Rules
- Salesforce Help: Customize Matching Rules
- Salesforce Help: Customize Duplicate Rules
- Salesforce Help: Things to Know About Duplicate Rules
- Salesforce Developers: REST API Duplicate Rule Header
- Salesforce Developers: SOAP API DuplicateRuleHeader
- Salesforce Help: Manage Duplicates Globally