RevOps HQ
← BACK TO BLOG
8/6/2026
Data & Governance

Duplicate Management in HubSpot - Prevention, Detection, Merge Semantics and Recovery

What HubSpot deduplicates automatically and what it does not, the matching criteria behind the duplicate tool, what a merge preserves and destroys, the 250-merge ceiling, and how to run remediation as a programme.

P

Paul Maxwell

AUTHOR

GET WEEKLY REVOPS INSIGHTS

No spam. Unsubscribe anytime.

Duplicates are usually treated as a cleanup task and are better treated as a system property — one of the data-quality responsibilities belonging to the function described in what revenue operations is. A CRM accumulates them continuously through form submissions, imports, integrations and manual entry, so the useful question is not how to remove the current set but which of those sources is producing them and what the platform will and will not resolve on its own.

This covers automatic deduplication, the native tool's matching criteria and limits, what a merge does to the underlying data, and how to structure remediation. Record matching across two systems is a related but separate problem, covered in matching HubSpot companies to QuickBooks customers.

Definitions

Automatic deduplication — resolution HubSpot performs without being asked, on a fixed identifier.

Duplicate management tool — the Professional and Enterprise feature listing likely duplicate pairs for review.

Primary record — in a merge, the record whose property values take precedence.

Merge — the permanent combination of two records into one. There is no undo.

Automatic Deduplication and Its Boundaries

HubSpot deduplicates contacts on email address without being configured to. Where a form submission arrives carrying an email that already exists, the submitted values are written onto the existing contact rather than creating a second record, as described in HubSpot's documentation on contact deduplication.

The boundaries of that behaviour define most of the duplicate problem. It operates on one identifier per object, so a contact submitting a personal address alongside a work address produces two records that the platform considers distinct and correct. Companies deduplicate on domain, which fails for organisations with several domains and for records created without one. Deals do not deduplicate automatically at all, which is why a re-run import creates a second pipeline rather than updating the first — the mechanism is set out in the CRM migration guide.

The Duplicate Management Tool

The tool is reached from the contacts or companies index through Actions → Manage Duplicates, and is available on Professional and Enterprise subscriptions. It presents pairs the matching model considers likely duplicates for manual review.

The model does not compare records field by field. For contacts it weighs first name, last name, email, IP-derived country, phone number, postcode and company name; for companies it weighs domain, company name, country or region, phone number and industry. HubSpot's deduplication documentation sets out the criteria, which are fixed rather than configurable.

That fixed criteria set is the tool's principal limitation. An organisation whose duplicates are distinguished by something outside that list — a customer number, an external system identifier, a location code — will find the tool surfaces pairs it considers similar while missing the ones that actually matter to the business.

The Display Ceiling

The tool displays a maximum of 2,000 duplicate pairs and refreshes periodically, so an account with more than that sees a truncated list which repopulates as pairs are resolved.

This changes how remediation should be planned. A backlog reported as 2,000 is a floor rather than a count, and progress cannot be measured by watching the number fall, because resolving pairs allows previously hidden ones to surface and the figure can remain flat or rise while real work is being done. Measure progress by pairs resolved rather than by pairs remaining.

Merge Semantics

What a merge does to the data determines whether a mistake is recoverable, and it largely is not.

Property values from the primary record take precedence, with the secondary record's value used only where the primary is empty. Choosing the wrong primary therefore silently discards the more accurate data on every field where both records hold a value, and the discarded value is not retained anywhere.

Timeline activities from both records survive and appear on the merged record, as do associations from both, so a merge does not lose history or relationships. HubSpot's documentation on merging records notes that activity synchronisation can take up to thirty minutes, which matters when verifying a bulk operation — a record inspected immediately after merging may appear to have lost activity that is still propagating.

The merge is permanent. There is no undo, and recovery means recreating the record from an export taken beforehand, which is only possible if such an export exists.

The 250-Merge Ceiling

Records cannot be merged once they have been involved in a combined total of 250 merges. At that point the operation is refused, and the remaining options are creating a new record or editing manually.

This constrains large remediation projects in a way that is easy to discover late. An account consolidating a heavily duplicated organisation — one company record absorbing hundreds of regional or import-created variants — can exhaust the allowance on the surviving record and find that the last several duplicates cannot be merged at all. Where a consolidation of that scale is expected, plan the sequence so that no single record approaches the ceiling.

Prevention at the Point of Entry

Remediation without prevention produces the same backlog again, and the four sources produce duplicates differently enough to need separate treatment.

Manual entry creates duplicates when a record is not found because the searcher used a different spelling. Requiring email on contact creation and domain on company creation gives the automatic deduplication something to work with.

Form submissions are largely handled by the email-address behaviour, with the exception of personal addresses on forms intended for business contacts.

Imports are the largest single source, and the mechanism is that rows without a Record ID create new records rather than updating existing ones. An import specification that includes an identifier column converts a corrective re-run from a duplication event into an update.

Integrations create duplicates when their matching rules are looser than the platform's. An integration writing companies without a domain value bypasses automatic deduplication entirely.

Failure Modes

Symptom: the duplicate count does not fall despite sustained merging. Cause: the tool displays at most 2,000 pairs, and resolving one allows another to surface. Fix: track pairs resolved rather than pairs remaining.

Symptom: a merged record holds worse data than one of its inputs. Cause: the primary record's values took precedence on every populated field. Fix: choose the primary on data quality rather than on creation date, and review field values during merge rather than accepting the default.

Symptom: a merge is refused with no obvious reason. Cause: one of the records has reached the 250-merge ceiling. Fix: create a new record or edit manually — the ceiling cannot be raised.

Symptom: activity appears to have been lost after a bulk merge. Cause: synchronisation can take up to thirty minutes. Fix: verify after that window rather than immediately.

Symptom: duplicates keep appearing from one source after remediation. Cause: prevention was not applied at that source. Fix: identify the creating source from the record's original source property before remediating further.

Scope Limitations

The duplicate management tool is a Professional and Enterprise feature, so accounts below that tier need export-based detection instead.

The matching criteria are fixed and not configurable, so an organisation whose duplicate definition differs from HubSpot's will need detection outside the tool.

Nothing here covers deal or ticket duplicates in depth, which behave differently because neither deduplicates automatically.

Verification Checklist

  • The creating source of duplicates has been identified before remediation began
  • Email is required on contact creation and domain on company creation
  • Import specifications include an identifier column
  • Integration matching rules have been checked against the platform's own
  • Primary record selection is made on data quality, not creation date
  • An export exists before any bulk merge, since merges cannot be reversed
  • Progress is measured by pairs resolved rather than pairs remaining
  • No single record in a large consolidation is approaching the 250-merge ceiling

Data quality remediation can be scoped on our store.

Our HubSpot Services

From implementation to optimization, we handle every aspect of your HubSpot journey

WEEKLY PROGRAM

RevOps Office Hours

A recurring weekly RevOps operating program. Live support plus hands-on HubSpot implementation work.

$1,500/mo
Monthly Operating Program
  • 1 live Office Hours session per week
  • 4 hours of hands-on implementation work per month
  • We determine how hours are allocated based on priorities
  • Recurring monthly cadence