RevOps HQ
← BACK TO BLOG
8/1/2026
Data & Governance

HubSpot Duplicate Management: What Deduplicates Automatically, What a Merge Destroys, and How to Prevent the Rest

HubSpot duplicate management: what deduplicates automatically, what the tool matches on, what a merge permanently discards, and how to prevent the rest.

P

Paul Maxwell

AUTHOR

GET WEEKLY REVOPS INSIGHTS

No spam. Unsubscribe anytime.

Duplicates get noticed twice, and the first time is cosmetic — two Sarah Kims in a search result, an awkward moment on a call. The second time is expensive: a revenue report double-counts because one company exists three times, or a re-run import quietly creates a second pipeline's worth of deals, and by then the portal contains work that has to be undone rather than corrected. The team's instinct at that point is a cleanup project, and a cleanup project that does not touch the source produces exactly the same backlog a quarter later.

This article explains what HubSpot resolves on its own, what it does not, and what to do about the difference. It starts with automatic deduplication, which turns on a single identifier per object and explains most of what gets through. It then covers the duplicate management tool — the criteria it matches on, the tiers it needs, and how many pairs it will show you — before turning to merge semantics, which decide whether a mistake is recoverable, and to the 250-merge ceiling that constrains large consolidations. From there: prevention at each of the four entry points, a remediation procedure in order, what prevention is worth against what it costs, and the symptoms that identify each fault.

Automatic deduplication is resolution HubSpot performs without being asked, on a fixed identifier. The duplicate management tool is the feature that lists likely duplicate pairs for review. In a merge, the primary record is the one whose property values take precedence, and the merge itself is permanent — HubSpot documents no way to reverse one.

Automatic Deduplication and Its Single Identifier

HubSpot deduplicates contacts on email address without being configured to. A form submission carrying an email that already exists writes its values onto the existing contact rather than creating a second record, as described in HubSpot's deduplication documentation. Companies deduplicate on company domain name.

Where duplicates enter, and the single identifier that stops some of themRecords enter a portal through manual entry, form submissions, imports and integrations. Contacts are matched automatically on email address and companies on company domain name, so a record carrying that identifier updates an existing record instead of creating a new one. Deals and tickets have no automatic matching at all, so every write creates a record. Anything arriving without the matching identifier becomes a duplicate.HOW RECORDS ARRIVEManual entryForm submissionsImportsIntegrationsWHAT THE PLATFORM CHECKS BEFORE CREATING ONEmatch on emailContactsa different email → a new recordmatch on domainCompaniesa second domain → a new recordno matchDealsevery write → a new recordno matchTicketsevery write → a new recordA record carrying the identifier updates the existing record.A record without it becomes a second record — which, on the data available, is correct.
Four entry points, one identifier gate per object: everything without that identifier passes straight through into a duplicate
What each object deduplicates on by itself, and the duplicates that rule lets through
ObjectContactsDeduplicates automatically onEmail addressWhat gets throughThe same person's personal address, or any submission with no email
ObjectCompaniesDeduplicates automatically onCompany domain nameWhat gets throughMulti-domain organisations, and records created without a domain
ObjectDealsDeduplicates automatically onNothingWhat gets throughEvery re-run import and every loose integration write
ObjectTicketsDeduplicates automatically onNothingWhat gets throughThe same issue reported through two channels

The boundaries of that behaviour define most of the duplicate problem. One identifier per object means a contact who submits a personal address alongside a work address becomes two records the platform considers distinct and correct — which, on the data it has, they are. Companies with several domains split the same way. Deals do not deduplicate at all, which is why a corrective re-import creates a second set of deals rather than updating the first; the identifier-column mechanics are set out in HubSpot's import documentation.

The Duplicate Management Tool and Its Fixed Criteria

The tool is reached from the contacts or companies index through Actions → Manage Duplicates, and it presents pairs a matching model considers likely duplicates for review.

That model does not compare records field by field. For contacts it uses first name, last name, email address, IP country, phone number, zip code and company name; for companies, company domain name, company name, country or region, phone number and industry. The criteria are fixed rather than configurable, and that is the tool's principal limitation: an organisation whose duplicates are distinguished by something outside that list — a customer number, an external system identifier, a location code — will find the tool surfacing pairs that look similar while missing the ones that matter to the business. Detection for those has to happen outside the tool, in an export.

What the tool shows also depends on the subscription. Professional and Enterprise portals manage duplicates individually and see up to 10,000 pairs. Data Hub Professional and Enterprise add bulk management and raise the ceiling to 30,000 and 100,000 pairs respectively, per HubSpot's duplicate management documentation. Results recalculate as records are created, and at least once a day when none are.

A displayed count at the ceiling is a floor, not a measure. Resolving pairs allows previously hidden ones to surface, so the number can stay flat or rise while real work is happening — which is demoralising for whoever is doing it and misleading for whoever is reporting on it. Count pairs resolved instead, in your own tally.

Merge Semantics, Field by Field

A merge is not a tidying operation but a permanent write, and its per-field rules decide whether a wrong choice can be recovered from.

How a merge resolves each field between the primary and secondary recordMerging two contacts field by field. Email addresses from both records are kept, the secondary one added as an additional value. Phone is empty on the primary, so the secondary's value fills the gap. Lifecycle stage keeps the furthest stage reached. Timeline activities from both records combine. Job title exists on both records, so the primary record's value is kept and the secondary record's value is permanently discarded with no record that it existed.MERGING TWO CONTACT RECORDSFIELDPRIMARYSECONDARYMERGEDRULEEmails.kim@acme.comsarah@personal.comboth keptadded as a secondary valuePhone617 555 0114617 555 0114secondary fills the gapLifecycle stageLeadSQLSQLfurthest stage reachedActivities14923both records' timelinesJob titleOps ManagerVP Revenue OperationsOps Managerprimary wins, other destroyedThe discarded value is not stored anywhere, and the merge cannot be reversed.
Field by field: the primary wins where both hold a value, the secondary fills what the primary left empty, and one value is destroyed with no record that it existed

Property values from the primary record take precedence by default, though the merge interface lets you choose which values to keep. Where the primary is empty, the secondary's value fills the gap. Where both hold a value, the secondary's is discarded and retained nowhere. Choosing the primary by creation date rather than by data quality therefore destroys the better record without saying so.

Three exceptions soften that. Contact email addresses and company domain names are added as secondary values rather than replaced, so the alternate identifier survives on the merged record. Timeline activities from both records appear on the merged record, as do associations from both, so history and relationships are not lost. And lifecycle stage keeps the furthest stage reached — the same forward-only behaviour that governs lifecycle stage generally.

HubSpot's merge documentation notes that activities can take up to thirty minutes to synchronise, which matters when verifying bulk work: a record inspected immediately after merging may look like it lost activity that is still propagating. The merge itself cannot be reversed, though for contacts and companies there is a partial recovery — delete the secondary email or domain from the merged record and create a new record with it — but the discarded property values are gone, and an export taken beforehand is the only real recovery path.

The 250-Merge Ceiling

Records cannot be merged once they have been involved in a combined total of 250 merges. At that point the operation is refused and the remaining options are creating a new record or editing by hand.

This constrains large consolidations in a way that is easy to discover late. A single company record absorbing hundreds of regional or import-created variants can exhaust the allowance on the survivor and leave the last several duplicates unmergeable — permanently, since the ceiling cannot be raised. Where a consolidation of that scale is expected, plan the sequence so no single record approaches it: merge in clusters, and let a cluster's survivor absorb the next cluster rather than pointing everything at one record from the start.

Prevention at the Point of Entry

Remediation without prevention regenerates the backlog, and the four entry points fail differently enough to need separate treatment.

Manual entry produces duplicates when the person creating a record did not find the existing one, usually through a spelling difference. Requiring email on contact creation and domain on company creation gives automatic deduplication something to act on, and does more than any amount of training about searching first.

Form submissions are largely handled by the email-address behaviour. The residue is personal addresses submitted to forms intended for business contacts, which is a form-design question — ask for a work email and say why — rather than a data-quality one.

Imports are the largest single source, and the mechanism is specific: rows without a Record ID create new records instead of updating existing ones. An import specification that mandates an identifier column converts a corrective re-run from a duplication event into an update, which is the difference between fixing a mistake and doubling it.

Integrations produce duplicates when their matching rules are looser than the platform's. An integration writing companies without a domain value bypasses automatic deduplication entirely, and it does so on every sync, quietly, at whatever volume the integration runs.

Remediation Procedure

  1. Export every object you intend to touch, unmodified, and keep the files. Merges cannot be reversed, so this export is the only recovery path that will exist.
  2. Identify the source producing the duplicates, using the record source property on a sample of recent pairs. Remediating while an unfixed source keeps writing is bailing with the tap on.
  3. Apply the prevention for that source — required identifiers for manual entry, an identifier column in the import specification, corrected matching in the integration — and confirm it took effect before continuing.
  4. Work the tool's queue, choosing each primary on data quality rather than creation date, and reviewing field values rather than accepting the defaults.
  5. Keep your own count of pairs resolved, since the displayed number is a floor.
  6. For any consolidation above a few dozen records into one survivor, plan the merge sequence against the 250-merge ceiling before starting.
  7. Verify at the end: confirm the queue is drained or triaged, wait past the thirty-minute synchronisation window, then spot-check a sample of merged records for the activities, associations and property values you expect. Checking inside that window produces false alarms and a lot of unnecessary panic.

The Business Case for Prevention Over Cleanup

What prevention buys is the end of a recurring cost. A cleanup project is paid for every time it is run, and it is run again as soon as the source that produced the backlog produces another one. Required identifiers and an import specification are configured once, and they work on records nobody is watching — which is where duplicates come from in the first place.

What it costs is friction at the point of entry, paid by the people least interested in data quality. Requiring an email address on contact creation stops a rep saving a record from a business card that does not carry one. Mandating an identifier column slows a marketer's import by however long it takes to find the Record IDs. Those objections are legitimate and they arrive from people with real work to do, so the standard needs a decision behind it rather than a preference.

The case is strongest where records arrive without human review — imports, integrations, form fills at volume — because that is precisely where nobody notices the second record being created. It is weakest in a small portal where one person enters everything and would recognise a duplicate on sight, and even there the discipline is worth having before the second person is hired rather than after.

Remediation Symptoms and Their Sources

The duplicate count does not fall despite sustained merging. The tool displays a bounded set of pairs, and resolving one lets a hidden one surface. Track pairs resolved rather than pairs remaining, or the work looks futile while it is succeeding.

A merged record holds worse data than one of its inputs. The primary's values took precedence on every populated field, and the secondary's were discarded without being recorded anywhere. Choose the primary on data quality and review the field values during the merge; on records already merged, the pre-merge export is the only source of the lost values.

A merge is refused with no obvious reason, which means one of the records has reached the 250-merge ceiling. Create a new record or edit manually — the ceiling cannot be raised, and no support request changes it.

Activity appears to have been lost after a bulk merge, but synchronisation takes up to thirty minutes, so verify after that window rather than inside it.

Duplicates keep appearing from one source after remediation. Prevention was applied somewhere other than where the records are being created. Read the original source property on a fresh pair and fix that entry point specifically.

Boundaries of This Article

This covers contacts and companies, which are the objects the tool handles and the ones automatic deduplication applies to. Deal and ticket duplicates are mentioned only as a boundary condition — neither deduplicates automatically, and both need import discipline instead of a tool. Matching records across two systems is a related but separate problem, worked through in the HubSpot–QuickBooks integration study, and the identity discipline behind all of it is the same one that governs a full CRM migration.

Tiers, pair ceilings, matching criteria and the merge limit are as documented in August 2026 and move with the product — the pair ceiling has already changed once — so treat HubSpot's documentation as authoritative over this text.

In Summary

HubSpot deduplicates contacts on email and companies on domain, and nothing else automatically at all. Everything arriving without that one identifier — a personal address, a second domain, a deal import, an integration write with an empty domain field — becomes a second record the platform considers correct. The duplicate management tool catches some of the rest using fixed criteria you cannot configure, and shows up to 10,000 pairs on Professional and Enterprise, more with Data Hub.

A merge is permanent, and the primary's values win on every field where both records hold one, the secondary's are discarded and unrecoverable, and only email addresses, domains, activities and associations survive from both sides. That makes primary selection a data-quality decision rather than an ordering one, and it makes the pre-merge export the only recovery path that will exist. Above 250 merges a record cannot be merged again, which is a planning constraint for any large consolidation.

The part worth acting on today is smaller than a cleanup project. Read the record source property on the last ten duplicate pairs in the portal, find which entry point produced them, and fix that one. Everything else in this article is remediation, and remediation is what you do after the tap is off.

HubSpot services

Onboarding, implementation, integration, migration, administration and training, each scoped and priced before the work begins

WEEKLY PROGRAM

RevOps Office Hours

A recurring weekly RevOps operating program. Live support plus hands-on HubSpot implementation work.

$1,500/mo
Monthly Operating Program
  • 1 live Office Hours session per week
  • 4 hours of hands-on implementation work per month
  • Hours allocated against priorities agreed at the start of each period
  • Recurring monthly cadence