Blog· HubSpot CRM 9 min read

Your HubSpot Reports Are Wrong: A Practical CRM Data Cleanup Guide

Almost every HubSpot portal over two years old has dashboards nobody quotes in a meeting. The cause is always the same handful of data problems — here is the order to fix them in.

Your HubSpot Reports Are Wrong: A Practical CRM Data Cleanup Guide — Hubstack HubSpot article cover
In this article
  • Step 1: Decide what the numbers are supposed to mean
  • Step 2: Audit properties and kill the dead ones
  • Step 3: Deduplicate contacts and companies properly
  • Step 4: Fix marketing versus non-marketing contacts
  • Step 5: Deal with the deals
  • Step 6: Rebuild the dashboards, then delete the old ones
  • Step 7: Keep it clean with automation, not willpower

There is a specific moment in most HubSpot portals where reporting stops being used. Someone quotes a number in a meeting, someone else contradicts it from a different dashboard, and from then on people go back to spreadsheets. The dashboards keep existing; nobody cites them.

The cause is almost never HubSpot. It is duplicate records inflating counts, lifecycle stages that mean different things to different teams, properties written by three sources at once, and deals that have been open since a rep who left in 2024 created them.

This is the cleanup sequence we run, in order, because order matters — deduplicating before you fix lifecycle definitions just means you merge records into the wrong stage. If you would rather hand the whole thing over, that is what HubSpot CRM setup and automation covers.

Step 1: Decide what the numbers are supposed to mean

Before touching a record, write down definitions for every lifecycle stage and every deal stage. A lead is what, exactly? Who moves a contact to MQL, and on what trigger? What has to be true for a deal to leave 'discovery'?

Do this in a shared document with sales and marketing in the room, and keep it to one line per stage. If two people give different answers, you have found the actual reason your reports disagree — and no amount of data hygiene fixes a definitional dispute.

Write exit criteria, not descriptions. 'Has agreed to a scoped proposal' is auditable. 'Is interested' is not.

Step 2: Audit properties and kill the dead ones

Export your property list and mark each one: actively used, historical, or dead. Most portals carry dozens of properties created for a campaign three years ago, half-populated, still appearing in filter dropdowns where somebody will eventually pick the wrong one.

For every surviving property, record who writes to it — form, import, integration, workflow, or human. A property with three writers and no owner is the single most common source of a broken report.

Archive rather than delete where history matters, and rename dead-but-needed fields with a clear prefix so nobody filters on them by accident.

The property audit table we keep for every portal
PropertyWritten byUsed inVerdict
Lifecycle stageWorkflow onlyAll pipeline reportingProtect — single writer
Original sourceHubSpot trackingAttribution reportingProtect — never overwrite
Lead source (custom)Form + import + repNothing currentRetire — duplicates original source
Deal stageRep, with exit criteriaForecastingProtect — needs written criteria
Campaign 2023 flagOne-off importNothingArchive

Step 3: Deduplicate contacts and companies properly

Run HubSpot's duplicate management tool first, then hunt the duplicates it misses: personal versus work email on the same person, company records differing by 'Ltd' versus 'Limited', and contacts created by an integration that keys on a field your forms do not collect.

Merge in the right direction. The surviving record should be the one with the richest activity timeline, because associations and engagement history are what your reporting reads.

Then close the tap. Deduplicating without fixing the source — an import without a matching key, a form collecting a second email field, an integration creating rather than updating — means you repeat this in four months.

Step 4: Fix marketing versus non-marketing contacts

Suppliers, job applicants, partners, form spam and long-dead imported lists should not be marketing contacts. This is both a data quality issue and a billing one, because marketing contact tiers price in blocks.

Set non-marketing as the default on import, then use a single workflow to promote contacts when a campaign genuinely needs them.

This one change routinely drops a portal a full tier at renewal with no loss of capability — see the HubSpot pricing breakdown for how the tiers behave.

Step 5: Deal with the deals

Open deals with no activity for 90 days are not pipeline; they are optimism. Build a list, review it with the owner, and close them as lost with a reason. Forecast accuracy improves immediately, and usually the forecast gets smaller and more believable.

Reassign deals owned by departed users before doing anything else, or your rep performance reporting stays permanently wrong.

Add required close reasons on lost deals going forward. Six months of honest loss reasons is more useful than any attribution report you will build this year.

Step 6: Rebuild the dashboards, then delete the old ones

Once definitions and data are fixed, build a small number of dashboards that answer named questions for named people. Three good dashboards beat fifteen that nobody opens.

Delete or archive the old ones. Leaving contradictory dashboards in place guarantees somebody quotes the wrong number in a meeting and the trust problem returns.

Put a documented owner and a review date on each surviving dashboard. Reporting decays without maintenance, exactly like a website does.

Step 7: Keep it clean with automation, not willpower

Set up hygiene workflows: normalise country and job title values, flag records missing required fields, alert on deals stalled beyond a threshold, and set non-marketing status on new records from non-campaign sources.

Add validation at the entry point — required fields on forms, restricted dropdowns instead of free text, and progressive fields rather than one long form that reps fill with placeholder text.

Then schedule a quarterly 90-minute audit: duplicates, unowned properties, stalled deals, seat usage. Cleanup is cheap when it is routine and expensive when it is a project. Our ongoing HubSpot support exists largely because of this.

FAQ

Questions people actually ask AI about this.

Why don't my HubSpot reports match what my team believes?

Usually because lifecycle and deal stages have no written exit criteria, so different people move records at different moments. Duplicates and multi-writer properties then inflate the gap. Fix definitions first, then data — reversing that order wastes the cleanup.

How do I find duplicate contacts in HubSpot?

Start with HubSpot's built-in duplicate management, then manually check for personal versus work emails on the same person, company name variants like Ltd versus Limited, and records created by an integration that keys on a field your forms never collect.

Which record should survive a merge?

The one with the richest activity timeline and associations, because that history is what your reporting reads. Property values from the secondary record can be filled in afterwards; a lost engagement history cannot be reconstructed.

What is the difference between marketing and non-marketing contacts?

Marketing contacts are the ones you can email and target, and they determine your Marketing Hub tier. Suppliers, applicants, partners and dead lists should be non-marketing — set that as the import default and promote contacts only when a campaign needs them.

How often should we clean our CRM data?

A quarterly 90-minute audit covering duplicates, unowned properties, stalled deals and seat usage keeps a portal healthy. Cleanup done routinely is trivial; the same work postponed for two years becomes a project.

Should I delete unused properties?

Archive rather than delete when historical values still matter, and rename dead-but-needed fields with a clear prefix so nobody filters on them by mistake. Genuinely unused, never-populated properties can go.

What do I do with deals that have been open for months?

Build a list of open deals with no activity for 90 days, review with the owner, and close as lost with a required reason. Forecast accuracy improves immediately, and the loss reasons become the most useful data set you collect all year.

How do I stop bad data coming back?

Close the source: required form fields, restricted dropdowns instead of free text, imports with a proper matching key, and integrations set to update rather than create. Then add hygiene workflows to normalise values and flag gaps automatically.

Does data cleanup affect our HubSpot bill?

It can, meaningfully. Marketing contact tiers price in blocks, so removing records that never needed marketing often drops a portal a full tier at renewal with no loss of capability.

Can Hubstack do a HubSpot data cleanup for us?

Yes. We run the sequence in this article — definitions, property audit, deduplication, marketing status, deal review, dashboard rebuild — then leave you with documented property ownership and hygiene workflows so it stays clean.

Proof

What this looks like when it's done right.

Want reports your leadership will actually quote?

Related

Technical SEO services

Redirect mapping, metadata baselines, canonical configuration, schema and indexing controls — handled as an ongoing programme rather than a launch-week scramble.

HubSpot technical SEO
Keep reading

What Actually Breaks SEO During a HubSpot Migration (And How We Prevent It)

Rankings rarely drop because you moved to HubSpot. They drop because of five specific technical failures during cutover — redirects, metadata, canonicals, schema and indexing settings.

Read article