78% fewer duplicate records. That’s the headline number vendors like Improvado and Hightouch are putting in front of CMOs right now, and it’s the kind of stat that makes procurement teams sit up and marketing ops teams get suspicious. Both reactions are correct. Identity resolution has quietly become the most consequential — and least understood — layer of the martech stack, and the 78% figure deserves a harder look before it lands in your next budget deck.
If your customer data platform still can’t tell you that “[email protected],” “Jane S.,” and device ID 8f3a-22b belong to the same person, you’re overspending on acquisition and underspending on retention. That’s the pitch. The question is whether the math behind it holds up.
What “78% Deduplication” Actually Means
Deduplication rate sounds simple. It isn’t. Vendors calculate it differently, and the difference changes what the number tells you.
Improvado’s claim, based on its marketing data pipeline product, refers to the reduction in duplicate contact and account records after running its AI matching layer across CRM, ad platform, and CDP exports. Hightouch’s number comes from a different angle: identity resolution inside its reverse ETL and audience-building workflow, where the same person shows up across Salesforce, a data warehouse, and multiple ad accounts under slightly different identifiers.
Both are measuring reduction from a “before” state to an “after” state. But the before state is the variable that matters most. If a brand’s baseline data hygiene is poor — say, five ad platforms feeding a warehouse with no normalization — then almost any matching layer will show dramatic gains. Reducing a genuinely dirty dataset by 78% is very different from reducing an already-decent dataset by 78%.
A 78% deduplication rate on a messy, unmanaged dataset is not the same claim as 78% on a dataset that already had basic hygiene rules in place. Ask which one you’re being sold.
This isn’t a knock on either vendor specifically. It’s a structural issue with how the entire identity resolution category reports performance. For more on how these headline numbers get built and where they break down, see our vendor due-diligence guide on match rate claims.
The Mechanics: How AI Matching Actually Works Here
Both platforms use probabilistic matching layered on deterministic rules. Deterministic matching relies on exact-match identifiers: email hash, phone number, customer ID. It’s reliable but limited — it only catches duplicates where a shared identifier exists.
Probabilistic matching is where the AI claim comes in. Machine learning models score the likelihood that two records represent the same person based on fuzzy signals: name variations, address proximity, device fingerprints, behavioral overlap, timing patterns. This is where the big deduplication gains show up, and it’s also where false positives sneak in.
Here’s the tension nobody puts on the sales slide: the more aggressive the probabilistic threshold, the higher your deduplication rate climbs — and the higher your risk of merging two different people into one profile. A vendor optimizing for an impressive dedup percentage has a built-in incentive to loosen the matching threshold. That’s not fraud. It’s just how the incentive structure works when “reduction rate” is the marketed KPI instead of “match accuracy.”
Improvado leans on its AI layer primarily for marketing data normalization across ad platforms — think matching a Meta campaign ID to a HubSpot contact to a Shopify order. Hightouch’s resolution engine sits closer to the warehouse, which means it has access to more first-party behavioral data but less native visibility into walled-garden platform IDs. Different architectures, different blind spots.
Where the ROI Case Gets Real (And Where It Doesn’t)
Let’s talk numbers that matter more than the headline stat. Duplicate records inflate ad spend by causing overlapping audience targeting — you’re paying to reach the same person twice under two different identities. They also corrupt attribution models, because a conversion event gets split across two “different” customer journeys instead of consolidating into one.
eMarketer has repeatedly flagged data fragmentation as a top-three constraint on marketing efficiency, and internal audits at mid-size retailers routinely find CRM duplication rates between 15% and 30% before any resolution tooling is applied. So a 78% reduction against that baseline is meaningful — if it’s real and if it’s measured correctly.
Where the ROI case gets shaky: neither vendor publishes independently audited case studies with before/after attribution accuracy, only internal customer testimonials. That’s standard for the industry, but it means the burden of verification falls on your team, not the vendor’s marketing department.
Ask for the raw counts, not just the percentage. A 78% reduction from 100,000 duplicate flags to 22,000 tells a very different operational story than 78% reduction from 500 to 110. Scale changes what the number costs you to fix and what it’s worth once resolved.
Compliance Is the Part Everyone Skips
Merging identity records isn’t just a data hygiene exercise. It’s a regulatory exposure point. Once you’ve matched “anonymous device ID” to “named customer with purchase history,” you’ve created a new PII record, and that record falls under whatever consent framework governed the original data collection.
The FTC has been increasingly active on data broker and identity matching practices, and the ICO in the UK treats probabilistic identity resolution as a form of profiling that can trigger additional disclosure obligations under UK GDPR. If your matching engine is merging records across consent boundaries — say, a user who opted into email marketing but not into cross-device tracking — you’ve got a compliance problem wrapped inside an efficiency win.
This is where governance has to catch up to the tooling. Our identity-based attribution governance piece walks through the specific controls brands need before turning on aggressive matching thresholds. If your legal and privacy teams aren’t in the room when you evaluate Improvado or Hightouch, get them there before contract signature, not after.
Comparing the Two Head to Head
Improvado’s strength is breadth. It connects to over 500 marketing and sales data sources, which means its deduplication claim is being tested across a genuinely wide surface area. That’s good for brands running a sprawling channel mix — paid social, affiliate, retail media, CRM — all feeding into one reporting layer.
Hightouch’s strength is depth within the warehouse. Because it operates as a reverse ETL tool syncing warehouse data out to activation platforms, its identity resolution benefits from whatever modeling your data team has already built in Snowflake, BigQuery, or Databricks. If your warehouse data is well-modeled, Hightouch’s matching has a stronger foundation to work from. If it’s not, you’re asking the AI to compensate for a modeling gap, and that’s a heavier lift than either vendor’s pitch deck suggests.
- Choose Improvado if: you need cross-channel marketing reporting normalized fast, with less warehouse maturity required upfront.
- Choose Hightouch if: your data team already has a solid warehouse model and you want resolution to feed directly into activation (ad audiences, personalization, lifecycle triggers).
- Pressure-test either vendor by: requesting a sample match-rate audit on a subset of your own data before full deployment, not just relying on their published benchmark.
This dynamic isn’t unique to these two vendors. We’ve seen the same pattern comparing Wunderkind, Cordial, and Klaviyo on identity resolution claims: the vendor with the most controlled data environment usually posts the cleanest match rates, and the vendor with the broadest integration footprint usually posts the biggest percentage gains. Neither is lying. They’re just measuring different problems.
What to Actually Ask Before You Sign
Skip the demo theater. Ask these questions in the vendor call:
- What’s the raw duplicate count before and after, not just the percentage?
- What confidence threshold does the matching model use, and can we adjust it?
- How are false-positive merges detected and reversed?
- Does the resolution process respect existing consent flags at the record level?
- Can you provide a match audit on our data before full contract commitment?
Any vendor confident in their number will answer all five without hesitation. If you get vague answers on question three (false-positive detection), that’s your signal to slow down. Merging two distinct customers into one profile is a silent failure — it doesn’t throw an error, it just quietly corrupts your CRM and your attribution model at the same time. For a deeper look at how attribution accuracy and identity resolution interact, our multi-touch vs algorithmic attribution comparison is a useful companion read alongside this evaluation.
Consider also how HubSpot’s research on data quality and CRM hygiene frames the downstream cost: bad data doesn’t just waste ad spend, it erodes trust in every dashboard your leadership team looks at. That trust is expensive to rebuild once a false-positive merge surfaces in a board report.
The Real Takeaway
The 78% figure isn’t fake, but it’s incomplete without context on baseline data quality, matching threshold aggressiveness, and false-positive controls. Before you cite that number in a budget request, get your own sample audit run against your actual data, and put your privacy team in the vendor evaluation from day one, not after the contract’s signed.
Frequently Asked Questions
Is the 78% deduplication claim from Improvado and Hightouch verified by a third party?
No independent, publicly audited study confirms this figure for either vendor. Both numbers come from internal benchmarking and customer case studies. Brands should request a sample audit against their own data before relying on the published percentage.
What’s the difference between deterministic and probabilistic identity matching?
Deterministic matching links records using exact identifiers like email or customer ID. Probabilistic matching uses AI models to score the likelihood that two records represent the same person based on fuzzy signals like name variations or device behavior, which increases match rates but also increases false-positive risk.
Can aggressive deduplication create compliance risk?
Yes. Merging records across different consent boundaries — for example, combining an anonymous device profile with a named customer record — can trigger obligations under frameworks enforced by regulators like the FTC and the ICO. Legal and privacy teams should review matching thresholds before deployment.
How is deduplication rate calculated, and why does baseline data quality matter?
Deduplication rate measures the reduction in duplicate records from a “before” state to an “after” state. A high percentage reduction on a poorly maintained dataset is far less impressive than the same percentage on an already-clean dataset, so raw counts matter as much as the percentage itself.
Should brands choose Improvado or Hightouch for identity resolution?
It depends on data infrastructure. Improvado suits brands needing broad cross-channel normalization with less warehouse maturity. Hightouch suits brands with a well-modeled data warehouse who want resolution to feed directly into activation platforms.
Top Influencer Marketing Agencies
The leading agencies shaping influencer marketing in 2026
Agencies ranked by campaign performance, client diversity, platform expertise, proven ROI, industry recognition, and client satisfaction. Assessed through verified case studies, reviews, and industry consultations.
Moburst
-
2

The Shelf
Boutique Beauty & Lifestyle Influencer AgencyA data-driven boutique agency specializing exclusively in beauty, wellness, and lifestyle influencer campaigns on Instagram and TikTok. Best for brands already focused on the beauty/personal care space that need curated, aesthetic-driven content.Clients: Pepsi, The Honest Company, Hims, Elf Cosmetics, Pure LeafVisit The Shelf → -
3

Audiencly
Niche Gaming & Esports Influencer AgencyA specialized agency focused exclusively on gaming and esports creators on YouTube, Twitch, and TikTok. Ideal if your campaign is 100% gaming-focused — from game launches to hardware and esports events.Clients: Epic Games, NordVPN, Ubisoft, Wargaming, Tencent GamesVisit Audiencly → -
4

Viral Nation
Global Influencer Marketing & Talent AgencyA dual talent management and marketing agency with proprietary brand safety tools and a global creator network spanning nano-influencers to celebrities across all major platforms.Clients: Meta, Activision Blizzard, Energizer, Aston Martin, WalmartVisit Viral Nation → -
5

The Influencer Marketing Factory
TikTok, Instagram & YouTube CampaignsA full-service agency with strong TikTok expertise, offering end-to-end campaign management from influencer discovery through performance reporting with a focus on platform-native content.Clients: Google, Snapchat, Universal Music, Bumble, YelpVisit TIMF → -
6

NeoReach
Enterprise Analytics & Influencer CampaignsAn enterprise-focused agency combining managed campaigns with a powerful self-service data platform for influencer search, audience analytics, and attribution modeling.Clients: Amazon, Airbnb, Netflix, Honda, The New York TimesVisit NeoReach → -
7

Ubiquitous
Creator-First Marketing PlatformA tech-driven platform combining self-service tools with managed campaign options, emphasizing speed and scalability for brands managing multiple influencer relationships.Clients: Lyft, Disney, Target, American Eagle, NetflixVisit Ubiquitous → -
8

Obviously
Scalable Enterprise Influencer CampaignsA tech-enabled agency built for high-volume campaigns, coordinating hundreds of creators simultaneously with end-to-end logistics, content rights management, and product seeding.Clients: Google, Ulta Beauty, Converse, AmazonVisit Obviously →
