Every few weeks someone tells me Clay got too expensive and they're thinking of ripping it out. Almost none of them have a pricing problem. They have an architecture problem. They opened the tool, started stacking enrichment columns, dragged in every provider they could find, and hit run on the whole table. Clay charges credits per enrichment call. Aimless clicking is, quite literally, spending money.
The fix isn't a cheaper tool. It's treating enrichment like a system with a defined sequence, hard gates, and cost controls, instead of a spreadsheet you poke at until emails appear. Here's the architecture we install for clients, and the mistakes it's designed to prevent.
The Default Setup Is the Expensive Setup
The out-of-the-box way people use Clay is: one table, every contact, every provider, run it all. That's the most expensive possible configuration and usually not even the most accurate one. A CRM without enriched data is a glorified Rolodex. But enriching everything, blindly, just means you're paying full price to confirm data you already had, on people you were never going to contact, with providers that don't cover your ICP.
Five changes turn that around.
The cheapest enrichment is the one you never run. Before a single credit is spent, the list gets checked against what's already in the CRM. We split this into a Sources Table (minimal identifiers only, deduped first) feeding a Data Table that does the actual enrichment. Match companies on domain, contacts on email. If the domain already exists in HubSpot, skip the company entirely. If the email exists, skip the contact but keep the company if it's net-new.
On one mid-market build the client's HubSpot already held ~3,700 companies and ~18,600 contacts. Enriching a new list without checking against that first would have meant paying to re-enrich thousands of records we already owned. Dedupe is not a cleanup step you do at the end. It's the first gate, and it's free.
"Waterfall enrichment" gets thrown around like it means "use lots of providers." It doesn't. A waterfall means providers run in sequence, and each one only fires if the previous one came back empty. Order matters enormously, because you pay per lookup. Put the cheapest, highest-hit-rate provider for your specific ICP first, and stop the moment you get a valid hit.
We ran a controlled cost test on a mid-market fashion-retail list in Europe: six contacts, full waterfall. The numbers make the point better than I can:
- PDL (People Data Labs): ~23.5 credits for the batch, 0 emails found for that ICP. Most expensive, worst performer.
- Hunter: ~1.5 credits, 4 of 6 emails found. Cheapest, best hit rate.
Run blindly, the full stack burned ~40 credits for six contacts. Ordered correctly (cheapest and most effective first, stop-on-first-hit) the same job lands around 5 to 15 credits per company. An earlier single-provider run with no waterfall logic cost 227 credits for 25 accounts and 70 contacts. Same data, an order of magnitude apart in spend, purely from sequencing.
The lesson isn't "Hunter good, PDL bad." It's that provider performance is ICP-specific, so you benchmark on your own list, rank by cost-per-successful-hit, and never let an expensive provider run first "just in case."
Discovery vs. reveal
Same discipline applies to people discovery. Use the free search step to find names and LinkedIn URLs; only pay for the reveal on the records that survive your filters. Treat certain providers as names-and-LinkedIn only and don't trust their emails or phone numbers for your region without verification. A provider that's great for US SaaS can be near-useless for European operations roles.
Not every record deserves the same spend. Before deep enrichment, every record gets a tier based on ICP fit: we surface these in the CRM as Hot / Warm / Consider / Ignore, and the tier rides along as a property through the whole pipeline. Ignore-tier records don't get the full waterfall. Hot-tier records get the deep, more expensive treatment because they're the ones your reps will actually work.
This is the single biggest lever on cost, because the default behavior (enrich everyone identically) spends your most expensive credits on the accounts least likely to convert. Enrich-by-tier flips the ratio. Your budget follows your ICP instead of your row count.
Here's a rule we never break: nothing writes to HubSpot without a verified email. Enrichment is not verification. Most providers hand you an email; they don't confirm it's deliverable. Clay-native enrichment finds the address, then a separate verification step, using LeadMagic, NeverBounce, or ZeroBounce, validates it before anything is allowed into the CRM.
Skip this and you don't just get messy data. You torch the client's sender reputation the day they start outbound. Unverified emails bounce, bounces tank your domain, and now the whole motion underperforms because of a step someone cut to save five minutes. The verified-email gate is a hard conditional in the pipeline, not a nice-to-have.
Clay is a workbench, not a warehouse. Leaving thousands of enriched rows sitting in tables is both a cost and a liability. Our pattern: enrich, push the clean record to the CRM, write the raw signals back to a store outside Clay (we use Supabase, keyed on LinkedIn URL plus source channel), then clear the Clay table. You keep every signal you paid for, you can re-use it without re-enriching, and you're not paying to warehouse data in the tool least suited to it.
Guardrails: Caps, Validation Columns, and Honest Signals
A few controls that keep the system from quietly bleeding credits or shipping garbage:
- Spending caps. For any large run, set a hard cap at roughly 150% of your estimated cost. You avoid runaway spend without cutting legitimate runs short.
- Validation columns. Carry explicit guardrail fields (size_valid, title_valid, region_valid) and flag anything guessed from an email pattern rather than confirmed. Bad enrichment that looks confident is worse than no enrichment.
- Honest signals. If a propensity signal (funding, expansion, new leadership, a competitor tool in use) can't be tied to a verifiable public source, it's marked unknown, never no, and never yes without a source in the notes. A signal you can't defend is noise you're paying to generate.
What Good Looks Like
Put together, the architecture is boring in the best way: dedupe against the CRM first, run an ICP-ordered waterfall that stops on first hit, spend deep only on high tiers, gate every write on a verified email, and warehouse signals outside the tool. The target we hold ourselves to is 90%+ coverage on the fields that actually matter, at a fraction of the credit spend of the run-everything approach.
Most GTM problems are broken systems, not bad tools. Clay isn't expensive. Enriching without an architecture is.
Fix the system that spends the credits, and the bill takes care of itself.
Burning credits and not sure where the spend is going?
Our Data Enrichment & List Building module installs this exact pipeline (ICP-ordered waterfalls, enrich-by-tier, and a verified-email gate) with a 90%+ ICP coverage guarantee. If your CRM data is already a mess underneath, start with the CleanOps Protocol cleanup, or run the CRM Decay Calculator to see what the bad data is costing you per month.
Let's talk about what your enrichment spend should actually look like.