What AI is actually good at in ABM, and where we stop using it
We use AI every day and never on the sentence that gets sent. The line is not about writing quality, it is about who pays when the output is wrong and nobody notices. Where we point it, where we do not, and the one test we use to decide.
We use AI every working day and never on the sentence that gets sent. That is not a compromise position and it is not squeamishness about the writing quality. The line falls where it does because of one question: when the output is wrong, who finds out, and how long does it take. Where being wrong is cheap and obvious, we automate freely. Where being wrong is expensive and invisible, a person does the work.
The shape of the buying process decides what AI should be pointed at
Sit with the implication, because it inverts the usual instinct. If most of the evaluation happens before you are in the conversation, the scarce resource is not message production. It is knowing what already happened. More emails do not buy you a seat in a process that ran without you. Knowing which account is in that process right now, and roughly what it is reading, does.
So the useful question is not "how do we make more with AI". It is "what could we know that we currently do not", and almost everything in that second category is a reading and sorting problem. Which is exactly what this technology is good at.
Where we use it, and why it is safe there
Reading things nobody has time to read. Annual reports, earnings call transcripts, every job posting a company has open, three years of product release notes, a competitor's documentation. Summarising a sixty-page filing down to the four things that bear on whether this account has our problem is the single highest-return use in the whole program. If the summary is wrong, we find out in the thirty seconds it takes to check the claim against the page it came from.
Spotting the account whose situation changed. Three types of intent data exist, and first-party intent from your own properties is the highest quality: fully attributable, in your control, and reflecting real engagement. The practical form is a scoring rule, and the same source gives a worked example: a pricing page visit at 30 points, an intent topic surge at 20, a blog post view at 5, with 50 points inside seven days raising an alert, against a narrow set of 5 to 15 intent topics that correlate with purchase. Note what that is doing. It sorts a queue. A person still decides what to do with the account at the top of it.
Turning the client's own words into a first pass. Their sales calls, their support threads, their internal docs, the emails that already worked. A model is good at extracting how a company actually describes its own product, which is usually sharper than its marketing site and always sharper than what an agency invents in week one.
Producing the artifact nobody wants to make. A consistent one-page research brief per account, in the same shape every time, so the person writing the email is reading a familiar document rather than twelve browser tabs. Consistency is genuinely hard for humans over two hundred accounts and trivial for a machine.
Structural work. Normalising job titles, mapping a buying committee, deduplicating a list, finding which of forty accounts have a named owner for the category. All checkable, all boring, all fine.
The pattern across all five: verification is faster than production. That is the test.
Where we stop
The email that gets sent. This is the one people push back on, so here is the evidence rather than the preference. After AI adoption, per-rep monthly volume went from roughly 1,150 emails to 7,400 while raw reply rates fell from 4.7% to 2.9%, with about 1 in 6 emails never reaching the inbox. And the bill lands on your infrastructure: 47% of attempted AI SDR deployments hit a domain-reputation problem within 90 days, with 18.7% of that mail going to spam in Microsoft 365 against 7.8% in Google Workspace. We took the longer version of this argument apart in why we do not use AI SDRs.
Note that neither of those numbers is about the prose being bad. They are about what happens when sending gets cheap. Bain Capital Ventures makes the substantive version of the point, that complex outbound still depends on timing, trust, tone, objection handling and contextual judgement, and its recommended modern sales team structure keeps humans in it.
Deciding who to target. A model will produce a confident, well-formatted account list. It will be built from the surface features of companies that resemble your existing customers, which is a different thing from the accounts that will buy, and the difference will not be visible for two quarters. The list is the whole program, and it is the one thing we build by hand.
Anything where being wrong is expensive and invisible. This is the actual rule, and the two above are instances of it. An incorrect summary is cheap and obvious. An email that confidently misstates a prospect's funding stage is expensive and invisible: it does not bounce, no dashboard flags it, and the only signal is a senior person at a named account quietly deciding you did not do your homework. You never see that. You just see a conversion rate.
Anything the recipient can check faster than we can. If we are asserting something about their business, one of them knows whether it is true. That asymmetry means we have to be right, not plausible.
The test, stated once
Before automating any step, ask: how long does it take a person to verify this output, compared with how long it took to produce?
If verification is faster, automate it and verify every time. If verification is slower, or if nobody will actually do it, do not automate it, because what you have built is a machine for generating unchecked claims at scale. The entire boundary in this post falls out of that single question, which is why we can apply it to tools that did not exist when we wrote this down.
There is a market-level version of the same warning. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027, for the pattern of activity without business value. Activity without business value is precisely what you get when production is automated and verification is not.
Where the tools sit on the line
The useful distinction is not AI versus no AI. It is whether software decides who to contact and then presses send. Enrichment, research and orchestration platforms stop short of that: Refonte describes Clay as unusually strong as a prospect-data, enrichment, research and orchestration layer, noting it now supports native campaigns too, which is exactly where a buyer has to look at the configuration rather than the category. An autonomous SDR is defined by crossing the line. Most of the stack is defined by not crossing it, and gets tarred anyway.
What we got wrong
We automated the research brief before we automated the reading, which meant our people were checking a tidy formatted document instead of the source it came from. A well-laid-out summary is more persuasive than a messy one and no more accurate, and for a couple of months that formatting did our checking for us. Now the brief carries the link next to each claim, and the rule is that anything appearing in a sent email gets opened first.
If you want to know where we would draw this line inside your program, tell us what you are trying to fix.
Tell us what you are trying to fix.
Four fields, one open question, and a reply from a person within one working day. If we are the wrong people for it we will say so and point you at what we would do instead.
Or take the email editor and the forty templates, free and with no account.