ai-sdroutboundabmdeliverability

Do AI SDRs work? What the public record says, and why we do not use them

The AI SDR category produced more raw replies, not fewer. We still do not use it, because the throughput it sells is not available in account-based work and the reputation it spends is your domain, not the vendor's. With the reporting, including the parts that cut against us.

September 14, 2026·9 min read·Draftship

We do not use AI SDRs on client work. Not because the software writes badly, and not because the vendors are frauds. We do not use them because the category was built for a volume problem and our clients bring us a precision problem. The public record from the last two years is now detailed enough to argue from instead of guess at, so here is the argument with the receipts attached, including the parts that cut against us.

The number the category sold on, and the number underneath it

The clearest summary of what happened to outbound after AI arrived sits in Salesmotion's 2026 AI SDR comparison: per-rep monthly outbound volume jumped from roughly 1,150 emails to 7,400 after AI adoption, raw reply rates fell from 4.7% to 2.9%, and about 1 in 6 emails now never reaches the inbox.

Most people quoting those figures stop at the reply rate, and they should not. Run the arithmetic. 1,150 emails at 4.7% is about 54 replies a month. 7,400 emails at 2.9% is about 215. On raw reply count the volume play wins, and it wins by a lot. Anyone telling you the AI SDR wave simply did not work is skipping that line, and we are not going to skip it.

The case against the category is not that it produces fewer replies. It is what it spends to get them, and whether that spend is available to you at all.

Six times the volume is not on the menu in account-based work

If your program is 120 named accounts with five or six people worth contacting at each, the entire addressable universe is around 700 humans. You cannot send 7,400 emails a month into that. You can only send the same people the same thing more often, which is the single variable reply rate is most sensitive to.

This is the part of the pitch that quietly does not apply to account-based programs. The machine's advantage is throughput, and in ABM throughput is capped on purpose. Salesmotion's agency guide puts the working ceiling plainly: a team of two to three marketers with the right tools can run effective ABM programs for 50 to 200 target accounts, and the gap is usually intelligence infrastructure rather than execution headcount. Buying a throughput machine for a job with a throughput ceiling is buying the wrong tool competently.

What the record says about the vendors, and what it does not

Refonte Learning's comparison of Clay, 11x and Artisan assembles the reporting on 11x in the order it deserves to be read:

  • Rework's June 3, 2026 analysis points to 11x after it raised more than $74 million from Andreessen Horowitz and Benchmark, describing substantial customer churn as a warning against assuming autonomous replacement produces durable returns.
  • TechCrunch reported in March 2025 that ZoomInfo ran a one-month trial and did not proceed, because it judged 11x's performance below that of its own SDR employees. Airtable told TechCrunch that its brief trial did not progress to production.
  • Sifted later reported sources alleging extremely high churn during part of 2024, with internal retention figures of roughly 20% to 30% in that period.
  • 11x disputed those allegations, described them as inaccurate, and said it had hundreds of customers receiving value.

That fourth bullet is not a formality. The Refonte write-up is explicit that the historical churn claims should not be presented as undisputed current retention data, and the same piece flags that Rework's estimate of 50% to 70% annual churn across the AI SDR category is attributed broadly to 2026 sales-tech analyses rather than to a transparent primary dataset. We are repeating those cautions because a post that quotes only the damaging half of a source is not evidence, it is selection.

What survives the caution is still substantial: two named companies whose pilots did not convert, a reported churn controversy, a vendor denial, and Gartner's forecast that more than 40% of agentic AI projects will be canceled by the end of 2027 for exactly this pattern of activity without business value.

We also cut three claims from the first draft of this post, because we went looking for the sources and could not stand them up. If we cannot link it, we do not print it.

The bill arrives at your domain, not the vendor's

The same Refonte comparison cites Digital Applied's synthesised sender data: 47% of attempted AI SDR deployments hit a domain-reputation problem within 90 days, with 18.7% of AI SDR mail landing in spam in Microsoft 365 environments against 7.8% in Google Workspace.

If you sell to businesses, the Microsoft number is the one describing your buyers. And it lands against a wall that got higher this year: compliant senders average about 89% inbox placement in 2026 while non-compliant senders see 22% to 34% of their mail routed to spam. We wrote the full set of checks up in the 2026 sender gates.

The asymmetry is the argument. The vendor's reputation is not at stake in your sending. Your domain is, and a damaged domain is not a campaign you rerun next quarter. It is the address your invoices and your renewal notices go out from.

Clay is a different thing and it is not fair to lump it in

Clay gets named in these conversations constantly and it does not belong in the same bucket. Refonte describes it as unusually strong as a prospect-data, enrichment, research and orchestration layer, while noting it now supports native email campaigns too, and Salesmotion's comparison calls it not an AI SDR out of the box.

That distinction is the one we actually care about, and it is not the distinction people think they are drawing. The question is not whether AI is in the stack. It is whether software decides who to contact and then presses send. Enrichment, research and orchestration all stop short of that line. An autonomous SDR is defined by crossing it.

What we do instead

The short version: the machine reads, a person writes, and nothing leaves the building unread.

  • The list is small and named. We work in the 50 to 200 range that the agency guide describes as runnable by two or three people, and we write down why each account is on it. That method is its own post, how to build an account list you can defend.
  • Research is generated, verified, then used. We will happily have a model read an annual report, a hiring page and a quarter of press coverage. A person checks the three facts that will actually appear in the email against the source before anyone drafts anything.
  • The sent sentence is written by a person. First drafts can come from the client's own words, their sales calls and their docs. The version that leaves is edited by someone who will be on the reply.
  • Volume is capped by the list, not by the tool. We do not have a throughput target, so we have nothing to spend reputation on.
  • The client owns the domain and the data. Including the replies that say no, which are the most useful thing a program produces in its first quarter.

Where we do use AI, and exactly where we stop, is set out in what AI is actually good at in ABM.

What we were slow to say

For most of 2025 our answer to "do you use AI SDRs" was a diplomatic "it depends on your motion". It does not depend. We hedged because the category was well funded and we did not want to sound like we were behind it, which is a bad reason to be vague with someone paying you for a recommendation. The honest answer, then and now, is that we would not run one on a domain we were responsible for.

None of this is a claim that autonomous outbound never works. It is a claim about which problem it solves. If the constraint on your pipeline is that you are not contacting enough people, a throughput machine is a reasonable thing to evaluate, and you should read the vendors' own defences alongside the reporting above. If the constraint is that the forty accounts that matter do not know who you are, more sending is not the fix and buying it will cost you the domain you need for the fix.

If you are weighing this for your own program, tell us what you are trying to fix and we will tell you what we would do.

Tell us what you are trying to fix.

Four fields, one open question, and a reply from a person within one working day. If we are the wrong people for it we will say so and point you at what we would do instead.

Start a conversation

Or take the email editor and the forty templates, free and with no account.