Lead-gen buyers face TCPA risk as study finds 27% of unwanted calls open with machines

An arXiv preprint reports that at least 26.9% of unwanted inbound calls open with machine voices, with synthetic speech concentrated in lead-generation spam rather than fraud. The paper also finds off-the-shelf detectors are inconsistent, complicating how marketers, carriers and compliance vendors p

Hannah Vogel ·

Lead-gen buyers face TCPA risk as study finds 27% of unwanted calls open with machines

In an arXiv preprint posted in September, the authors report that at least 26.9% of unwanted inbound calls in their sample opened with a machine voice, as measured by a disclosed pipeline combining audio fingerprinting, a commercial synthetic-speech detector, and blinded human listeners. The study is not peer reviewed. It places a business problem squarely on the desk of marketing, procurement and compliance leaders who buy or oversee outbound lead-generation: the Federal Communications Commission classified AI-generated voices under the Telephone Consumer Protection Act (TCPA) in February 2024, the preprint notes, yet the dominant tools to identify synthetic voices are inconsistent and the behavior concentrates in lead-gen campaigns, not just outright fraud.

The study’s pipeline makes a concrete claim about who is using synthetic voices and where they appear

The researchers instrumented an interactive voice honeypot—language-model personas on real U.S. numbers, recording the caller on its own track—and captured 10,987 calls over 66 days, setting aside 11 days when the system answered silently. Of 7,233 calls the persona greeted on normal days, 13.8% opened with a recording that also appeared on other calls, and 13.1% opened with fresh audio labeled synthetic by a commercial detector. A further 9.9% opened with a caller who never spoke after the greeting, 54.2% with fresh audio the detector labeled human, and 9.0% could not be scored. By the study’s construction, machine-voiced openings are therefore at least 26.9%, with an additional tenth of calls presenting as silent connections the authors interpret as machine-placed. The paper states synthetic openings concentrate in lead-generation spam (33.8%) more than in fraud (21.1%), and that only 0.44% of the calls disclosed automation.

For operators, the location of the behavior matters more than the topline rate. If synthetic openers cluster in lead-gen, then brand and performance marketing budgets are underwriting the behavior—and the TCPA risk lives in master service agreements and affiliate contracts, not just in security or fraud teams. The paper’s finding that campaigns outlast numbers—one recorded compliance notice opened calls in six campaigns, and one synthetic voice served nine—points to reusable assets and networks of vendors, not one-off bad actors.

Detectors disagree with themselves, and humans only confirm half of what the model flags

The preprint is explicit about instrument limits. Across scored openings (6,192 of the calls), the detector labeled 29.3% as synthetic. Replays of a recording made up 45% of the detector’s own synthetic rate, according to the paper’s calculation, underscoring how much “AI voice” in the wild is not a model generating new speech but a recorded asset redeployed across campaigns. More concerning for compliance proofs, the same waveform played on two separate calls fell on opposite sides of the detector’s threshold 13.6% of the time, and eleven blinded listeners confirmed only 54.4% of the detector’s flags.

That creates a practical problem. Carriers, call analytics providers and enterprise compliance teams have sold AI voice detection as a filter and a shield; the paper suggests neither function is reliable enough to carry legal weight on its own. If one and the same audio file can be both above and below threshold depending on context, automated blocking and after-the-fact attestations look shaky. For marketers who rely on downstream vendors to certify compliance, “detector says human” will not survive discovery if the same vendor’s tool calls it synthetic on Tuesday and human on Wednesday.

The risk sits in lead-gen procurement, not just with telcos or fraud teams

The paper’s segmentation is the most operationally relevant data point: “Synthetic openings concentrate in lead-generation spam (33.8%), not fraud (21.1%).” In plain terms, the campaigns using AI voice are the same ones marketing budgets commission to fill pipelines and the same networks that resell and recycle calls. The study also finds that prevalence tracks how long a bait number has circulated—59% versus 19% in the same weeks—implying that list age and circulation through lead markets, not calendar time, explains the rising trend on some numbers.

If you buy calls or leads, that means your risk is not linear with spend; it is nonlinear with how your vendors source and recycle numbers. Buyers who specify “TCPA-compliant live-agent calls only” without auditing number provenance, script assets and subaffiliate behavior could be marketing into a synthetic-first supply chain. With only 0.44% of calls disclosing automation in this sample, disclosure as a safe harbor is not showing up in practice.

Silent connections are not dead air—they are a machine cost that becomes your liability

The authors classify 9.9% of greetings as followed by a caller who never speaks, a pattern they read as machine-placed connections. Marketers and contact-center operators will recognize the behavior: high-volume dialers probing for live answers, then routing to an agent or dropping. Under the FCC’s view of AI voice within the TCPA, and with increasing scrutiny of call completion and attestation, these silent connects are part of the same machine-placed universe. They also drive consumer complaints and brand damage before a script even begins. If your attribution stack optimizes to “live answer” or “agent connect,” the instrument may reward the very vendors whose machine-driven connects load the top of the funnel with low-quality, high-risk attempts.

Campaign assets are recycled across vendors; fingerprinting exposes the network—and raises the bar for deniability

Audio fingerprinting in the study surfaced two kinds of reuse: one recorded compliance notice reused across six campaigns, and one synthetic voice serving nine. For compliance officers and outside counsel, that cuts both ways. On the one hand, repeated assets allow brands and carriers to flag and block faster at scale. On the other, if a brand’s approved script or compliance notice is found across distinct campaigns, it becomes harder to argue that a downstream affiliate acted entirely without direction. The reuse pattern suggests a supply chain for sound itself: voices and notices are content, passed from one shop to another. A brand that cannot map who holds which asset will struggle to police its own exposure.

The obvious counter: this is a honeypot, not your customer base—does it generalize?

Skeptics will point out that the data come from an interactive voice honeypot, not a consumer panel, and that 11 days of silent answering were excluded. The authors themselves present their method and denominators, and the numbers are the paper’s own. The stronger objection for enterprise readers is not whether 26.9% is the magic number for their vertical, but whether the instrument’s inconsistencies undermine any enforcement or filtering at all. The detector’s threshold flips on identical waveforms 13.6% of the time, and human confirmation hits 54.4%—figures that argue for corroborating evidence, not for ignoring the behavior. If anything, the honeypot’s ability to capture and compare repeated assets across campaigns is the point: repetition, not individual calls, is what procurement and compliance can act on.

What changes now for marketers, carriers and compliance vendors

For marketers who commission lead-gen calls, the near-term change is contractual. If synthetic openings concentrate in lead-gen and detectors are inconsistent, then the enforceable handle is upstream: number provenance, asset inventories (recorded voices and compliance notices), and a requirement to disclose automation where used. Buyers will need to specify audit rights over subaffiliates and insist on asset-level logs, not just aggregate “human-only” attestations. Expect more brands to pause or restructure pay-per-call arrangements until they can see both the content and the circulation path of the calls they buy.

For carriers and call analytics providers, the business implication is product posture. “AI voice detection” marketed as a binary filter will meet pushback if it cannot demonstrate stable thresholds across contexts. Tools that combine fingerprinting of repeated assets with probabilistic scoring, and that expose error bands rather than one label, will be more defensible to enterprise customers and regulators. The paper’s finding that campaigns outlast numbers suggests that network-level asset maps, not per-call judgments, are the product marketers will pay for.

For compliance vendors, the sales motion shifts from promising clean rooms to selling process: corroboration across instruments, human audit on edge cases, and a playbook that recognizes silent connects as part of machine-placed behavior. With only 0.44% of calls disclosing automation in the sample, vendors who build and measure disclosure rates in scripts will have a clearer, enforceable KPI to sell into legal and brand risk teams.

Signals to watch over the next two quarters

If the paper’s segmentation is right, expect contract language in lead-gen to evolve first: disclosures of automation and explicit prohibitions on synthetic openers without consent. Watch for any public enforcement or demand letters that cite synthetic speech under the TCPA, which would validate that regulators and plaintiffs’ firms are acting on the February 2024 classification the paper references. On the technology side, look for call analytics providers to publish model update notes that acknowledge threshold instability and to add fingerprinting of repeated assets to their marketing copy. Finally, track whether brands begin to report the rate of disclosed automation in their call scripts; even a small shift from 0.44% would indicate that disclosure is moving from theory into practice.

This is single-source, unaudited research: an arXiv preprint with its own disclosed methods and limitations, and no outside parties consulted here. But for operators who buy or oversee lead-gen calls, the direction is clear enough to act on today: the risk sits where the budget does, and the tools you were sold to detect it are not yet a shield you can rely on.

More stories