LMSYS chat audit finds 0.90% user abuse; apologies raise immediate risk

A new study of 777K chatbot conversations reveals that 0.90% of user turns involve mistreatment. Learn how apologies impact user hostility levels.

Hannah Vogel ·

LMSYS chat audit finds 0.90% user abuse; apologies raise immediate risk

In a v1 arXiv preprint dated September 2026 (arXiv:2609.13579v1), the authors audit 777K English conversations from LMSYS-Chat-1M and report that two independent detectors together mark about 5% of user turns for either toxic content solicitation or assistant-directed harassment, with a precision-adjusted estimate of 0.90% specifically for mistreatment aimed at the assistant. The paper emphasizes these absolute rates describe arena-style evaluation traffic and should not be read as deployment-wide base rates. No one in the reported packet is on the record.

The numbers are small, but the cost line they touch is not

The study uses two signals that capture different, weakly overlapping phenomena: a lexicon-based detector for eight categories of assistant-directed hostility (insults, threats, jailbreak coercion among them), and the dataset’s own moderation flags, which the authors say are dominated by toxic-content solicitation. Together, these mark around 5% of user turns in the sample; when narrowed to assistant-directed harassment and adjusted for measured precision, the estimate is 0.90%. In a consumer arena that sounds minor. In a support queue or lead-qualification funnel running millions of turns a month, 0.90% is a predictable, billable stream of escalations to human agents, extra moderation passes, or automated rate-limits—all of which cost money and degrade measured resolution times. The authors’ caveat about generalizability underscores the budgeting challenge: a procurement team cannot lift 0.90% into a forecast, but it also cannot round it to zero if every escalation triggers a human handoff or adds tokens under a consumption contract.

The study’s framing matters for finance leaders because most enterprise vendors now price chat assistants on consumption—tokens, requests, or session minutes—not seats. That shifts risk from a fixed license to a variable cost line. If abusive or coercive interactions lengthen sessions, force additional refusals, or require extra redaction, those costs accrue immediately. The precise rate in your environment will differ from LMSYS arena traffic, but the direction of travel is clear: hostile turns are not just a moderation problem, they are a utilization problem.

Apologies appear to invite abuse in-session, even as more apologetic models get less hostility overall

A striking within-conversation finding is that assistant apologies are consistently associated with higher odds of next-turn hostility under both detectors. The effect holds even after the authors restrict to non-refused prior turns and exclude jailbreaks, and it is positive in 20 of 23 models tested. At the same time, the cross-model comparison shows that more apologetic models receive less hostility overall. For operating teams, the combination points to a trade-off. In-session, apologizing may nudge some users toward more adversarial behavior, raising the likelihood of a costly escalation. In the aggregate, a more apologetic tone correlates with calmer traffic.

Policy and copy matters here. “We’re sorry for the inconvenience” is a brand-safe reflex; it may also be an operationally expensive trigger in a subset of sessions. Conversely, clear boundary-setting without apology may reduce immediate adversarial probing while risking higher overall hostility if the model’s persona feels curt. The preprint does not prescribe a persona; it flags a measurable interaction effect that support and marketing leaders can test in their own flows. Expect copy A/B tests to move from click-through and CSAT to “apology-to-hostility transition rate” alongside first-contact resolution.

Moderation flags aren’t measuring what you think they are

The paper’s two-detector approach is a practical warning for buyers and sellers of conversational systems. The dataset’s moderation signal is “dominated by toxic-content solicitation rather than hostility at the model,” while the lexicon catches insults, threats and coercive jailbreak attempts directed at the assistant. Those are related but distinct risks. If your dashboard only tracks generic toxicity, you may be carrying a gap in brand safety and agent welfare—as well as a blind spot in forecasting marginal cost under a consumption plan—because assistant-directed harassment is not the same signal as content toxicity.

For procurement, that implies new questions for RFPs and SLAs: ask vendors to report assistant-directed mistreatment per 1,000 turns, not just toxicity rates; require separate telemetry for jailbreak coercion attempts; and specify response policies that cap token spend or trigger human handoff when adversarial pressure persists. For vendors, it implies a sales motion shift: you will be asked to disaggregate moderation metrics and prove that your system’s persona and refusal style manage both brand risk and cost of service under real traffic.

Who you attract may matter more than how you behave

The authors find that user hostility varies 13-fold across models and argue that differences are “driven largely by who each model attracts rather than by model behaviour.” They note that first-turn hostility spreads far wider than post-response hostility, and that more than fifteenfold separates the extremes even after deduplicating opening prompts. If that holds in production channels, it reframes model selection and distribution as go-to-market choices. A bot embedded in a gaming forum or a developer sandbox will attract different opening salvos than one in a loyalty app. Marketing placement and audience gating, not just safety tuning, will drive your abuse profile—and your moderation and compute bills.

For platform owners, it also suggests a channel strategy: front-load defenses and ID gating where opening prompts are adversarial, and spend more on persona tuning where affective hostility accumulates over a session. For vendors, it offers a caution about comparative claims. “Our model elicits less abuse” is probably a statement about your user acquisition funnel as much as your refusal policy. The study’s finding is a reminder to separate model traits from channel mix in both demos and pilots.

The limits of the evidence, and the objection you will hear

This is a single-source research preprint auditing a public, arena-style dataset. The authors themselves caution that these absolute rates “should not be read as deployment-wide base rates.” Real deployments feature identity, entitlements, rate limits, purpose-built prompts, and channel norms that an open arena lacks. The detectors themselves pick up different phenomena, and the lexicon-based approach will miss euphemism and sarcasm while the moderation flags will overcount toxic requests that are not assistant-directed hostility. A skeptic will reasonably argue that production hostility rates are either much lower (because of gating) or much higher (because of complaint-heavy support contexts) than the arena measure, and that apology effects will be confounded by workflow design.

That objection is not a reason to dismiss the findings; it is a reason to localize them. The paper surfaces measurable transitions—apology-to-hostility, first-turn adversarial openings—that your team can instrument. Even if the base rates diverge, the direction of the operational trade-offs will appear in your telemetry within a week of adding the counters.

What changes for buying, selling and operating over the next budget cycle

For buyers—support leaders, CMOs running chat in acquisition journeys, and procurement—the near-term change is contractual. Consumption pricing is now standard for LLMs and many CCaaS stacks. If hostility lengthens sessions or increases refusal-and-retry loops, you pay more. RFPs should now ask for: (i) reporting of assistant-directed mistreatment per 1,000 turns and jailbreak coercion attempts; (ii) configurable apology and refusal policies with roll-back and guardrails; (iii) hard caps or rate-limit controls tied to adversarial detections; and (iv) human handoff triggers that balance cost with brand risk. Without those, you are accepting an unpriced variable exposure.

For vendors, the sales narrative will shift toward the composition of traffic as much as model traits. The study’s cross-model variance being driven by “who each model attracts” puts the onus on channel design and persona alignment. Expect savvy buyers to ask for in-tenant A/Bs—your model versus a baseline—with their audience and intent mix, not generic arena benchmarks. The apology finding will also pull product teams to expose finer-grained controls: tune apology frequency by intent class, separate refusal templates for safety versus capability limits, and provide a turnkey way to measure the next-turn effect buyers now know to watch.

For operators, the immediate workflow change is escalation hygiene. If apologies can increase next-turn hostility in-session, the copy that follows a refusal should blend boundary-setting with resolution options instead of open-ended contrition. Because the paper finds affective hostility accumulates over a session while coercive openings front-load first turns, guard your first impression: constrain opening prompts in channels with adversarial audiences, and monitor longer sessions for sentiment drift.

What to watch in the next two quarters

Watch product changelogs and release notes for apology and refusal controls surfaced to admins; if vendors believe the within-session finding is real, they will make this tunable and measurable. Expect to see at least some RFP templates add an assistant-directed hostility KPI and a separate “jailbreak coercion” counter; if sellers start responding with demo dashboards, you will know this has become a sales conversation, not just a safety blog post. Finally, observe where chat is embedded: platforms with adversarial opening prompts—developer arenas, gaming communities—will trial stronger gating or identity checks at session start, while loyalty and commerce apps may bias toward persona tuning to manage cumulative affective hostility. Any of these would validate that operators are acting on precisely the interaction patterns the preprint highlights.

More stories

Latest news