ArXiv preprint finds manager agents add cost and reduce clarity in LLM teams

Hierarchical multi-agent LLM setups consume 51.5% more tokens and reduce clarity. Research suggests supervisory loops only help when verifying facts.

Hannah Vogel ·

ArXiv preprint finds manager agents add cost and reduce clarity in LLM teams

In an arXiv preprint with identifier 2609.14767v1 posted in September 2026, the authors run a paired experiment on a business‑intelligence reporting task and report a counterintuitive result: a hierarchical agent team, in which a Manager can send work back for revision, performs worse on utility and writing clarity than the same team coordinated flat—with the supervisory layer adding 51.5% more token cost for no quality gain. This is a single‑source research preprint, not peer‑reviewed, and no one in the reported packet is on the record. Still, for operators buying or building agentic workflows into marketing, sales enablement and analytics, it challenges the most common production default.

The only link they varied was managerial authority, and it made reports worse

The preprint holds constant five LLM agents, their roles, prompts, tools, models and data, and varies one thing: whether the Manager can reject a worker’s output and oblige a revision. Across 43 paired products and 86 runs, a five‑model judge panel and a deterministic specification check score each report’s Utility and Writing Clarity. The flat organization scores higher on Utility (d = 0.42, p = 0.009) and Writing Clarity (d = 0.34, p = 0.030). The authors say the hierarchical reports hedge 53% more, each revision loop is associated with a 0.14‑point drop in Writing Clarity, and the hierarchical Writer’s first draft is indistinguishable from the flat report—the gap opens during the revision loop. Specification accuracy is at ceiling in both organizations, meaning the task’s factual constraints were satisfied either way, and the managerial layer did not add correctness.

The cost side is equally explicit: the supervisory tier costs 51.5% more tokens for no measurable quality gain in this setup. The authors’ own summation is blunt: “A supervisor pays for itself when it can verify and becomes a liability when it can only opine.” For teams instrumenting agent‑based drafting of briefs, summaries or sales collateral, that is not an abstract claim—it’s a budget line.

Why this runs against the industry’s default hierarchy

Most production agent frameworks ship with hierarchical orchestration templates: a Manager delegates, collects drafts and issues critique, and Workers revise. Classical organizational theory predicts the authority link speeds convergence on decisive output; much vendor marketing adopts that logic for “quality assurance.” The preprint’s paired design isolates the authority link and finds the opposite on an open‑ended task: manager‑driven critique increases hedging and degrades clarity even when facts are correct. That dovetails with documented sycophancy and “degeneration‑of‑thought” dynamics in LLMs: when a model is asked to defend or comply with an authoritative critique without hard ground truth checks, it tends to soften claims and meander, not sharpen them.

For enterprise buyers, the mechanism matters. If the Manager can only opine on style or tone, not verify against a source of truth, each loop appears to nudge the model away from crispness. The preprint also observes that first drafts in both organizations are comparable; the divergence happens inside the managerial loop. That points the finger squarely at the coordination pattern, not the base capability of the models or tools.

The procurement implication: pay for verification, not for opinionated loops

Enterprises pay for agent orchestration in two ways: platform or framework fees, and variable token spend that scales with steps taken and messages exchanged. A hierarchical Manager who critiques and triggers revisions on subjective grounds inflates the latter with no measured lift in utility or clarity in this study. In other words, buyers are paying a premium for a feeling of QA where the only checks are opinionated, not verifiable.

Procurement language can change quickly here. Statements of work and internal runbooks should constrain supervisory loops to verifiable checks—schema conformity, linkable citations, consistent metrics—where the Manager function calls tools that can pass/fail specific criteria. For open‑ended writing and synthesis tasks without ground truth, the preprint suggests running flat, with a single pass through automated specification checks and then human review only at the end. That keeps variable spend down and preserves clarity. Vendors selling “manager agents” as quality engines will be pressed to show the verification harness, not the persona.

What this changes for vendors and teams selling agentic workflows

If your price meter rolls with steps, the default to hierarchy is attractive to your P&L. But this study gives enterprise customers a clean, testable argument to push back: show me the quality delta net of token cost, or default me to flat orchestration with verify‑only gates. Expect savvy customers to ask for template libraries where the Manager role is scoped to deterministic checks and for usage dashboards that break out tokens by role so they can see exactly what they are paying for.

Sales and marketing teams pitching agentic content operations will also need to change proof points. “Manager‑driven critique improves quality” is no longer a safe claim on open‑ended work—at least not without task‑level verification. Demonstrations should shift from cinematic back‑and‑forth critique to before/after comparisons on clarity and utility scores scored by pre‑committed judge panels or human raters, with token‑per‑artifact disclosed.

Inside teams, the org‑chart consequence is straightforward: move the intelligence from managerial persona to spec. Build stronger, machine‑checkable acceptance criteria (links required, numbers reconciled to source tables, reading level targets enforced by tool) and let the writer agent ship once it passes. Where verification is impossible—thought leadership drafts, executive memos—treat the manager loop as a cost center that should be minimized, not celebrated.

The skeptic’s read: one task, one setup, and ceiling effects

There are meaningful limitations. This is a single‑source arXiv preprint; it has not been peer‑reviewed, and no external replication is provided. The task is a business‑intelligence reporting setup; results might not generalize to code generation, safety‑critical workflows, or domains where style conformity is itself a requirement. Specification accuracy being at ceiling means the spec may have been too lenient to expose any factual lift the manager could have produced; on harder, more brittle specs, a supervisory check might matter more. And frameworks differ widely in prompt patterns and tool use. Enterprises should treat the claim as a strong hypothesis to test against their own tasks, not a blanket rule.

What to watch in the next budgeting cycle

If this finding is directionally right, three signals should show up quickly. First, vendors of agent frameworks will add or promote flat‑team templates and “verify‑only manager” modes in their documentation and demos, and de‑emphasize cinematic critique loops for open‑ended writing. Second, enterprise RFPs will begin to require per‑artifact token accounting by role and verifiable QA checkpoints, making it harder to bury revision‑loop costs in aggregate usage. Third, internal AI ops teams will report lower token‑per‑artifact and faster throughput on open‑ended content after removing authority‑driven revision loops, while maintaining or improving clarity scores.

Those are all observable within a quarter. The managerial impulse is hard to unlearn, but in agent land it looks like a paid habit. If the supervisor cannot verify, the preprint’s evidence suggests it should get out of the way.

More stories

Latest news