RadGuide AI claims auditable LLM workflows will reshape hospital imaging procurement
A v1 medRxiv preprint reports RadGuide AI, a tool-augmented LLM framework for nuclear medicine, integrating 55 specialized tools and reporting 88.
Edward Mullen ·

The prevailing wisdom suggests that artificial intelligence for diagnostic imaging will be adopted based purely on its performance metrics. However, this view overlooks the fundamental role of trust and accountability in healthcare procurement. The true differentiator will not be raw accuracy, but rather the auditable, transparent systems that allow hospitals to understand and verify every AI-generated decision.
What the preprint actually builds and measures
The authors present a modular LLM orchestration that connects the language model to 55 specialist tools intended to make radiopharmaceutical decision workflows "traceable" and auditable, according to the preprint's abstract and technical description. The paper emphasizes structured calls to tooling for tasks such as radiochemistry specification and dosimetry calculations rather than free-form narrative outputs.
It reports a headline performance number of 88.5% for its evaluation, presented as evidence the architecture can standardize complex workflows.
Why the headline metrics need scrutiny before procurement decisions A v1 preprint is an unvalidated claim and the paper's 88.5% figure requires contextual interrogation: the medRxiv document does not furnish a clear baseline comparator in the reporting packet, nor does it disclose whether that metric is measured on simulated workflows, retrospective chart data, or prospective clinical cases. The preprint also does not make public the hardware, dataset splits, or failure modes the number conceals.
For a hospital procurement team, a percentage without clear in-distribution vs out-of-distribution delineation is insufficient for risk acceptance.
Why procurement panels will prefer auditable, tool-oriented stacks Procurement for diagnostic imaging is driven by liability, integration, and explainability as much as raw accuracy. An architecture that records explicit tool calls and produces machine-readable audit trails maps directly onto contract and compliance needs: logs can be included as deliverables in service-level agreements; deterministic tool outputs can be independently re-run; and audit trails can be part of a validation dossier for internal governance.
That alignment means procurement committees will have a concrete checkbox — auditable tool integration — that black-box models cannot satisfy without significant extra engineering.
The skeptical read: what the paper does not resolve A reasonable counter is that hospitals will prioritize immediate clinical uplift and vendor time-to-market over auditability, especially where vendors package validated datasets and indemnities. The preprint does not address total-cost-of-ownership, vendor lock-in risks, or how its logs would integrate with existing electronic health record or PACS audit systems, leaving open whether the proposed transparency is practically consumable by hospital IT.
This is the omission the authors do not answer in the reported packet.
What changes inside hospital org charts and procurement processes If purchasers accept auditability as a procurement requirement, hospital RFPs will add new roles and checkpoints: legal and compliance will demand schema for audit logs; clinical engineering will require deterministic replay tests; procurement will add acceptance criteria for tool-level interfaces and versioned traceability. That will shift decision authority away from pure radiology chiefs toward cross-functional committees that can interpret technical audit artifacts.
Vendors who can present an end-to-end traceable workflow—linked model, tool, and dataset versions—will find it easier to win enterprise contracts.
Signals that will falsify or strengthen this thesis in the next 6–12 months Watch whether public hospital tenders start listing "auditable decision trails" or explicit tool-integration requirements in their AI RFPs; whether pilot procurements demand replayable tool logs as deliverables; and whether peer-reviewed clinical evaluations replicate the preprint's 88.5% result on held-out, real-world cases rather than internal benchmarks. Conversely, if major procurement frameworks continue to accept proprietary black-box outputs without traceability, that will falsify the thesis.
These signals are monitorable inside contract archives, procurement notices, and subsequent peer-reviewed validations—none of which the current medRxiv packet includes.
RadGuide AI's preprint stakes a claim: tool-augmented LLMs can make nuclear medicine decision workflows traceable. Whether that technical architecture becomes a procurement must-have depends less on the single reported accuracy number and more on whether hospitals and regulators treat auditable evidence as a contractual requirement.
The paper provides a prototype argument; it omits the purchasing, integration, and compliance plumbing that will determine whether auditable stacks outcompete black-box models in practice.