Atlas360
Markets

KampoBench tests LLMs on 74 Japanese medicine cases

The physician-reviewed framework found top models still trailed reference answers on patient-specific Kampo consultations.

Jurgen Goldmeier
KampoBench tests LLMs on 74 Japanese medicine cases

Hideaki Takata and five co-authors introduced KampoBench, a 74-scenario physician-reviewed benchmark for large language models, text-generating AI systems, in Japanese Kampo medicine, in a medRxiv preprint posted Oct. 11.

The Oct. 11 preprint matters for clinical AI evaluation because the top model scored 74.8% to 86.3% across graders, below 87.9% to 92.4% for physician-revised reference answers.

How it works

KampoBench tests whether a model can handle consultation-style judgment, not just answer exam questions. The authors said earlier evaluations focused on examination accuracy, while no benchmark had measured whether models could apply Kampo knowledge to patient-specific consultations.

The benchmark contains 310 criteria linked to individual conversations, according to the preprint. Scenarios, scoring criteria and reference answers were first drafted with an LLM, then revised by board-certified Kampo physicians.

Takata and co-authors described the intended standard this way: "Evaluation in traditional medicine should measure the judgments a consultation requires, not recall alone."

The authors evaluated 10 current models from three vendors. Because LLM graders may favor systems from their own model family, each response was scored by one grading model from each vendor, according to the preprint.

The benchmark follows the design of HealthBench, the authors said. The source material does not name the 10 models or the three vendors in the supplied text.

Performance gaps

The main result was a gap between knowing Kampo concepts and using them in a case. The authors wrote that the shortfall was in applying knowledge to the consultation, rather than in possessing the underlying information.

Models performed best on criteria involving advice to seek care, where reliability reached 83.6%, according to the preprint. They performed worst when answers had to connect a known crude-drug risk to the specific patient case, with reliability of 17.5%.

That weakness was not simply a failure to mention the relevant ingredient. The authors said 79.2% of responses in that category named the drug, suggesting the harder task was linking a known risk to the particular facts of the consultation.

Rankings were consistent across grading models. The preprint reported Spearman’s rho, a statistic that measures rank correlation, of 0.952 to 0.976 across graders.

The authors also measured whether a grading model favored systems from its own family. They reported a family-favoring deviation of 0.4 to 0.6 percentage point, which they said was below the bootstrap standard deviation of one measurement, a measure of sampling variability.

Grader leniency still varied by 14.3 points, according to the preprint. That finding led the authors to say a KampoBench score is usable only when reported together with the grader that produced it.

Limits and access

The study is a medRxiv preprint and has not been certified by peer review. medRxiv says preprints report medical research that has not yet been evaluated and should not be used to guide clinical practice.

The authors disclosed competing interests. Tatsuya Nogami received lecture fees and collaborative research funding from Tsumura & Co., and lecture fees from Kracie Pharmaceuticals, within the past 36 months.

Hiroki Inoue and Kazushi Uneda received lecture fees from Tsumura & Co. in the same 36-month period, according to the preprint. Tetsuhiro Yoshino is employed at Keio University for collaborative research with Tsumura & Co.; Hideaki Takata and Yoshinao Harada declared no competing interests.

The dataset was deposited at Zenodo under a CC BY 4.0 license, but the 74 scenarios and 310 criteria are restricted until the article is published in a journal. The authors said the same record contains per-criterion grading outcomes for all 10 models and reference answers under the three graders.

The authors said model responses and grader explanations were not deposited, to reduce the chance the benchmark becomes part of future training data. They said those materials are available from the corresponding author on request, while evaluation code, measurement conditions and dataset provenance are available on GitHub.

Source: academic preprint, medRxiv, Oct. 11, 2026

More stories

Latest news