SPEAK bioRxiv preprint claims labs will spend more on expert spatial data
A new bioRxiv preprint introduces SPEAK, an LLM-based approach that relies on 'expert aligned knowledge' to identify tissue domains in spatial transcriptomics.
Edward Mullen ·

The prevailing narrative in bio-AI spotlights ever more complex algorithms as the singular driver of progress. However, a recent preprint detailing SPEAK, an LLM-based method, quietly upends this orthodoxy.
Its reliance on "expert aligned knowledge" for identifying tissue domains in spatial transcriptomics data shifts the value proposition. The new battleground for R&D dollars and competitive advantage will be in securing and meticulously annotating proprietary spatial datasets, not solely in architectural brilliance.
If SPEAK
A preprint that roots model performance in human expertise
The dominant read misses the margin shift
What the paper reports, and what it doesn’t
Why this becomes a procurement problem within 12 months
The skeptic’s case: open datasets and model cleverness could suffice
The unpriced variable: rights, access, and curation labor
Who benefits, who is exposed, and the under-noticed middle
How this could be wrong — and what to watch next Three observable signals would falsify this margin-shift thesis. First, if top labs run head-to-head studies in the next year showing minimal loss on open, un-annotated SRT corpora versus expert-aligned setups, then expert data scarcity is not the bottleneck. Second, if leading biotechs publicly report falling data-acquisition costs for spatial omics thanks to abundant, non-proprietary high-quality datasets, then the rights-and-annotation premium is overstated. Third, if buyers and investors price higher for companies shipping novel spatial model architectures than for those selling curated datasets and annotation services, then algorithms — not data — remain the locus of value. Until one of those shows up, procurement teams should assume that expert-aligned data, not raw model ingenuity, is the gating resource for deploying SPEAK-class methods into real pipelines.
The preprint frames SPEAK as “a new large language model (LLM)-based method for identifying spatial domains in spatially resolved transcriptomic (SRT) data,” integrating LLM prompting with human expert input through two stages. In plain terms: the core technical idea is not just pattern finding in gene-expression maps, but steering that pattern finding with domain knowledge encoded up front, so the model’s segmentation of tissue regions lines up with what experts consider biologically meaningful.
That makes the data substrate — expert-labeled regions, controlled vocabularies, and curated references — a load-bearing component rather than an optional enhancement.
The knee-jerk reading is that this is another notch in the belt for LLMs conquering bio-data tasks, with differentiation flowing from ever more intricate architectures. SPEAK’s own setup points the other way: if performance depends on “expert aligned knowledge,” then the scarce input is not accelerator time or clever decoders — it’s the availability and rights to use expertly annotated spatial datasets across tissues, disease states, and platforms.
For operators, that moves margin to whoever controls annotation pipelines and dataset licensing, not whoever publishes the next preprint model variant.
From the public summary, SPEAK combines a two-stage prompting procedure with expert knowledge to identify spatial domains. As a preprint, it should be treated as preliminary: the summary does not enumerate baseline comparisons, hardware, or cross-cohort generalization tests an enterprise buyer would need to evaluate reproducibility outside the authors’ environment.
Without those specifics, executives should assume sensitivity to prompt wording, training corpus provenance, and domain shift (new tissues, slide prep differences) until shown otherwise. This is not a knock on the approach; it is a procurement reality when the performance driver is the quality and fit of human-aligned inputs.
If expert alignment is the lever, budgets move toward acquiring three things: first, high-quality SRT data with appropriate rights attached; second, structured annotations produced or validated by domain specialists; third, the ontologies and controlled vocabularies that make those annotations interoperable. That spending profile advantages hospitals and specialist labs that can bundle data custody with expert review, as well as CROs that can contract for annotation at scale.
It puts pressure on algorithm-only vendors without proprietary data access, who will face compressed pricing against buyers skeptical of model claims that lack clear, transferable performance tied to specific, licensed datasets.
A fair counterargument is that sufficiently capable models trained on broad, open-weights corpora may extract useful spatial domains from public, minimally annotated SRT datasets — making heavy expert alignment optional. If major groups soon post comparable results on open, un-annotated corpora, the budget shift predicted here would stall.
That is the test to watch: whether SPEAK-class methods deliver clear, repeatable deltas that only appear when expert-curated prompts and labels are present, or whether general-purpose LLMs plus light curation close most of the gap.
Even if the algorithm is free to run, the inputs are not. Spatial transcriptomics involves human samples, institutional review, and lab-specific protocols that complicate data rights and re-use.
The preprint does not discuss the downstream economics of acquiring, normalizing, and annotating SRT at the scale implied by multi-tissue, multi-disease applications. If SPEAK-like systems are brittle without carefully aligned prompts and references, then the hidden cost line is the curation labor — pathologists, molecular biologists, and informatics teams — whose time must be compensated and whose institutions may assert control over derivative dataset use.
Institutions that own both tissue data and access to expert annotators are well positioned to package data-plus-alignment as a premium product, capturing a larger share of bio-AI value than pure-play model vendors. Startups whose pitch is primarily architectural novelty in spatial modeling, without privileged data access or annotation services, will feel pricing pressure as buyers ask, explicitly, what curated datasets come with the software.
The least obvious winner is the ontology and tooling layer: whoever standardizes the expert-aligned vocabularies SPEAK-type systems depend on will become a default dependency embedded in many labs’ workflows.