Supply chain AI buyers can cut carbon without quality loss, arXiv study says

An arXiv preprint proposes a carbon-aware AI procurement playbook for enterprise supply chains and reports that the largest models did not top performance in its six-model, 520-task benchmark. The authors say a routing approach can meet quality targets with far lower estimated emissions per task, su

Hannah Vogel ·

Supply chain AI buyers can cut carbon without quality loss, arXiv study says

In an arXiv preprint (v1) posted in September 2026, the authors argue that enterprises deploying AI into supply chain decision flows are over-paying in environmental terms by defaulting to the largest general-purpose model. Their benchmark of six large language models across 520 supply chain tasks measures both decision quality and estimated generation-related operational carbon, and it undercuts the common procurement heuristic that “bigger is better.” Within their bounded sample, the reported quality span is 0.497–0.723, and models with the largest disclosed parameter totals did not achieve the highest scores. The paper proposes a Carbon-Aware AI Procurement Framework (CAAPF) to operationalize “benchmark first, select green” inside enterprise buying. This is single-source research on arXiv , not peer-reviewed or independently verified.

The research measures quality against estimated grams of CO2 per task, not just leaderboard scores

The authors report a paired evaluation: task-level decision quality and an estimate of generation-related operational carbon intensity per task (gCO2/task). That framing matters for buyers because it puts an environmental unit next to an outcome metric, rather than treating sustainability as a parallel narrative divorced from performance. In their results, a category-by-tier calibrated GreenRoute proof of concept reaches 0.733 mean out-of-sample quality at an estimated 0.402 gCO2/task. A static model labeled Haiku reaches 0.699 at 0.022 gCO2/task, while Sonnet reaches 0.723 at 0.401 gCO2/task. Their stated conclusion is that the “preferred strategy depends on the organization’s quality requirement”: once a quality floor is set, a buyer can pick a lower-carbon model or a router that meets it, instead of reflexively selecting the largest available option. The headline claim is bounded by the authors’ own caution that their design does not isolate size, provider, architecture, or benchmark construction effects.

The default procurement heuristic is the target: stop buying the maximum model to mitigate risk

The practical business target here is not a new model, but a change in buying behavior. The preprint claims that enterprises “commonly default to the largest available language model,” a risk-averse procurement habit that can ignore empirical fit and environmental cost. CAAPF, rooted in the Technology-Organization-Environment framework and described as a Green IS design artifact, is intended to move buyers from brand-size proxies to instrumented selection. In their sample, the authors say the largest disclosed parameter totals did not deliver the highest quality scores, implying that size is an unreliable purchasing proxy for this class of tasks. If accepted, this reframes the vendor conversation from logo and model scale to measured quality-at-carbon for the specific decision categories in scope.

What changes for enterprise buyers if this approach is adopted

If procurement teams take this seriously, two steps follow. First, “benchmark first”: require vendors and internal teams to validate model quality on representative supply chain task categories before committing to a model family, and to report results in the format used in the preprint — a quality metric paired with estimated gCO2/task. Second, “select green”: once a quality threshold is fixed by the business owner (for example, planners setting an acceptable accuracy for lead-time classification), pick the lowest estimated carbon option that clears the bar — whether that is a smaller static model or a routing strategy like the reported GreenRoute. The study’s numbers are illustrative rather than definitive, but they show non-trivial spreads: 0.723 versus 0.699 quality, with orders-of-magnitude differences in estimated carbon per task (0.401 versus 0.022 gCO2/task in the cited cases). That is enough to warrant adding carbon intensity to RFP scoring where AI inference is part of the production path.

The carbon number becomes a governance and assurance knob, not just marketing copy

By providing a per-task carbon estimate alongside decision quality, the paper gives procurement and risk teams a comparable unit to govern. The authors explicitly tie their approach to SDG 12 and SDG 13 and position CAAPF as sustainable digital infrastructure governance. In practice, that means a buyer could set class-by-class acceptance criteria: for routine classification tiers, accept models under a certain gCO2/task threshold; for escalated decisions, allow a higher threshold if the measured quality lift merits it. This shifts environmental responsibility from a generic supplier code to a measurable selection rule inside the AI system design, and it makes the tradeoff reviewable by internal audit. The paper does not price carbon or quantify costs, but the presence of a unit-rate estimate opens the door to internal carbon pricing or disclosure, if a company already uses those tools in broader sustainability governance.

The denominator problem: this is a bounded sample with undisclosed cost, vendor, and architecture effects

The authors are explicit about the limits: their design “does not isolate size, provider, architecture, or benchmark-construction effects.” That caveat is load-bearing. The reported quality range (0.497–0.723) and the specific pairings of quality and estimated gCO2/task are bound to their task set and estimation method. The preprint also focuses on generation-related operational carbon, not embodied emissions from model training, data center construction, or end-to-end supply chain impacts. Buyers should therefore treat the paper as a decision framework and a proof of concept, not a universal table of model merit or a definitive carbon accounting method. The implication for procurement is to demand like-for-like measurements on your workload, not to adopt the authors’ reported numbers as constants.

This is a procurement and governance problem, not an engineering trophy hunt

The paper’s most usable idea for operators is organizational: move the decision right to where purchasing power sits. The authors propose a procurement framework (CAAPF) rather than a purely technical ranking. That implicitly puts the RFP and vendor-selection gate at the center of the change, not a one-off model bake-off in a lab. For supply chain leaders, this suggests drafting category-by-tier model matrices with quality floors, specifying carbon measurement methods vendors must use, and writing contract language that allows the buyer to route tasks dynamically to lower-carbon options that maintain the agreed quality. The “benchmark first, select green” principle is the lever — and it is auditable.

What vendors will have to disclose if buyers move this way

Should CAAPF-like criteria spread, model providers, integrators, and cloud marketplaces will feel pressure to publish standardized per-task or per-token carbon-intensity estimates and to support routing across model tiers. The authors’ GreenRoute result (0.733 quality at 0.402 gCO2/task) is positioned as a proof of concept rather than a product, but the procurement ask is clear: disclose measured quality on the buyer’s tasks and the associated estimated gCO2/task. Vendors who only offer a single “flagship” model may find themselves outmaneuvered by those who package calibrated routers across multiple tiers that can hit the buyer’s quality requirement with lower estimated emissions. Again, the authors caution that their design does not isolate provider effects, so any vendor claims will need to be validated against the buyer’s workload.

The skeptic’s read is straightforward: without cost, latency, and legal terms, this may not move procurement

No one in the reported packet is on the record, and the preprint does not present cost, latency, or contractual constraints — the other three pillars that procurement will weigh alongside quality and environmental metrics. A counter-argument is that absent price-per-task and service-level commitments, a gCO2/task figure is not yet a buy/no-buy criterion. Legal will also want to understand whether carbon estimates are assured and how disputes over measurement are handled in contracts. The paper does not address those questions; its contribution is a decision framework and a paired metric demonstration. For this to affect deals, buyers will need vendors to supply the same paired metrics under test conditions specified in the RFP, and to accept those measurements into service-level governance.

The near-term test is whether RFPs start to require paired quality and carbon metrics

The authors present CAAPF as “sustainable AI governance” aligned with SDG 12 and SDG 13. The test of whether this moves from research to buying behavior is simple and observable: do enterprise RFPs for AI-enabled supply chain systems start to require benchmarked decision quality on the buyer’s tasks and an estimated gCO2/task for the proposed model(s) or routing strategy? If so, procurement teams will gain a new lever to push vendors toward smaller or routed model tiers where they meet quality thresholds. If not, the default to the largest model will persist, and any environmental claims will remain marketing copy rather than contract criteria.

Disclosure: This story relies on a single source — an arXiv preprint . Its claims are unaudited and not independently verified.

More stories