EU AI Act buyers get executable tests, says arXiv paper on governance-as-code
An arXiv preprint proposes turning EU AI Act obligations for generative systems into machine-checkable tests run in CI/CD.
Hannah Vogel ·

In an arXiv preprint posted in September 2026, the authors propose “Governance-as-Code” for the EU AI Act: 43 machine-checkable acceptance criteria across six modules that run in a CI/CD pipeline and emit Article-indexed audit evidence for generative AI. The paper is single-source — arXiv only, not peer-reviewed — and claims the framework matched a manual expert audit on two enterprise deployments (a high-risk advisory chatbot and a limited-risk content generator) and cut audit labor by roughly 75%. No one in the reported packet is on the record. This is research, not a regulatory finding. [S1]
The claim: turn open-textured obligations into numbers and policy code
The authors argue that Articles 8–15 of Regulation 2024/1689 were drafted for predictive systems and leave seven gaps when applied to generative AI: non-deterministic data governance, training-data provenance, continuous conformity, human oversight, open-ended robustness, emergent risk, and generative fairness. Their Governance-as-Code (GaC) approach encodes acceptance criteria in policy (they show actual Rego code, not just descriptions) and asserts that open-textured standards such as “appropriate levels” or “possible biases” can be operationalized as declared, auditable numbers. In their framing, robustness thresholds are derived from a provider’s documented baseline and a state-of-the-art floor, while “framing bias” is collapsed into eight measurable proxies tested via counterfactual demographic probing. The modules run in the release pipeline and produce evidence keyed to specific Articles, designed to be inspection-ready for auditors. [S1]
Why this is a procurement problem, not just an audit one
The paper reassigns who owes what. Under Article 25 and Chapter V, a downstream deployer relies on an upstream provider’s Article 53 training-data summary and documents only the layers it controls; GaC thus verifies that summary rather than demanding per-sample documentation the deployer never had. If this becomes the norm, procurement stops accepting static “responsibility matrices” and starts asking vendors for machine-verifiable artifacts at the contract boundary: executable tests, versioned Article-indexed logs, and attested training-data summaries. RFPs and SOWs would migrate from binder-ready controls to a suite of acceptance tests that must pass for a model, fine-tune, or retrieval layer to ship. That moves budget to compliance engineering inside product and platform teams, and it drags legal and risk functions directly into the CI/CD gate where “conformity” becomes a release condition, not a post-hoc review. [S1]
The load-bearing caveats the paper itself can’t resolve
The authors validate on two enterprise deployments and say the GaC framework reproduces all manual audit findings, including three penalty-triggering violations, with roughly 75% less labor. But the baseline is a manual expert audit of unspecified design; the paper does not identify the enterprises, the auditors, or whether any regulator accepted machine-generated evidence for conformity. “Penalty-triggering” is their interpretation of the Act’s thresholds, not a decided case. And while the framework claims continuous conformity, non-deterministic outputs raise stability questions: repeated runs can produce different outputs, so pass/fail can fluctuate with seed, sampling parameters, or prompt phrasing. The authors address this by deriving thresholds from a provider baseline and a SOTA floor, but they do not show whether those thresholds survive model updates, distribution shift, or emergent behaviors over time without frequent re-tuning — an ongoing cost that procurement will need to price in. [S1]
Second-order effect: ship cadence becomes a function of compliance gates
Embedding Article-indexed tests in CI/CD turns compliance into a throttle on release velocity. Teams will schedule model updates, retrieval corpus changes, and prompt library revisions around gate stability. Expect “red builds” that represent fairness-proxy regressions or robustness drops against the declared baseline to escalate like Sev1 incidents — with legal in the on-call tree. The framework’s counterfactual demographic probing encourages standardized test suites for fairness, but that also invites teams to optimize to the test, risking gaps where proxies fail to capture real-world harms. Compute costs for running test batteries on large models during release may become material at scale, and buyers will push for vendor-provided, article-mapped test fixtures to reduce duplicate work. The winners will be those who can keep gates tight without starving deployment frequency; the losers will be teams whose CI/CD pipelines can’t absorb the compliance workload. [S1]
Where vendor lock could quietly creep in
Although the paper publishes Rego policies, the real lock-in risk sits one layer up: the choice of acceptance criteria, their thresholds, and the evidence schema keyed to specific Articles. If buyers standardize on a particular policy dialect and evidence format, upstream model providers and tools vendors will be pressured to furnish exactly those artifacts. That can harden into a de facto specification for EU AI Act conformity — authored by whoever ships the first widely adopted test suite. For deployers, swapping vendors later may mean rewriting policies, retraining teams, and remapping logs to Article indices. For providers, failing to supply an Article 53 training-data summary in the expected structure becomes a commercial blocker, not merely a documentation lapse. [S1]
The obvious skeptic’s read — and what would prove this wrong
Skeptics will argue that compliance cannot be fully automated because the Act embeds human-oversight duties and context-sensitive judgments that resist codification. They will question whether national market surveillance authorities will accept machine-emitted evidence for high-risk systems, and whether fairness proxies and robustness baselines can keep pace with emergent risks in generative models. Those objections stand until a regulator signals acceptance of executable evidence, or a major buyer writes Article-indexed tests into standard contracts. Conversely, if large buyers adopt this approach, audit firms will be compelled to read machine logs alongside narratives, and platforms will productize Article-aware gates. Watch whether RFP templates start citing Article 8–15 acceptance tests, whether MLOps vendors ship native Rego-based compliance gates, and whether any early enforcement action references “continuous conformity” failures specific to generative systems. [S1]
What changes in the next renewal cycle if this sticks
If Governance-as-Code migrates from paper to practice, generative AI procurement will resemble security procurement: a pre-negotiated test suite, signed evidence formats, and fail-closed release gates. Buyers will demand upstream training-data summaries and versioned model cards aligned to Article indices; deployers will scope contracts around which layers they can verifiably control; and internal GRC platforms will integrate with policy engines so audit trails fall out of normal build runs. Budget will shift from annual third-party audits toward continuous compliance engineering, with audit firms focusing on sampling, scope assurance, and spot checks of the machine evidence rather than re-performing the tests. Vendors that can’t emit Article-indexed evidence will find themselves excluded before pricing is discussed. [S1]