CDT report: Fine-tuning foundation models can create unpredictable safety drift
A new report from the Center for Democracy & Technology and MIT researchers finds that fine-tuning general-purpose AI models for specialized tasks can prod…
Sophie McAlister ·

The Center for Democracy & Technology (CDT), a Washington, D.C.-based digital-rights nonprofit, published a report co-authored with researchers from MIT’s Algorithmic Alignment Group warning that fine-tuning foundation models for downstream tasks can lead to unpredictable changes in model behavior. The analysis notes that adapting general-purpose models for specialized uses may shift safety properties in unexpected ways.
Authors Emaan Bilal and Dylan Hadfield-Menell, working with CDT, examined how common fine-tuning workflows used to create specialized applications — from medical transcription to legal contract analysis — can produce what the report terms "safety drift." The paper argues that safety controls and behaviors observed in a base model do not always transfer reliably after tuning and that new risks can emerge when models are repurposed.
The Center
The report frames these technical findings as governance challenges, saying that regulators, developers and procurement teams need clearer tools to evaluate how model behavior changes across development cycles. It highlights the difficulty of predicting downstream risks solely from base-model assessments and flags the need for improved documentation, testing and post-deployment monitoring to catch emergent harms.
CDT’s analysis does not prescribe a single regulatory fix but points to operational and policy gaps — including provenance tracking, standardized evaluation metrics for tuned models, and greater transparency from model providers and downstream developers. The authors call for further research and coordinated oversight to ensure safety properties persist as models are adapted for specific, high-stakes tasks.