Text-to-image buyers face skewed gender outputs across Stable Diffusion generations, arXiv study says

An arXiv preprint examines 8,000 images from four Stable Diffusion generations and reports that 76.4% depict men, with no model reaching gender parity.

Hannah Vogel ·

Text-to-image buyers face skewed gender outputs across Stable Diffusion generations, arXiv study says

In an arXiv preprint (version v1) examining 8,000 images across four Stable Diffusion generations, the authors report that 76.4% of outputs depict men, with bias worsening from SD 1.5 to SDXL before partially recovering in SD 3 Medium. This is single-source — an arXiv preprint only, not independently verified, and unaudited — but it quantifies a risk creative and marketing buyers have felt anecdotally: default text-to-image outputs skew male, even for historically female-coded occupations, and no model tested achieved gender parity. That finding, if it holds, does not live in ethics decks; it lands in procurement checklists, brand guidelines and the warranty page of vendor order forms.

The paper's numbers matter because they mirror everyday creative prompts, not niche edge cases

Across 20 occupations, 5 prompt templates and 4 Stable Diffusion model generations (SD 1.5, SD 2.1, SDXL, SD 3 Medium), the paper says it generated 8,000 images — n = 100 per occupation-model cell (5 prompts x 20 images) — and classified them using DeepFace. The authors report that 76.4% of the images show male subjects (95% CI [75.1%, 78.7%], p < 2.2 x 10^-16, Benjamini–Hochberg adjusted). For historically female-coded occupations, 57.6% of outputs still depict men (raw p = 3.43 x 10^-22, BH-adjusted p = 1.71 x 10^-21). When benchmarked to U.S. Bureau of Labor Statistics workforce data, models underrepresent women by 20–46 percentage points on average, with large deviations for near gender-balanced roles: scientist (48% female in BLS, 82–99% male in outputs) and cleaner (46% female in BLS, 80–92% male in outputs), the preprint reports. The authors add that bias does not improve linearly with model updates: it worsens from SD 1.5 to SDXL before a partial recovery in SD 3 Medium. No model achieves gender parity. All nine significant tests in the paper survive BH correction across 10 tests. A preliminary comparison with GPT-image-1 on five occupations suggests lower bias than the open-source models, though the practical effect is small (Cramer's V = 0.080) and the comparison is explicitly exploratory. [S1]

Why this shifts from ethics to procurement: unscripted prompts become compliance and brand risks at scale

Marketing and creative teams increasingly use text-to-image tools as first-draft engines for ads, landing pages, and recruiting visuals. The paper’s design — default occupational prompts across multiple templates — approximates how non-specialists actually work. If the default output for “a scientist” is 82–99% male across model generations while the U.S. workforce baseline is 48% female, then unattended use can systemically misrepresent a brand’s audience or workforce. That creates three immediate operational consequences: added human review steps, new prompt standards to explicitly balance representation, and contract provisions that push fairness and QA obligations onto vendors. Each adds cycle time and cost, and each is measurable in service-level terms buyers can put into a statement of work. [S1]

The denominator problem is the brief, not only the model — but the numbers say defaults still dominate

A common pushback is that careful prompt engineering (“a woman scientist,” “a diverse team of…”), style controls, or negative prompts can correct skew. The preprint’s point is that across five generic prompt templates and four model generations, a majority of outputs default to male representation even where the underlying labor force is close to balanced. In creative operations, most requests will not carry detailed demographic instructions; the brief is “we need a hero image for the career site” rather than “ensure representation reflects BLS ratios by role.” That asymmetry is why defaults matter. If defaults over-index male across roles, buyers either standardize prompts to counterweight the model or add post-generation curation — both of which turn into process and spend lines that procurement now has to model and vendors now have to price. [S1]

What the paper does and doesn’t show — limitations that buyers should translate into contract language

The findings are strongly stated but carry caveats. The preprint classifies gender via DeepFace, an automated classifier that will have error rates and may not capture non-binary or culturally diverse presentations; buyers should assume some misclassification and test on their own assets. The occupation set is 20 roles; risk could differ in categories not covered (e.g., caregivers, technical trades). The benchmarking is to U.S. BLS data; global campaigns will require different baselines. The GPT-image-1 comparison is preliminary and small-sample across five occupations; the reported effect size is small (Cramer’s V = 0.080). Finally, this is a preprint — not peer-reviewed, unaudited — and no one in the reported packet is on the record. In commercial terms, those limitations translate into concrete asks: vendors should disclose their own test design, the occupations covered, the baseline source, and classifier used; they should provide confidence intervals, not just point estimates; and they should agree to a re-test protocol on customer prompts. [S1]

Sales collateral will have to carry parity metrics, not just safety badges

Trust-center pages for generative tools today tend to emphasize content safety, IP indemnity and data handling. If buyers begin to treat representation imbalance as a measurable operational risk analogous to safety (and the preprint’s numbers give them a basis to do so), sales motions will change. Expect RFPs to ask vendors to demonstrate occupational parity testing across a named prompt set; expect warrant language that commits to publishing model-level bias metrics by version; and expect service credits if monthly sample tests exceed a representation deviation threshold against a declared baseline. Vendors that can show configuration controls (for example, a toggle or template library that targets parity by role) will have an easier time converting pilots into enterprise seats because they reduce the QA burden on the customer’s side. [S1]

The open-source versus closed question: small measured differences won’t spare anyone from measurement duty

The preprint reports that a preliminary comparison with GPT-image-1 suggests lower bias than the open-source models tested, but the authors also stress that the effect is small and the comparison exploratory. That nuance matters for buying decisions. A small measured gap will not excuse any vendor, open or closed, from providing parity evidence at the version level. If anything, the paper’s generation-over-generation result — bias worsened before recovering — argues against assuming that newer equals fairer. Buyers should treat each model version as a fresh object that requires re-testing and a re-baselining of prompts and QA. Vendors who publish per-version representation metrics alongside safety notes will convert more skeptical enterprise procurement teams than those who argue that “the latest model is better” without numbers. [S1]

The operations consequence: more steps, more time, and a budget line for fairness testing

This is less an ethics debate than a workflow change. Creative operations teams will need to add three practical steps if they adopt text-to-image broadly: a template library that encodes balanced prompts by role, an internal sampling plan (say, 100 images per high-traffic role each month) classified by a documented method, and a human review gate before publication for recruiting and brand-forward assets. Procurement teams will ask vendors to share their own occupational parity sample results and to align on baselines (U.S. BLS for U.S.-only campaigns, or regional datasets otherwise). That increases variable cost on both sides. It will also affect renewals: tools that can show lower QA overhead because their defaults are closer to parity will keep seats; those that can’t will be demoted to pilot-only or single-seat use where manual curation is cheap. [S1]

The skeptical read and what would disprove this procurement shift

Skeptics will argue that DeepFace-based classification and a 20-occupation sample cannot justify contract-level changes, and that teams can fix output with better prompts. They will also argue that creative style and audience fit matter more than workforce parity baselines for many campaigns. The preprint offers one rebuttal: even near gender-balanced roles skewed heavily male in defaults, across multiple prompt templates, and no model hit parity. What would prove the skeptics right in practice is straightforward: vendors begin publishing versioned parity metrics that show near-benchmark representation on default prompts, and enterprise buyers stop asking for fairness annexes because their own month-on-month samples match the vendor’s. Until then, the paper gives procurement enough numbers to formalize what many teams have already started doing informally: measure, document, and assign responsibility. [S1]

More stories

Latest news