LLM content policies that swap pronouns for descriptions still encode gender

An arXiv preprint introduces a dataset showing that 'neutral' physical descriptors carry structured gender associations for human readers and that large…

Hannah Vogel ·

LLM content policies that swap pronouns for descriptions still encode gender

In a preprint on arXiv, researchers introduce GAPA, a dataset of 316 physical attributes paired with 14,706 gender-association ratings from 304 US-based annotators, and test 16 large language models against those human judgements. The authors report that seemingly neutral descriptors like “short hair” or “a defined jawline” carry structured, graded gender associations for human readers, and that models partially reproduce these associations while showing systematic misalignment: compressed rating distributions, weaker alignment for associations with men, and an abstention pattern that disproportionately affects the non-binary category. They also release a proxy model trained to predict humans’ judgements and apply it to character descriptions in a literary dataset. This is a preprint, not peer-reviewed, but it names a practical problem customers and vendors have treated as solved: swapping pronouns for “objective” traits does not yield gender-neutral communication. [S1]

Neutral descriptions leak gender signals for readers and models, which weakens a popular mitigation

The preprint’s core contribution is empirical. The team compiled GAPA—316 common physical attributes aggregated from diverse sources—and collected 14,706 ratings from 304 US-based annotators on how those attributes associate with women, men or non-binary identities. In aggregate, the ratings show that people read gender into physical descriptors, with more consistent and distinctive associations for women and men than for non-binary identities. The authors then evaluate 16 models spanning different families, sizes and post-training variants against those human ratings. The models “partially recover human associations” but deviate in consistent ways, including “compressed rating distributions,” “weaker alignment for associations with men,” and “asymmetric abstention that disproportionately targets the non-binary category.” The study’s scope is limited—US-based raters and a curated attribute list—but the finding punctures a straightforward mitigation many enterprises use when prompting models for HR copy, support transcripts, creative briefs or summarisation: remove “she/he/they” and describe. The descriptors themselves carry signal. [S1]

For sales, marketing and HR operations, policy that bans pronouns is not enough and may backfire

Many content governance playbooks for LLM deployments have settled on a simple rule: strip identity labels, forbid gendered pronouns, and rely on descriptive language. This paper shows that such guidance can still encode gender in outputs as read by audiences, which matters commercially. A job ad generator that leans on “strong,” “assertive,” or “soft-spoken” can systematically tilt perception. A support summary that replaces “she” with “short-haired customer” does not neutralise reader inference. The authors also report model abstention asymmetries that “disproportionately target the non-binary category.” In practice, safety-tuned models may refuse to assign or discuss non-binary associations while still producing gender cues via physical descriptors, producing an odd mix of blocked explicit mention and permitted implicit signalling. For brand leaders and HR teams whose compliance hinges on inclusive, non-discriminatory communication, that combination can expose the organisation to complaints even if explicit labels are banned. The gap between policy and perception becomes a governance problem rather than a prompting trick. [S1]

Procurement should shift from blanket attestations to attribute-level evaluation across categories

Enterprise buyers have leaned on vendor attestations—“our model avoids sensitive attributes”—and on generic fairness language in master services agreements. The GAPA preprint argues, implicitly, that category-specific evaluation is needed. Attribute-level benchmarking can test whether a model’s outputs track human-associated signals and whether misalignment is consistent across women, men and non-binary categories. The paper highlights two concrete failure modes to test for: compressed rating distributions, which flatten differences and can mask skew, and asymmetric abstention that over-suppresses certain categories. Buyers should expect vendors to show, at minimum, that their content filters do not disproportionately abstain on non-binary identity-related prompts while still letting physical descriptors leak signal, and that alignment for associations with men is not systematically weaker than for women. Because the dataset’s annotators are US-based and the attribute list curated, procurement should push for locale- and domain-specific validation rather than treating GAPA as universal. But the mechanism—measuring how descriptors function as gender proxies—generalises to many enterprise contexts. [S1]

Safety tuning that over-indexes on abstention can degrade inclusivity in non-obvious ways

The reported “asymmetric abstention” is a practical warning about how safety layers are perceived in the wild. If models more readily refuse to engage with non-binary category associations but still produce descriptors that audiences map to male or female, the result is a surface-level compliance posture that marginalises one group’s visibility. Customer support agents guided by such models may find non-binary mentions struck from summaries while other gendered cues remain. Marketing teams may see creative tools decline brief content about non-binary personas while still outputting trait lists that audiences read as male or female. That double standard is not only an inclusivity failure; it creates operational friction as teams workaround abstentions, wastes cycles on prompt engineering, and complicates downstream audit trails. The study’s claim that models show “weaker alignment for associations with men” adds another operational wrinkle: if tuning or datasets skew the mapping for male-associated descriptors, campaigns or job copy targeting men can drift in tone without teams realising why. The misalignment is a measurable, testable property; treating abstention counts as a safety KPI is not enough. [S1]

Vendors will be pushed to evidence category-aware alignment, not just redact pronouns

The researchers release a “best-performing proxy model” trained to predict human gender associations of descriptive language and demonstrate its use on character descriptions in LitBank. For vendors, that points to a productisation path: use proxy models to pre-screen outputs for implicit gender signalling and to quantify alignment gaps by category. But the preprint also underscores the cost of relying on a single cultural context (US-based annotators) and a fixed attribute list. Any proxy will encode those choices; portability claims should be treated as unverified unless vendors provide locale-specific evaluations. For customers, the demand should shift from high-level bias claims to documented tests: show the distribution of ratings by category, show abstention rates by identity, and show how safety policies handle descriptive proxies, not merely explicit pronouns. Absent that evidence, “gender-neutral” policies are marketing claims, not operational safeguards. [S1]

Expect RFPs to add descriptor-level tests and for model release notes to acknowledge proxy risks

If this preprint influences practice, the next six months should bring procurement questionnaires that ask vendors to report descriptor-level evaluation results and abstention by category, and internal policy updates that replace “no pronouns” with tests for descriptor proxies in HR, support and marketing templates. Model release notes and trust-and-safety pages may begin to acknowledge this specific risk: that physical descriptors function as gender proxies and that safety tuning can create asymmetric abstention harms. The counter-argument is obvious and should be weighed: the dataset reflects US annotator judgements and the tested models may have already evolved; the proxy model is a research artefact, not an industry standard; and any single benchmark can be gamed. But the operational point stands. If customers’ audiences read gender into physical description, vendors who claim neutrality through pronoun swaps—without demonstrating descriptor-level alignment—are selling comfort, not control. [S1]

More stories

Latest news