Author ORCID Identifier
https://orcid.org/0009-0005-2171-5501
Document Type
Conference Paper
Disciplines
Computer Sciences, Women's and gender studies, Social issues
Abstract
Gendered language is the use of words or phrases that indicate an individual's gender. Although useful in some contexts, gendered language can reinforce gender stereotypes and introduce bias, particularly in machine learning models for tasks involving people such as recruitment or occupation classification. When textual content about individuals, such as biographies, contain gender cues, models can learn spurious associations between gender and other characteristics of individuals, such as profession, potentially resulting in unfair outcomes such as reduced hiring opportunities for women.
To address this challenge, we propose GenWriter 2.0, a hybrid approach that integrates Case-Based Reasoning (CBR) with Large Language Models (LLMs) to rewrite textual content about people in a manner that reduces implicit gender cues while preserving semantic meaning. We aim to mitigate bias at its source by transforming the text people write, enabling the generation of content that obscures gender and can also serve as less biased training data for machine learning systems. GenWriter 2.0 retrieves semantically similar content describing people of a different gender from a casebase and uses an LLM for adaptation to generate controlled, context-aware rewrites.
We evaluate GenWriter 2.0 by measuring gender bias in an occupation classification task, before and after rewriting the biographies used for training the occupation classification model. Results show that models trained on GenWriter 2.0-rewritten biographies achieve higher classification accuracy (up to 98.85%) and significantly reduce gender bias by 78% in nurse and 63% in surgeon biographies, while preserving semantic consistency and maintaining balanced lexical variation. In contrast, LLM-only rewriting, despite introducing greater lexical variation, leads to lower accuracy and smaller bias reduction in all cases. Further analysis demonstrates that moderate casebase reduction and increased casebase diversity preserve the performance of GenWriter 2.0, with casebase reduction improving rewriting efficiency and casebase diversity enhancing rewriting coverage. Overall, GenWriter 2.0 provides a robust and effective approach for rewriting text to mitigate gender bias at the source, achieving a strong balance between gender bias reduction, classification performance, semantic preservation and lexical variation.
DOI
https://doi.org/10.1007/978-3-032-33865-5_14
Recommended Citation
Soundararajan, Shweta and Delany, Sarah Jane, "GenWriter 2.0: A Hybrid Case-Based and LLM Rewriting Approach for Mitigating Implicit Gender Cues in Text" (2026). Conference papers. 460.
https://arrow.tudublin.ie/scschcomcon/460
Funder
Technological University Dublin
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Included in
Gender and Sexuality Commons, Gender, Race, Sexuality, and Ethnicity in Communication Commons, Human-Computer Interaction Commons
Publication Details
https://link.springer.com/chapter/10.1007/978-3-032-33865-5_14
doi:10.1007/978-3-032-33865-5_14