Introduction Gender biases in Generative Artificial Intelligence (GenAI) and Large Lan-guage Model (LLM) systems have been observed when the training datasets do not reflect the heterogeneity of social groups [1]. Texts and images outputs may perpetuate stereotyped gender roles, skills and characteristics. The human inter-action with biased models may reinforce them and enhance existing social ine-qualities, thereby reducing the identification and mitigation of such models [1-3]. This article presents the preliminary results of a systematic review with the aim of mapping and synthesising evidence related to gender bias in texts and images generated by AI tools. Research design and methods Following PRISMA criteria, a systematic research was conducted in Scopus, Web of Science and PubMed databases using the following query: ("gender" AND bias*) AND (text* OR image* OR languag*) AND (stereotyp* OR prejudic* OR discriminat* OR "sexism") AND ("machine learning" OR "artificial intelligence" OR generative* OR "deep learning" OR "neural networks" OR "large language model" OR "natural language processing" OR chatbot*). The inclusion criteria were limited to English-language articles explicitly ref-erencing LLM, Natural Language Processing (NLP) or GenAI, or which included references to GenAI tools. Studies exclusively focusing on dataset design for bias detection or reduction, as well as reviews, meta-analyses, and conference reports, were excluded. The methodological quality was assessed using the New-castle-Ottawa Scale (NOS), which indicated an overall moderate risk of bias across the selected literature. Results A total of 926 records were identified through databases researching. During screening 312 duplicates were removed, 378 records were conference papers or other types of publication format, 118 records were not pertinent, and three rec-ords were excluded for other reasons. A final list of 115 records was obtained. After full-texts assessment, 81 articles were included. Preliminary findings reveal the prevalence of studies from European (31) geo-graphical area and a notable research interest in this subject area worldwide. OpenAI and Google tools were the most used systems. Articles were organized into a framework of five thematic clusters. Profiling Texts (PTs) and Profiling Images (PIs) involved studies with GenAI outputs by prompts aimed to design a specific outcome (e.g. “Create a doctor”). Gender Texts (GTs) and Gender Images (GIs) included studies that created a generic person (e.g. “Create a person”). Machine Translation Texts (MTTs) involved studies designed to analyse how automated systems select a specific gender fol-lowing a translation from a gender-neutral language to a binary language. The-matic analysis revealed that men are overrepresented in leadership and high-status professional roles within the PI cluster. Furthermore, GenAI tends to asso-ciate domestic domains with women within GTs, while MTTs studies have shown how automated systems often default to male pronouns when translating from gender-neutral languages. Conclusions The utilisation of AI systems is becoming prevalent in a wide range of applica-tions. It is critical to recognise the implications of GenAI utilisation on the rein-forcement of conventional gender stereotypes. Their perpetuation may influence access to essential services [3, 4] or worsen interpersonal relationships, thereby contributing to phenomena including gender-based violence and homophobic bullying [5, 6]. Although our findings may facilitate debiasing strategies to pro-mote equity principles and to consider variability of social groups in the GenAI products, the literature selected were limited to English-language. Future re-search should adopt multilingual search strategies to provide a more inclusive understanding of the cultural biases identified.
Gender Biases in texts and images produced by Artificial Intelligence: preliminary results of a systematic review
Sabino Mutino
;Valeria Gabellone;Fabiana Nuccetelli
2026-01-01
Abstract
Introduction Gender biases in Generative Artificial Intelligence (GenAI) and Large Lan-guage Model (LLM) systems have been observed when the training datasets do not reflect the heterogeneity of social groups [1]. Texts and images outputs may perpetuate stereotyped gender roles, skills and characteristics. The human inter-action with biased models may reinforce them and enhance existing social ine-qualities, thereby reducing the identification and mitigation of such models [1-3]. This article presents the preliminary results of a systematic review with the aim of mapping and synthesising evidence related to gender bias in texts and images generated by AI tools. Research design and methods Following PRISMA criteria, a systematic research was conducted in Scopus, Web of Science and PubMed databases using the following query: ("gender" AND bias*) AND (text* OR image* OR languag*) AND (stereotyp* OR prejudic* OR discriminat* OR "sexism") AND ("machine learning" OR "artificial intelligence" OR generative* OR "deep learning" OR "neural networks" OR "large language model" OR "natural language processing" OR chatbot*). The inclusion criteria were limited to English-language articles explicitly ref-erencing LLM, Natural Language Processing (NLP) or GenAI, or which included references to GenAI tools. Studies exclusively focusing on dataset design for bias detection or reduction, as well as reviews, meta-analyses, and conference reports, were excluded. The methodological quality was assessed using the New-castle-Ottawa Scale (NOS), which indicated an overall moderate risk of bias across the selected literature. Results A total of 926 records were identified through databases researching. During screening 312 duplicates were removed, 378 records were conference papers or other types of publication format, 118 records were not pertinent, and three rec-ords were excluded for other reasons. A final list of 115 records was obtained. After full-texts assessment, 81 articles were included. Preliminary findings reveal the prevalence of studies from European (31) geo-graphical area and a notable research interest in this subject area worldwide. OpenAI and Google tools were the most used systems. Articles were organized into a framework of five thematic clusters. Profiling Texts (PTs) and Profiling Images (PIs) involved studies with GenAI outputs by prompts aimed to design a specific outcome (e.g. “Create a doctor”). Gender Texts (GTs) and Gender Images (GIs) included studies that created a generic person (e.g. “Create a person”). Machine Translation Texts (MTTs) involved studies designed to analyse how automated systems select a specific gender fol-lowing a translation from a gender-neutral language to a binary language. The-matic analysis revealed that men are overrepresented in leadership and high-status professional roles within the PI cluster. Furthermore, GenAI tends to asso-ciate domestic domains with women within GTs, while MTTs studies have shown how automated systems often default to male pronouns when translating from gender-neutral languages. Conclusions The utilisation of AI systems is becoming prevalent in a wide range of applica-tions. It is critical to recognise the implications of GenAI utilisation on the rein-forcement of conventional gender stereotypes. Their perpetuation may influence access to essential services [3, 4] or worsen interpersonal relationships, thereby contributing to phenomena including gender-based violence and homophobic bullying [5, 6]. Although our findings may facilitate debiasing strategies to pro-mote equity principles and to consider variability of social groups in the GenAI products, the literature selected were limited to English-language. Future re-search should adopt multilingual search strategies to provide a more inclusive understanding of the cultural biases identified.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


