The linguistic trajectory of terms such as queer, dyke, or bitch exemplifies the sociocultural process of semantic reclamation, whereby historically stigmatizing labels are reappropriated by marginalized groups and resignified as markers of identity, solidarity, or resistance (Brontsema, 2004; Galinsky et al., 2013). From a performative perspective, such processes can be understood as practices through which subjects negotiate recognition and agency within regimes of power that are themselves discursively constituted (Butler, 1997). Pragmatic accounts of slurs additionally emphasize that derogatory force and (re)appropriative potential are context-dependent, shaped by speaker identity/positioning, uptake, and interactional history (Bianchi, 2014; Croom, 2013). Digital environments – and Reddit in particular – provide a significant site for observing these dynamics, as they host both intra-community uses that may signal empowerment and inter-group uses that may reproduce insult, thereby rendering visible the contextual and interactional conditions under which meaning shifts occur in online discourse (Jane, 2014; Massanari, 2015). More broadly, subreddits can be understood as networked publics whose discourse is shaped by platform affordances and imagined audiences, with potential context collapse across audiences that influences how contested language is interpreted and contested (Marwick & Boyd, 2011; Limone & Toto, 2022). Content moderation and other forms of platform governance further co-produce community norms and visibility, affecting the circulation of both harmful and supportive speech (Gillespie, 2018; Roberts, 2019). Previous research in computational linguistics has shown the importance of reclaimed language when training automated hate-speech detection systems to build safer online spaces (Draetta et al., 2026). At the same time, online gaming communities represent a unique environment: they are often ripe with sexist or homophobic discourse (Di Leo et al., 2025), but they can also configure themselves as safe spaces of care and support (Di Leo, 2025). Therefore, analysing the interactions between their users can provide valuable linguistic perspectives into how specific terms are mediated and contested. We created a corpus of Reddit posts and comments taken from major gaming communities. To do so, we manually selected a list of polysemic English lemmas (nouns and adjectives) from Hatebase (2017), a large linguistic resource on hateful speech, specifically targeting reclaimed slurs and terms which very recently underwent semantic extension in online discourse. This data was annotated by adapting the guidelines proposed by Palmer et al. (2020), identifying, for each post and comment, the presence of offensive intent and the meaning of the target lemmas. Candidate meanings were taken from BabelNet (Navigli & Ponzetto, 2010) and Wiktionary (Wikimedia Foundation, 2002). A balanced subset of the data was manually annotated by three researchers independently, reaching a high inter-annotator agreement. This constituted the gold standard for the subsequent automatic annotation of the rest of the corpus, which was executed by an open-source Large Language Model (DeepSeek V4 Pro; DeepSeek-AI, 2026) provided with detailed annotation guidelines and few-shot prompting examples (Brown et al., 2020). Preliminary findings show that different communities employ largely different communication strategies, with some being highly aggressive, and others welcoming and supportive. We also trace which terms are more often used in offensive posts and comments, and which ones in reclamatory contexts. Based on this data, we reflect on the optimal pedagogical strategies to educate students on safe use of contested language (Toto & Limone, 2022). In doing so, we draw on critical language awareness and critical (digital) media literacy to foreground how power, audience, and platform context shape meaning-making and the harms/benefits of using contested terms (Buckingham, 2013; Fairclough), and we situate this reflection within culturally sustaining pedagogy, emphasizing attention to community language practices while engaging questions of safety and justice (Paris & Alim, 2017). The final, annotated dataset will be released openly to enable further investigations into the linguistic properties of the target terms, and, more broadly, into dynamics of reclamation and aggression in online gaming communities. Specifically, we aim to provide a useful resource to train automated hate-speech detection systems.

From “flame” to “frame”. When context flips meaning online

Michele Ciletti
Data Curation
;
Alice Rizzi
Membro del Collaboration Group
;
Nadia Di Leo
Conceptualization
;
Martina Rossi
Supervision
2026-01-01

Abstract

The linguistic trajectory of terms such as queer, dyke, or bitch exemplifies the sociocultural process of semantic reclamation, whereby historically stigmatizing labels are reappropriated by marginalized groups and resignified as markers of identity, solidarity, or resistance (Brontsema, 2004; Galinsky et al., 2013). From a performative perspective, such processes can be understood as practices through which subjects negotiate recognition and agency within regimes of power that are themselves discursively constituted (Butler, 1997). Pragmatic accounts of slurs additionally emphasize that derogatory force and (re)appropriative potential are context-dependent, shaped by speaker identity/positioning, uptake, and interactional history (Bianchi, 2014; Croom, 2013). Digital environments – and Reddit in particular – provide a significant site for observing these dynamics, as they host both intra-community uses that may signal empowerment and inter-group uses that may reproduce insult, thereby rendering visible the contextual and interactional conditions under which meaning shifts occur in online discourse (Jane, 2014; Massanari, 2015). More broadly, subreddits can be understood as networked publics whose discourse is shaped by platform affordances and imagined audiences, with potential context collapse across audiences that influences how contested language is interpreted and contested (Marwick & Boyd, 2011; Limone & Toto, 2022). Content moderation and other forms of platform governance further co-produce community norms and visibility, affecting the circulation of both harmful and supportive speech (Gillespie, 2018; Roberts, 2019). Previous research in computational linguistics has shown the importance of reclaimed language when training automated hate-speech detection systems to build safer online spaces (Draetta et al., 2026). At the same time, online gaming communities represent a unique environment: they are often ripe with sexist or homophobic discourse (Di Leo et al., 2025), but they can also configure themselves as safe spaces of care and support (Di Leo, 2025). Therefore, analysing the interactions between their users can provide valuable linguistic perspectives into how specific terms are mediated and contested. We created a corpus of Reddit posts and comments taken from major gaming communities. To do so, we manually selected a list of polysemic English lemmas (nouns and adjectives) from Hatebase (2017), a large linguistic resource on hateful speech, specifically targeting reclaimed slurs and terms which very recently underwent semantic extension in online discourse. This data was annotated by adapting the guidelines proposed by Palmer et al. (2020), identifying, for each post and comment, the presence of offensive intent and the meaning of the target lemmas. Candidate meanings were taken from BabelNet (Navigli & Ponzetto, 2010) and Wiktionary (Wikimedia Foundation, 2002). A balanced subset of the data was manually annotated by three researchers independently, reaching a high inter-annotator agreement. This constituted the gold standard for the subsequent automatic annotation of the rest of the corpus, which was executed by an open-source Large Language Model (DeepSeek V4 Pro; DeepSeek-AI, 2026) provided with detailed annotation guidelines and few-shot prompting examples (Brown et al., 2020). Preliminary findings show that different communities employ largely different communication strategies, with some being highly aggressive, and others welcoming and supportive. We also trace which terms are more often used in offensive posts and comments, and which ones in reclamatory contexts. Based on this data, we reflect on the optimal pedagogical strategies to educate students on safe use of contested language (Toto & Limone, 2022). In doing so, we draw on critical language awareness and critical (digital) media literacy to foreground how power, audience, and platform context shape meaning-making and the harms/benefits of using contested terms (Buckingham, 2013; Fairclough), and we situate this reflection within culturally sustaining pedagogy, emphasizing attention to community language practices while engaging questions of safety and justice (Paris & Alim, 2017). The final, annotated dataset will be released openly to enable further investigations into the linguistic properties of the target terms, and, more broadly, into dynamics of reclamation and aggression in online gaming communities. Specifically, we aim to provide a useful resource to train automated hate-speech detection systems.
2026
9791282455046
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11369/485732
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact