Beyond Toxicity: Evaluating Language Models for Ableism Detection in Online Social Media

  • Annie Amorim UFF
  • Gabriel Assis UFF
  • Daniele Pimenta UFF
  • Débora Pina UFRJ
  • Aline Paes UFF
  • Daniel de Oliveira UFF

Resumo


Ableism is a form of discrimination directed at people with disabilities that can appear through explicit offenses, but also through subtle, indirect, and context-dependent expressions that are difficult to identify automatically. This challenge remains underexplored in content moderation, particularly in Brazilian Portuguese. This paper investigates how different levels of prompt information affect the performance of Large Language Models (LLMs) in detecting ableist comments on social media. To this end, we build a dataset from YouTube comments related to disability-related content, followed by manual annotation. Four LLMs are evaluated using five progressively detailed prompts, ranging from a simple classification instruction to prompts that include definitions, examples, contextual information, and the video transcript. Results show that more informative prompts tend to improve classification performance, although the effect varies across models. Qualitative inspection further reveals that LLMs often struggle to distinguish ableist discourse from criticism of ableism, as well as to separate ableism from other forms of offensive or discriminatory language. These findings highlight the importance of contextualized prompting for sensitive classification tasks and point to the need for further research on ableism detection in Portuguese.

Palavras-chave: ableism detection, large language models, prompt engineering, social media, text classification

Referências

Assis, G. et al. Explorando técnicas de aprendizado em modelos de linguagem para classificação de discurso de Ódio e ofensivo em português. Linguamática 16 (2): 91–113, Dez., 2024a.

Assis, G. et al. Exploring Portuguese Hate Speech Detection in Low-Resource Settings: Lightly Tuning Encoder Models or In-context Learning of Large Models? In Proc. of 16th PROPOR. ACL, Santiago de Compostela, 2024b. Cohere. Command A: An Enterprise-Ready Large Language Model, 2025.

DeepSeek. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature 645 (8081), 2025.

Google DeepMind. Gemma 4: Our most capable open models to date. Google Blog, 2026. Accessed: 2026-05-07.

Guo, K. et al. An Investigation of Large Language Models for Real-World Hate Speech Detection, 2024.

Li, L., Fan, L., Atreja, S., and Hemphill, L. “HOT” ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive, and Toxic Comments on Social Media. ACM Trans. Web 18 (2), Mar., 2024.

Li, R., Kamaraj, A., Ma, J., and Ebling, S. Decoding ableism in large language models: An intersectional approach. In Proceedings of the Third Workshop on NLP for Positive Impact. Association for Computational Linguistics, Miami, Florida, USA, pp. 232–249, 2024.

Oliveira, A. et al. Toxic Speech Detection in Portuguese: A Comparative Study of Large Language Models. In Proc.s of the 16th PROPOR. ACL, Santiago de Compostela, pp. 108–116, 2024.

OpenAI. gpt-oss-120b and gpt-oss-20b model card, 2025.

Paiva, L., Assis, G., Amorim, A., Dias, L. G., Paes, A., and de Oliveira, D. Domínio delimitado, Ódio exposto: O uso de prompts para identificação de discurso de Ódio online com llms. In Anais do XL Simpósio Brasileiro de Bancos de Dados. SBC, Porto Alegre, RS, Brasil, pp. 493–506, 2025.

Phutane, M., Seelam, A., and Vashistha, A. “cold, calculated, and condescending”’: How AI identifies and explains ableism compared to disabled people. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. FAccT ’25. Association for Computing Machinery, New York, NY, USA, pp. 1927–1941, 2025.

Rizvi, N. et al. AUTALIC: A dataset for anti-AUTistic ableist language in context. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Vienna, Austria, pp. 20999–21015, 2025.

Zeng, W. et al. ShieldGemma: Generative AI content moderation based on Gemma, 2024.
Publicado
19/10/2026
AMORIM, Annie; ASSIS, Gabriel; PIMENTA, Daniele; PINA, Débora; PAES, Aline; OLIVEIRA, Daniel de. Beyond Toxicity: Evaluating Language Models for Ableism Detection in Online Social Media. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 65-72. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.32055.