Context Matters, Bias Matters Too: Assessing Fairness in Profile Embeddings in the HateBR Dataset

  • Vinícius Moitinho Universidade Federal de Sergipe (UFS)

Resumo


Offensive language detection is challenged by discourse’s contextual nature, especially in Brazilian political contexts. This work examines how domain-specific information affects classifying offensiveness toward public figures, using the HateBR corpus. We compare three BERTimbau-based approaches: a domain-agnostic baseline, a model with textual metadata injection, and a hybrid with latent profile embeddings. Results show profile embeddings outperform other approaches in accuracy and convergence, indicating the model learns target-specific patterns. However, fairness analysis reveals a trade-off: specialization increases False Positive Rate variance across profiles, suggesting the model absorbs dataset biases. Findings reinforce the need for models integrating equity and context sensitivity in online debate.
Palavras-chave: Hate Speech Detection, Algorithmic Fairness, Political Discourse, BERTimbau, Profile Embeddings, Bias in NLP

Referências

Badjatiya, P., Gupta, S., Gupta, M., and Varma, V. (2017). Deep learning for hate speech detection in tweets. In Proceedings of the 26th International Conference on World Wide Web Companion, pages 759–760.

Badri, N., Kboubi, F., and Chaibi, A. H. (2022). Combining fasttext and glove word embeddings for offensive and hate speech text detection. Procedia Computer Science, 207:769–778.

Chen, Y., Zhou, Y., Zhu, S., and Xu, H. (2012). Detecting offensive language in social media to protect adolescent online safety. In 2012 International Conference on Privacy, Security, Risk and Trust, pages 71–80. IEEE.

Dai, W., Li, W., Li, R., Ji, D., and Li, C. (2020). Kungfupanda at semeval-2020 task 12: Bert-based multi-task learning for offensive language detection. arXiv preprint arXiv:2004.13432.

Davidson, T., Warmsley, D., Macy, M., and Weber, I. (2017). Automated hate speech detection and the problem of offensive language. In Proceedings of the 11th International AAAI Conference on Web and Social Media (ICWSM), pages 512–515.

de Araújo Miranda, A. L. and de Oliveira Rodrigues, C. M. (2025). Uma abordagem integrada para detecção de discurso de ódio em mídias sociais utilizando vetorização de textos e emojis. In Anais do Workshop sobre as Implicações da Computação na Sociedade (WICS), pages 247–255. SBC.

Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.

Fortuna, P. and Nunes, S. (2018). A survey on automatic detection of hate speech in text. ACM Computing Surveys, 51(4):1–30.

Haddad, D., Monteiro, R., and Gonçalves, T. (2021). Hatebr: An annotated corpus for hate speech detection in brazilian portuguese. In Proceedings of the 11th Brazilian Symposium in Information and Human Language Technology (STIL).

Parafino, C., Tomaszewski, J., de Freitas, L. A., and Santana, B. S. (2025). Discursos de ódio e notícias falsas: Estudo preliminar da coocorrência entre esses fenômenos em textos em português. In Anais da Escola Regional de Aprendizado de Máquina e Inteligência Artificial da Região Sul (ERAMIA-RS), pages 180–183. SBC.

Pelle, R. M. and Moreira, V. P. (2017). Offensive comments in the brazilian web: a dataset and analysis. In Proceedings of the 8th Brazilian Symposium on Information and Human Language Technology (STIL).

Sap, M., Card, D., Gabriel, S., Choi, Y., and Smith, N. A. (2019). The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678.

Schmidt, A. and Wiegand, M. (2017). A survey on hate speech detection using natural language processing. In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, pages 1–10.

Silva, G., Silva, P., Peixoto, M., Moreira, G., and Luz, E. (2026). A multitask transformer for offensive language detection and target identification in hatebr. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 1049–1054.

Talat, Z. and Hovy, D. (2016). Hateful symbols or hateful people? predictive features for hate speech detection on twitter. In Proceedings of the NAACL Student Research Workshop, pages 88–93.
Publicado
08/09/2026
MOITINHO, Vinícius. Context Matters, Bias Matters Too: Assessing Fairness in Profile Embeddings in the HateBR Dataset. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 550-560. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249257.