Um Dataset Sintético Multi-Cenário para Avaliação de Fairness em Sistemas de Recomendação de Recrutamento

Resumo


A ampla adoção de Sistemas de Recomendação no e-recrutamento levantou preocupações significativas quanto à justiça (fairness) e viés contra grupos demográficos minoritários. A falta de conjuntos de dados padronizados que incorporem métricas objetivas de adequação e cenários explícitos de viés dificulta o desenvolvimento e a avaliação de técnicas de mitigação. Neste artigo, apresentamos um dataset sintético multicenário para e-recrutamento, com controle explícito de injeção de viés em interseções demográficas (gênero, raça e localização), com versões variando em tamanho e condições de viés. Descrevemos a metodologia de geração, apresentamos estatísticas descritivas, discutimos limitações do dataset e demonstramos sua utilidade para o benchmarking de algoritmos de fairness em recursos humanos. O dataset e o código de geração estão publicamente disponíveis em nossos repositórios no GitHub.
Palavras-chave: Sistemas de Recomendação, E-recrutamento, Recrutamento Automatizado, Fairness, Justiça Algorítmica, Viés Interseccional, Dataset Sintético

Referências

Angwin, J., Larson, J., Mattu, S., and Kirchner, L. (2016). Machine bias. ProPublica.

Barocas, S., Hardt, M., and Narayanan, A. (2023). Fairness and Machine Learning: Limitations and Opportunities. MIT Press.

Becker, B. and Kohavi, R. (1996). Adult data set. UCI Machine Learning Repository.

Bird, S., Dudík, M., Edgar, R., Horn, B., Lutz, R., Milan, V., Sameki, M., Wallach, H., and Walker, K. (2020). Fairlearn: A toolkit for assessing and improving fairness in AI. In Microsoft Research Technical Report.

Bogen, M. and Rieke, A. (2018). Help wanted: An examination of hiring algorithms, equity, and bias. Upturn.

Burkardt, J. (2014). The truncated normal distribution. Department of Scientific Computing, Florida State University, Technical Report.

Celis, L. E., Straszak, D., and Vishnoi, N. K. (2018). Ranking with fairness constraints. In Proceedings of the 45th International Colloquium on Automata, Languages, and Programming (ICALP). Schloss Dagstuhl.

Choi, K., Grover, A., Singh, T., Shu, P., and Ermon, S. (2020). Fair generative modeling via weak supervision. In Proceedings of the International Conference on Machine Learning (ICML).

Cowgill, B. and Tucker, C. E. (2020). Bias and productivity in humans and algorithms: Theory and evidence from resume screening. Working Paper 28243, National Bureau of Economic Research (NBER).

Crenshaw, K. (1989). Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. University of Chicago Legal Forum, 1989(1):139–167.

Dastin, J. (2018). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters.

Dwork, C., Hardt, M., Pitassi, T., Reingold, O., and Zemel, R. (2012). Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference (ITCS), pages 214–226. ACM.

González, S., Frery, A. C., and Rosales-Salas, J. (2023). Synthetic data generation for fairness-aware machine learning: A survey. arXiv preprint arXiv:2307.10306.

Hofmann, H. (1994). Statlog (German credit data). UCI Machine Learning Repository.

IBGE (2022). Pesquisa nacional por amostra de domicílios contínua – TIC 2022. Technical report, Instituto Brasileiro de Geografia e Estatística.

Kearns, M. and Roth, A. (2019). The Ethical Algorithm: The Science of Socially Aware Algorithm Design. Oxford University Press.

Kenthapadi, K., Le, B., and Venkataraman, G. (2017). Personalized job recommendation system at LinkedIn: Practical challenges and lessons learned. In Proceedings of the 11th ACM Conference on Recommender Systems (RecSys), pages 346–347. ACM.

Little, R. J. A. and Rubin, D. B. (2019). Statistical Analysis with Missing Data. John Wiley & Sons, 3rd edition.

Lund, B. D. and Wang, T. (2023). Chatting about ChatGPT: How may AI and GPT impact academia and libraries? Library Hi Tech News, 40(3):26–29.

Lundberg, S. M. and Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (NeurIPS), volume 30.

Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1–35.

Menon, A. K. and Williamson, R. C. (2018). The cost of fairness in binary classification. Proceedings of the 1st Conference on Fairness, Accountability and Transparency (FAccT), pages 107–118.

Nature Editorial (2023). Tools such as ChatGPT threaten transparent science; here are our ground rules for their use. Nature, 613(7945):612.

Raghavan, M., Barocas, S., Kleinberg, J., and Levy, K. (2020). Mitigating bias in algorithmic hiring: Evaluating claims and practices. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAccT), pages 469–481.

Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pages 1135–1144. ACM.

Singh, A. and Joachims, T. (2018). Fairness of exposure in rankings. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD), pages 2219–2228. ACM.

van Breugel, B., Qian, Z., and van der Schaar, M. (2023). Beyond privacy: Navigating the opportunities and challenges of synthetic data. arXiv preprint arXiv:2304.03722.

Xu, D., Yuan, S., Zhang, L., and Wu, X. (2018). FairGAN: Fairness-aware generative adversarial networks. In 2018 IEEE International Conference on Big Data (Big Data), pages 570–575. IEEE.

Xu, L., Skoularidou, M., Cuesta-Infante, A., and Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. In Advances in Neural Information Processing Systems (NeurIPS), volume 32.
Publicado
08/09/2026
CARVALHO, Luiz H.; FORTES, Reinaldo Silva. Um Dataset Sintético Multi-Cenário para Avaliação de Fairness em Sistemas de Recomendação de Recrutamento. In: DATASET SHOWCASE WORKSHOP (DSW), 8. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 102-113. DOI: https://doi.org/10.5753/dsw.2026.249573.