Severe Name Homonymy: A Learning to Rank Approach for Civil Registry Entity Linking

  • Marco T. Dutra Universidade Federal de Minas Gerais (UFMG) https://orcid.org/0009-0000-3865-5799
  • Lucas G. L. Costa Universidade Federal de Minas Gerais (UFMG) https://orcid.org/0009-0002-8898-4237
  • Felipe D. S. Martins Universidade Federal de Minas Gerais (UFMG)
  • Gabriela M. M. Freitas Universidade Federal de Minas Gerais (UFMG)
  • Leonardo C. C. Castro Universidade Federal de Minas Gerais (UFMG)
  • Pedro G. B. Vieira Universidade Federal de Minas Gerais (UFMG)
  • Michele A. Brandão Institudo Federal de Minas Gerais (IFMG)
  • Gisele L. Pappa Universidade Federal de Minas Gerais (UFMG)

Resumo


This study addresses the challenge of transforming string-based kinship references into explicit entity relationships within large-scale civil registries. In scenarios where a record provides only the name of a relative (e.g., a mother’s name), identifying the correct unique identifier among numerous homonymous candidates constitutes a complex Known-Item Search problem. We propose a supervised Learning to Rank framework designed to disambiguate and rank these candidates by integrating kinship-based relational features, conditional probability signals, demographic indicators, and BERT-based semantic similarity. The approach was evaluated on a large-scale private dataset from Brazil and four public datasets: IPUMS (covering USA, Norway, and Iceland) and Wikidata. Experimental results demonstrate that our method provides a robust solution for identity resolution under high name ambiguity, achieving a Recall@5 above 0.90 across all evaluated datasets and enabling the construction of accurate entity links in civil databases.
Palavras-chave: Information Retrieval, Entity Linking, Learning to Rank, Known-Item Search, Civil Registry

Referências

Albuquerque, D. L., Santos, V., Nack, P., et al. (2025). Language models are not a panacea: Combining them with domain knowledge and efficient indexes for entity linking. In Anais do XL Simpósio Brasileiro de Bancos de Dados, Porto Alegre, RS, Brasil. SBC.

Ali Omar, Z., Zamzuri, Z. H., Mohd Ariff, N., and Abu Bakar, M. A. (2023). Training data selection for record linkage classification. Symmetry, 15(5).

Archive, T. D., Centre, N. H. D., and Center, M. P. (2008). National Sample of the 1900 Census of Norway, Version 2.0.

Arora, A. and Dell, M. (2024). LinkTransformer: A unified package for record linkage with transformer language models. In Cao, Y., Feng, Y., and Xiong, D., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 221–231, Bangkok, Thailand. Association for Computational Linguistics.

Chen, T. and Guestrin, C. (2016). Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, pages 785–794.

Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer.

Garðarsdóttir, (n.d.). 1729 Census of Iceland, Version 1.0.

Ke, G., Meng, Q., Finley, T., et al. (2017). Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems, 30.

Liu, T.-Y. (2009). Learning to rank for information retrieval. Found. Trends Inf. Retr., 3(3):225–331.

Lundberg, S. M., Erion, G., Chen, H., et al. (2020). From local explanations to global understanding with explainable ai for trees. Nature machine intelligence, 2(1):56–67.

Lundberg, S. M. and Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in neural information processing systems, 30.

Marone, M., Weller, O., Fleshman, W., Yang, E., Lawrie, D., and Durme, B. V. (2025). mmbert: A modern multilingual encoder with annealed language learning.

Mon Myint, K. S. and Naing, W. W. (2023). Historical census data linkage: Graph-based household matching method. In 2023 IEEE Conference on Computer Applications (ICCA), pages 351–356.

Mrini, K., Nie, S., Gu, J., et al. (2022). Detection, disambiguation, re-ranking: Autoregressive entity linking as a multi-task problem. In Muresan, S., Nakov, P., and Villavicencio, A., editors, Findings of the Association for Computational Linguistics: ACL 2022, pages 1972–1983, Dublin, Ireland. Association for Computational Linguistics.

Nanayakkara, C. and Christen, P. (2022). Locality sensitive hashing with temporal and spatial constraints for efficient population record linkage. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, CIKM ’22, page 4354–4358, New York, NY, USA. Association for Computing Machinery.

Nations, U. (2014). Principles and Recommendations for a Vital Statistics System, Revision 3. United Nations.

Ornstein, J. T. (2025). Probabilistic record linkage using pretrained text embeddings. Political Analysis, page 1–12.

Price, J., Buckles, K., Van Leeuwen, J., and Riley, I. (2021). Combining family history and machine learning to link historical records: The census tree data set. Explorations in Economic History, 80:101391.

Ruggles, S., Cleveland, L. L., Lovatón Dávila, R., et al. (2025a). IPUMS International: Version 7.7 [dataset].

Ruggles, S., Flood, S., Sobek, M., et al. (2025b). IPUMS USA: Version 16.0 [dataset].

Santos, M. and Nascimento, D. (2023). Avaliando fatores de influência sobre algoritmos de aprendizado de máquina na etapa de classificação da resolução de entidades. In Anais do XXXVIII Simpósio Brasileiro de Bancos de Dados, pages 63–75, Porto Alegre, RS, Brasil. SBC.

Scharpf, P., Breitinger, C., Spitz, A., et al. (2026). Entity linking with wikidata: A systematic literature review. ACM Comput. Surv., 58(9).

Schütze, H., Manning, C. D., and Raghavan, P. (2008). Introduction to information retrieval, volume 39. Cambridge University Press Cambridge.

Shen, W., Wang, J., and Han, J. (2014). Entity linking with a knowledge base: Issues, techniques, and solutions. IEEE Transactions on Knowledge and Data Engineering, 27(2):443–460.

Song, J., Nanayakkara, C., and Christen, P. (2025). Privately evaluating sensitive population record linkage without ground truth data. International Journal of Data Science and Analytics, 20:2971–2986.

Tian, S., Chen, Q., Comeau, D. C., et al. (2024). Pubmed computed authors in 2024: an open resource of disambiguated author names in biomedical literature. Bioinformatics, 40(11).

Vrandecic, D. and Krötzsch, M. (2014). Wikidata: A free collaborative knowledgebase. Communications of the ACM, 57(10):78–85.

Wu, L., Petroni, F., Josifoski, M., et al. (2020). Scalable zero-shot entity linking with dense entity retrieval. In Webber, B., Cohn, T., He, Y., and Liu, Y., editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6397–6407, Online. Association for Computational Linguistics.

Zhang, T., Kishore, V., Wu, F., et al. (2019). Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.

Zhou, Q., Chen, W., Zhao, P.-P., et al. (2024). Towards effective author name disambiguation by hybrid attention. Journal of Computer Science and Technology, 39(4):929–950.

Ítalo Pereira and Ferreira, A. (2024). E-bela: Enhanced embedding-based entity linking approach. In Proceedings of the 30th Brazilian Symposium on Multimedia and the Web, pages 115–123, Porto Alegre, RS, Brasil. SBC.
Publicado
08/09/2026
DUTRA, Marco T.; COSTA, Lucas G. L.; MARTINS, Felipe D. S.; FREITAS, Gabriela M. M.; CASTRO, Leonardo C. C.; VIEIRA, Pedro G. B.; BRANDÃO, Michele A.; PAPPA, Gisele L.. Severe Name Homonymy: A Learning to Rank Approach for Civil Registry Entity Linking. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 617-630. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249274.