Large Language Model como Suporte à Tarefa de Entity Matching
Resumo
Fundamental para a integração de dados, a tarefa de Entity Matching visa identificar registros que referenciam a mesma entidade no mundo real. Modelos pré-treinados dominam o estado da arte, mas enfrentam limitações quanto à necessidade de grandes volumes de dados rotulados, enquanto que Large Language Models (LLMs) demonstram capacidade superior de generalização, sem a necessidade de grandes volumes, apesar do alto custo computacional. Diante deste cenário, o presente trabalho propõe uma abordagem híbrida para equilibrar o melhor dos dois mundos de modelos, realizando novas classificações com os LLMs apenas em amostras em que o modelo pré-treinado apresente baixa confiança, gerando um aumento de até 8% de F1-Score nos resultados.
Palavras-chave:
Entity Matching, Pre-Trained Models, Large Language Models
Referências
Akbarian Rastaghi, M., Kamalloo, E., and Rafiei, D. (2022). Probing the robustness of pre-trained language models for entity matching. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, CIKM ’22, page 3786–3790, New York, NY, USA. Association for Computing Machinery.
Araújo, T. B., Efthymiou, V., Christophides, V., Pitoura, E., and Stefanidis, K. (2025a). Treats: Fairness-aware entity resolution over streaming data. Information Systems, 129:102506.
Araújo, T. B., Efthymiou, V., and Stefanidis, K. (2025b). Fairness and explanations in entity resolution: An overview. IEEE Access.
Araújo, T. B., Stefanidis, K., Santos Pires, C. E., Nummenmaa, J., and Da Nóbrega, T. P. (2020). Schema-agnostic blocking for streaming data. In Proceedings of the 35th Annual ACM Symposium on Applied Computing, pages 412–419.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020). Language models are few-shot learners.
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., Ye, W., Zhang, Y., Chang, Y., Yu, P. S., Yang, Q., and Xie, X. (2024). A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol., 15(3).
Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., and Stefanidis, K. (2019). End-to-end entity resolution for big data: A survey. arXiv preprint arXiv:1905.06397.
Donato, R. B. and Araújo, T. B. (2025). Entity matching com large language models: estudo comparativo com abordagem de entity blocking. In Anais do XL Simpósio Brasileiro de Bancos de Dados, pages 956–962, Porto Alegre, RS, Brasil. SBC.
Ebraheem, M., Thirumuruganathan, S., Joty, S., Ouzzani, M., and Tang, N. (2018). Distributed representations of tuples for entity resolution. Proc. VLDB Endow., 11(11):1454–1467.
Li, Y., Li, J., Suhara, Y., Doan, A., and Tan, W. (2020). Deep entity matching with pre-trained language models. CoRR, abs/2004.00584.
Maciejewski, J., Nikoletos, K., Papadakis, G., and Velegrakis, Y. (2025). Progressive entity matching: A design space exploration. Proc. ACM Manag. Data, 3(1).
Mestre, D. G., Pires, C. E. S., Nascimento, D. C., de Queiroz, A. R. M., Santos, V. B., and Araujo, T. B. (2017). An efficient spark-based adaptive windowing for entity matching. Journal of Systems and Software, 128:1–10.
Papadakis, G., Efthymiou, V., Thanos, E., Hassanzadeh, O., and Christen, P. (2023). An analysis of one-to-one matching algorithms for entity resolution.
Peeters, R., Der, R. C., and Bizer, C. (2023). Wdc products: A multi-dimensional entity matching benchmark.
Peeters, R., Steiner, A., and Bizer, C. (2024). Entity matching using large language models.
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. (2022). Emergent abilities of large language models.
Xia, Y., Chen, J., Li, X., and Gao, J. (2024). Aprompt4em: Augmented prompt tuning for generalized entity matching.
Zeakis, A., Papadakis, G., Skoutas, D., and Koubarakis, M. (2025). An in-depth analysis of pre-trained embeddings for entity resolution: An in-depth analysis of pre-trained embeddings for entity resolution: A. zeakis et al. VLDB Journal International Journal on Very Large Data Bases, 34(1).
Zhang, Z., Groth, P., Calixto, I., and Schelter, S. (2025). A deep dive into cross-dataset entity matching with large and small language models. In Advances in Database Technology - EDBT, number 3 in Advances in Database Technology - EDBT, pages 922–934. OpenProceedings.org.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. (2025). A survey of large language models.
Araújo, T. B., Efthymiou, V., Christophides, V., Pitoura, E., and Stefanidis, K. (2025a). Treats: Fairness-aware entity resolution over streaming data. Information Systems, 129:102506.
Araújo, T. B., Efthymiou, V., and Stefanidis, K. (2025b). Fairness and explanations in entity resolution: An overview. IEEE Access.
Araújo, T. B., Stefanidis, K., Santos Pires, C. E., Nummenmaa, J., and Da Nóbrega, T. P. (2020). Schema-agnostic blocking for streaming data. In Proceedings of the 35th Annual ACM Symposium on Applied Computing, pages 412–419.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020). Language models are few-shot learners.
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., Ye, W., Zhang, Y., Chang, Y., Yu, P. S., Yang, Q., and Xie, X. (2024). A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol., 15(3).
Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., and Stefanidis, K. (2019). End-to-end entity resolution for big data: A survey. arXiv preprint arXiv:1905.06397.
Donato, R. B. and Araújo, T. B. (2025). Entity matching com large language models: estudo comparativo com abordagem de entity blocking. In Anais do XL Simpósio Brasileiro de Bancos de Dados, pages 956–962, Porto Alegre, RS, Brasil. SBC.
Ebraheem, M., Thirumuruganathan, S., Joty, S., Ouzzani, M., and Tang, N. (2018). Distributed representations of tuples for entity resolution. Proc. VLDB Endow., 11(11):1454–1467.
Li, Y., Li, J., Suhara, Y., Doan, A., and Tan, W. (2020). Deep entity matching with pre-trained language models. CoRR, abs/2004.00584.
Maciejewski, J., Nikoletos, K., Papadakis, G., and Velegrakis, Y. (2025). Progressive entity matching: A design space exploration. Proc. ACM Manag. Data, 3(1).
Mestre, D. G., Pires, C. E. S., Nascimento, D. C., de Queiroz, A. R. M., Santos, V. B., and Araujo, T. B. (2017). An efficient spark-based adaptive windowing for entity matching. Journal of Systems and Software, 128:1–10.
Papadakis, G., Efthymiou, V., Thanos, E., Hassanzadeh, O., and Christen, P. (2023). An analysis of one-to-one matching algorithms for entity resolution.
Peeters, R., Der, R. C., and Bizer, C. (2023). Wdc products: A multi-dimensional entity matching benchmark.
Peeters, R., Steiner, A., and Bizer, C. (2024). Entity matching using large language models.
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. (2022). Emergent abilities of large language models.
Xia, Y., Chen, J., Li, X., and Gao, J. (2024). Aprompt4em: Augmented prompt tuning for generalized entity matching.
Zeakis, A., Papadakis, G., Skoutas, D., and Koubarakis, M. (2025). An in-depth analysis of pre-trained embeddings for entity resolution: An in-depth analysis of pre-trained embeddings for entity resolution: A. zeakis et al. VLDB Journal International Journal on Very Large Data Bases, 34(1).
Zhang, Z., Groth, P., Calixto, I., and Schelter, S. (2025). A deep dive into cross-dataset entity matching with large and small language models. In Advances in Database Technology - EDBT, number 3 in Advances in Database Technology - EDBT, pages 922–934. OpenProceedings.org.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. (2025). A survey of large language models.
Publicado
08/09/2026
Como Citar
DONATO, Rodolfo Bolconte; ARAÚJO, Tiago Brasileiro; PIRES, Carlos Eduardo S.; MESTRE, Demetrio Gomes.
Large Language Model como Suporte à Tarefa de Entity Matching. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 645-658.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249284.
