Entity Linking in Invoices: Reducing Large Language Model Calls by Using Knowledge-Guided Heuristic Filtering
Resumo
Accurately identifying invoice items is crucial for procurement auditing, but challenging due to noisy, non-standardized descriptions. Large Language Models (LLMs) alone solve this problem with limited performance, and entail high computational costs and privacy concerns. Thus, we propose a hybrid entity linking framework that combines index-based retrieval to match entity surface names in a knowledge base, knowledge-guided heuristic filtering of EL candidates, and an on-premise LLM to solve only persistent ambiguities. Evaluated on Brazilian medicine invoices, our filters reduced LLM inferences by 49.60% and tokens per call by 42.4%. Using qwen2.5:7b, our approach achieved 84.48% accuracy, outperforming a filterless baseline by over 21%.
Palavras-chave:
Entity Linking, Rules, Product Attribute Value Extraction (PAVE), Large Language Models
Referências
Albuquerque, D. L., Santos, V., Nack, P., Fileto, R., and Dorneles, C. (2025). Language models are not a panacea: Combining them with domain knowledge and efficient indexes for entity linking. In Simp. Brasileiro de Bancos de Dados (SBBD), pages 479–492, Porto Alegre, RS, Brasil. SBC.
Ayoola, T., Tyagi, S., Fisher, J., Christodoulopoulos, C., and Pierleoni, A. (2022). ReFinED: An efficient zero-shot-capable approach to end-to-end entity linking. In Loukina, A., Gangadharaiah, R., and Min, B., editors, Conf. of the North American Chapter of the ACL: Human Language Technologies: Industry Track, pages 209–220, Hybrid: Seattle, Washington + Online. Association for Computational Linguistics (ACL).
Balog, K. (2018). Entity-Oriented Search, volume 39 of The Information Retrieval Series. Springer International Publishing.
Ding, Y., Poudel, A., Zeng, Q., Weninger, T., Veeramani, B., and Bhattacharya, S. (2025). Entgpt: Entity linking with generative large language models. arXiv preprint.
Heisler, G., Beckhauser, W., Santos, V., and Fileto, R. (2025). Detection of vehicle purchases in various invoices using large-scale language models. In Anais da I Escola Regional de Aprendizado de Máquina e Inteligência Artificial da Região Sul, pages 164–167, Porto Alegre, RS, Brasil. SBC.
Li, Y., Galimov, A., Ganapaneni, M. D., Thejaswi, P., Meng, D., Kumar, P., and Potdar, S. (2025). Leveraging the power of large language models in entity linking via adaptive routing and targeted reasoning. In Potdar, S., Rojas-Barahona, L., and Montella, S., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 871–882, Suzhou (China). Association for Computational Linguistics.
Liu, X., Liu, Y., Zhang, K., Wang, K., Liu, Q., and Chen, E. (2024). Onenet: A fine-tuning free framework for few-shot entity linking via large language model prompting. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 13634–13651, Miami, Florida, USA. Association for Computational Linguistics.
Romero, P., Han, L., and Nenadic, G. (2025). INSIGHTBUDDY-AI: Medication extraction and entity linking using pre-trained language models and ensemble learning. In Ebrahimi, A., Haider, S., Liu, E., Haider, S., Leonor Pacheco, M., and Wein, S., editors, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop), pages 18–27, Albuquerque, USA. Association for Computational Linguistics.
Rynkiewicz, A. A., Palma, R., and Formanowicz, P. (2025). Universal entity linking. Engineering Applications of Artificial Intelligence, 161:112185.
Shen, W., Wang, J., and Han, J. (2015). Entity linking with a knowledge base: Issues, techniques, and solutions. IEEE Transactions on Knowledge and Data Engineering, 27(2):443–460.
Shlyk, D., Groza, T., Mesiti, M., Montanelli, S., and Cavalleri, E. (2024). REAL: A retrieval-augmented entity linking approach for biomedical concept recognition. In Demner-Fushman, D., Ananiadou, S., Miwa, M., Roberts, K., and Tsujii, J., editors, Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, pages 380–389, Bangkok, Thailand. Association for Computational Linguistics.
Tedeschi, S., Conia, S., Cecconi, F., and Navigli, R. (2021). Named Entity Recognition for Entity Linking: What works and what’s next. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t., editors, Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2584–2596, Punta Cana, Dominican Republic. Association for Computational Linguistics.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. (2023). Llama: Open and efficient foundation language models.
Vollmers, D., Zahera, H., Moussallem, D., and Ngonga Ngomo, A.-C. (2025). Contextual augmentation for entity linking using large language models. In Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., and Schockaert, S., editors, Proceedings of the 31st International Conference on Computational Linguistics, pages 8535–8545, Abu Dhabi, UAE. Association for Computational Linguistics.
Wang, F., Tao, Z., Wang, M., Hu, M., and Bai, X. (2025). AELC: Adaptive entity linking with LLM-driven contextualization. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V., editors, Findings of the Association for Computational Linguistics: EMNLP 2025, pages 4313–4327, Suzhou, China. Association for Computational Linguistics.
Xin, A., Qi, Y., Yao, Z., Zhu, F., Zeng, K., Xu, B., Hou, L., and Li, J. (2025). Llmael: Large language models are good context augmenters for entity linking. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management, CIKM ’25, page 3550–3559, New York, NY, USA. Association for Computing Machinery.
Ye, C. and Mitchell, C. S. (2025). LLM as entity disambiguator for biomedical entity-linking. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T., editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 301–312, Vienna, Austria. Association for Computational Linguistics.
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., and Chi, E. (2023). Least-to-most prompting enables complex reasoning in large language models.
Zhou, K., Li, Y., Wang, Q., Qiao, Q., and Li, Q. (2024). GenDecider: Integrating “none of the candidates” judgments in zero-shot entity linking re-ranking. In Duh, K., Gomez, H., and Bethard, S., editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 239–245, Mexico City, Mexico. Association for Computational Linguistics.
Ayoola, T., Tyagi, S., Fisher, J., Christodoulopoulos, C., and Pierleoni, A. (2022). ReFinED: An efficient zero-shot-capable approach to end-to-end entity linking. In Loukina, A., Gangadharaiah, R., and Min, B., editors, Conf. of the North American Chapter of the ACL: Human Language Technologies: Industry Track, pages 209–220, Hybrid: Seattle, Washington + Online. Association for Computational Linguistics (ACL).
Balog, K. (2018). Entity-Oriented Search, volume 39 of The Information Retrieval Series. Springer International Publishing.
Ding, Y., Poudel, A., Zeng, Q., Weninger, T., Veeramani, B., and Bhattacharya, S. (2025). Entgpt: Entity linking with generative large language models. arXiv preprint.
Heisler, G., Beckhauser, W., Santos, V., and Fileto, R. (2025). Detection of vehicle purchases in various invoices using large-scale language models. In Anais da I Escola Regional de Aprendizado de Máquina e Inteligência Artificial da Região Sul, pages 164–167, Porto Alegre, RS, Brasil. SBC.
Li, Y., Galimov, A., Ganapaneni, M. D., Thejaswi, P., Meng, D., Kumar, P., and Potdar, S. (2025). Leveraging the power of large language models in entity linking via adaptive routing and targeted reasoning. In Potdar, S., Rojas-Barahona, L., and Montella, S., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 871–882, Suzhou (China). Association for Computational Linguistics.
Liu, X., Liu, Y., Zhang, K., Wang, K., Liu, Q., and Chen, E. (2024). Onenet: A fine-tuning free framework for few-shot entity linking via large language model prompting. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 13634–13651, Miami, Florida, USA. Association for Computational Linguistics.
Romero, P., Han, L., and Nenadic, G. (2025). INSIGHTBUDDY-AI: Medication extraction and entity linking using pre-trained language models and ensemble learning. In Ebrahimi, A., Haider, S., Liu, E., Haider, S., Leonor Pacheco, M., and Wein, S., editors, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop), pages 18–27, Albuquerque, USA. Association for Computational Linguistics.
Rynkiewicz, A. A., Palma, R., and Formanowicz, P. (2025). Universal entity linking. Engineering Applications of Artificial Intelligence, 161:112185.
Shen, W., Wang, J., and Han, J. (2015). Entity linking with a knowledge base: Issues, techniques, and solutions. IEEE Transactions on Knowledge and Data Engineering, 27(2):443–460.
Shlyk, D., Groza, T., Mesiti, M., Montanelli, S., and Cavalleri, E. (2024). REAL: A retrieval-augmented entity linking approach for biomedical concept recognition. In Demner-Fushman, D., Ananiadou, S., Miwa, M., Roberts, K., and Tsujii, J., editors, Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, pages 380–389, Bangkok, Thailand. Association for Computational Linguistics.
Tedeschi, S., Conia, S., Cecconi, F., and Navigli, R. (2021). Named Entity Recognition for Entity Linking: What works and what’s next. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t., editors, Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2584–2596, Punta Cana, Dominican Republic. Association for Computational Linguistics.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. (2023). Llama: Open and efficient foundation language models.
Vollmers, D., Zahera, H., Moussallem, D., and Ngonga Ngomo, A.-C. (2025). Contextual augmentation for entity linking using large language models. In Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., and Schockaert, S., editors, Proceedings of the 31st International Conference on Computational Linguistics, pages 8535–8545, Abu Dhabi, UAE. Association for Computational Linguistics.
Wang, F., Tao, Z., Wang, M., Hu, M., and Bai, X. (2025). AELC: Adaptive entity linking with LLM-driven contextualization. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V., editors, Findings of the Association for Computational Linguistics: EMNLP 2025, pages 4313–4327, Suzhou, China. Association for Computational Linguistics.
Xin, A., Qi, Y., Yao, Z., Zhu, F., Zeng, K., Xu, B., Hou, L., and Li, J. (2025). Llmael: Large language models are good context augmenters for entity linking. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management, CIKM ’25, page 3550–3559, New York, NY, USA. Association for Computing Machinery.
Ye, C. and Mitchell, C. S. (2025). LLM as entity disambiguator for biomedical entity-linking. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T., editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 301–312, Vienna, Austria. Association for Computational Linguistics.
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., and Chi, E. (2023). Least-to-most prompting enables complex reasoning in large language models.
Zhou, K., Li, Y., Wang, Q., Qiao, Q., and Li, Q. (2024). GenDecider: Integrating “none of the candidates” judgments in zero-shot entity linking re-ranking. In Duh, K., Gomez, H., and Bethard, S., editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 239–245, Mexico City, Mexico. Association for Computational Linguistics.
Publicado
08/09/2026
Como Citar
AZEVEDO, Pedro H.; FILETO, Renato; ZIBETTI, André W.; WERNER, Simone S..
Entity Linking in Invoices: Reducing Large Language Model Calls by Using Knowledge-Guided Heuristic Filtering. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 659-672.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249286.
