Recuperação de documentos para enriquecimento de dados tabulares
Resumo
Técnicas de descoberta de dados focam principalmente na descoberta de dados por meio de palavras-chave, busca por similaridade entre tabelas ou consultas estruturadas. Tais abordagens pouco exploram a conexão entre tabelas e documentos de texto, que pode viabilizar o enriquecimento de dados tabulares a partir de fontes textuais. Neste trabalho, propomos o problema table-to-documents: dada uma tabela como entrada e uma coleção de documentos em um data lake, recuperar os top-k documentos mais relevantes para essa tabela. Nossas contribuições incluem: uma base de dados adaptada à esse problema; uma solução inicial que explora diversas representações tabela-documento; e a avaliação experimental dessa solução.
Referências
Cong, T., Hulsebos, M., Sun, Z., Groth, P., and Jagadish, H. V. (2023). Observatory: Characterizing embeddings of relational tables. PVLDB, 17(4):849–862.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT 2019, pages 4171–4186.
Formal, T., Piwowarski, B., and Clinchant, S. (2021). Splade: Sparse lexical and expansion model for first stage ranking. In SIGIR ’21, pages 2288–2292.
Freitag, D., Cadigan, J., Sasseen, R., and Kalmar, P. (2022). Valet: Rule-based information extraction for rapid deployment. In Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022), pages 524–533.
Hulsebos, M. (2024). Table Representation Learning. Phd thesis, Universiteit van Amsterdam. Defended 23 February 2024.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
Manning, C. D., Raghavan, P., and Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press, Cambridge.
Paton, N. W., Chen, J., and Wu, Z. (2023). Dataset discovery and exploration: A survey. ACM Comput. Surv., 56(4).
Sadia, M., Yang, Z., Xiao, Y., Chen, A., and Roy Chowdhury, A. (2025). SQUiD: Synthesizing relational databases from unstructured text. In EMNLP 2025, pages 31987–32012.
Suhara, Y., Li, J., Li, Y., Zhang, D., Demiralp, c., Chen, C., and Tan, W.-C. (2022). Annotating columns with pre-trained language models. In SIGMOD 2022, pages 1493–1503.
Theologitis, M., Dammu, P. P. S., Shah, C., and Suciu, D. (2026). Claimdb: A fact verification benchmark over large structured data.
Wu, Y., Agarwal, P. K., Li, C., Yang, J., and Yu, C. (2014). Toward computational fact-checking. PVLDB, 7(7):589–600.
Yin, P., Neubig, G., Yih, W.-t., and Riedel, S. (2020). TaBERT: Pretraining for joint understanding of textual and tabular data. In ACL 2020, pages 8413–8426.
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., Zhang, Z., and Radev, D. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In EMNLP, Brussels, Belgium.
Zhang, T., Wang, S., Yan, S., Jian, L., and Liu, Q. (2023). Generative table pre-training empowers models for tabular prediction. In EMNLP 2023, pages 14836–14854.
