Multimodal RAG and Parsing: Data Architectures for Semantic Retrieval in Complex Industrial Layouts
Resumo
Traditional RAG systems struggle with industrial documents by neglecting structural semantics in complex layouts. To mitigate this, we propose a formal Visual Document Understanding ingestion methodology and a Semantic Chunking framework. By orchestrating Vision-Language Models to generate high-fidelity captions for visual elements and grouping them with adjacent text, our pipeline preserves context-rich retrieval units. Experiments on an industrial dataset demonstrate 88.0% accuracy, a 0.91 Hit Rate@5, and 0.89 Context Recall, significantly outperforming OCR baselines (74.0% accuracy) and matching state-of-the-art solutions.
Palavras-chave:
Retrieval-Augmented Generation, Visual Document Understanding, Document Parsing
Referências
An, Z., Ding, X., Fu, Y.-C., Chu, C.-C., Li, Y., and Du, W. (2024). Golden-retriever: High-fidelity agentic retrieval augmented generation for industrial knowledge base. arXiv preprint arXiv:2408.00798.
DeepMind, G. (2025a). Gemini 3 flash: Frontier intelligence built for speed. [link].
DeepMind, G. (2025b). Gemini 3 pro: State-of-the-art multimodal ai model. [link].
Gao, S., Zhao, S., Jiang, X., Duan, L., Chng, Y. X., Chen, Q.-G., Luo, W., Zhang, K., Bian, J.-W., and Gong, M. (2025). Scaling beyond context: A survey of multimodal retrieval-augmented generation for document understanding. arXiv preprint arXiv:2510.15253.
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., and Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
Gu, J., Kuen, J., Morariu, V. I., Zhao, H., Barmpalios, N., Jain, R., Nenkova, A., and Sun, T. (2021). Unified pretraining framework for document understanding. In Advances in Neural Information Processing Systems (NeurIPS 2021).
Kim, G., Hong, T., Yim, M., Nam, J., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S. (2022). Ocr-free document understanding transformer. In Proceedings of the 17th European Conference on Computer Vision (ECCV 2022), pages 498–517, Berlin, Heidelberg. Springer-Verlag.
Mandal, S., Talewar, A., Ahuja, P., and Juvatkar, P. (2025). Nanonets-ocr-s: A model for transforming documents into structured markdown with intelligent content recognition and semantic tagging. [link].
McKie, J. X. (2024). PyMuPDF: Python bindings for the MuPDF library. Version 1.24.x.
Mei, L., Mo, S., Yang, Z., and Chen, C. (2025). A survey of multimodal retrieval-augmented generation. arXiv preprint arXiv:2504.08748.
Qwen (2025). Qwen2.5-vl-3b: Instruction-tuned vision-language model. [link].
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., and Zhou, M. (2020). Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, pages 1192–1200, New York, NY, USA. Association for Computing Machinery.
DeepMind, G. (2025a). Gemini 3 flash: Frontier intelligence built for speed. [link].
DeepMind, G. (2025b). Gemini 3 pro: State-of-the-art multimodal ai model. [link].
Gao, S., Zhao, S., Jiang, X., Duan, L., Chng, Y. X., Chen, Q.-G., Luo, W., Zhang, K., Bian, J.-W., and Gong, M. (2025). Scaling beyond context: A survey of multimodal retrieval-augmented generation for document understanding. arXiv preprint arXiv:2510.15253.
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., and Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
Gu, J., Kuen, J., Morariu, V. I., Zhao, H., Barmpalios, N., Jain, R., Nenkova, A., and Sun, T. (2021). Unified pretraining framework for document understanding. In Advances in Neural Information Processing Systems (NeurIPS 2021).
Kim, G., Hong, T., Yim, M., Nam, J., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S. (2022). Ocr-free document understanding transformer. In Proceedings of the 17th European Conference on Computer Vision (ECCV 2022), pages 498–517, Berlin, Heidelberg. Springer-Verlag.
Mandal, S., Talewar, A., Ahuja, P., and Juvatkar, P. (2025). Nanonets-ocr-s: A model for transforming documents into structured markdown with intelligent content recognition and semantic tagging. [link].
McKie, J. X. (2024). PyMuPDF: Python bindings for the MuPDF library. Version 1.24.x.
Mei, L., Mo, S., Yang, Z., and Chen, C. (2025). A survey of multimodal retrieval-augmented generation. arXiv preprint arXiv:2504.08748.
Qwen (2025). Qwen2.5-vl-3b: Instruction-tuned vision-language model. [link].
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., and Zhou, M. (2020). Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, pages 1192–1200, New York, NY, USA. Association for Computing Machinery.
Publicado
08/09/2026
Como Citar
MIYAJI, Renato; LUZ, Arthur; MOULIN, Renato; MACHADO, Leonardo; CORRÊA, Pedro L. P..
Multimodal RAG and Parsing: Data Architectures for Semantic Retrieval in Complex Industrial Layouts. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 1008-1014.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249655.
