MulitaMiner: A Multi-Version Evaluation of LLM-Based Vulnerability Report Extraction
Resumo
Security teams face 131 new CVEs per day, yet the National Vulnerability Database fails to enrich 44% of new submissions, leaving practitioners without the structured metadata required to prioritize remediation. MulitaMiner closes this gap: an LLM-based pipeline for structured vulnerability extraction from heterogeneous scanner PDFs, hardened across three successive versions (V1–V3) that evolve from a functional baseline, to scanner-aware segmentation, and finally to per-LLM tuning with few-shot prompting. We evaluate five LLMs (DeepSeek, GPT-4, GPT-5, LLaMA 3, LLaMA 4) on three manually curated OpenVAS baselines across 450 independent runs. Without changing the underlying models, Exact Record Match rises from 37.5% to 90.4%, vulnerability-level omission falls by an order of magnitude (20.5% → 1.7%), and cross-model variance in field-level omission collapses from 41.8 to 2.9 percentage points, all with non-overlapping 95% bootstrap confidence intervals. These results establish that pipeline engineering, not model substitution, is the primary lever for extraction quality: within the scope evaluated here, a well-engineered pipeline turns LLM choice into a question of cost and latency rather than accuracy.
Referências
Croft, R., Babar, M. A., and Kholoosi, M. M. (2023). Data quality for software vulnerability datasets. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), pages 121–133. IEEE.
Dabholkar, C. S. (2023). Integrating open-source vulnerability scanning tools reports with openai api for automated report generation. Master’s thesis, National College of Ireland, Dublin, Ireland.
Dong, Y., Chattopadhyay, S., Aung, Y. L., and Zhou, J. (2025). Chatiot: Large language model-based security assistant for internet of things with retrieval-augmented generation. arXiv preprint arXiv:2502.09896.
Dunn, A., Dagdelen, J., Walker, N., Lee, S., Rosen, A. S., Ceder, G., Persson, K. A., and Jain, A. (2024). Structured information extraction from scientific text with large language models. Nature Communications, 15:1418.
Ferrag, M. A., Al-Hawawreh, M., Battah, A., et al. (2025). Generative AI in cybersecurity: A comprehensive review of LLM applications and vulnerabilities. arXiv preprint arXiv:2405.12750.
Foreman, P. (2019). Vulnerability Management. Auerbach Publications.
Gao, P., Liu, X., Choi, E., Ma, S., Yang, X., and Song, D. (2022). Threatkg: An ai-powered system for automated open-source cyber threat intelligence gathering and management. arXiv preprint arXiv:2212.10388.
Garcia, A. M., Beunza, J.-J., et al. (2024). Enhanced medical data extraction: Leveraging llms for accurate retrieval of patient information from medical reports. Preprints.org.
Ghane, S., Sawant, R., et al. (2024). Langchainiq: Intelligent content and query processing. International Journal of Management, Technology, and Social Sciences (IJMTS), 9(3):34–43.
Guo, Y., Bettaieb, S., and Casino, F. (2024). A comprehensive analysis on software vulnerability detection datasets: trends, challenges, and road ahead. International Journal of Information Security, 23(5):3311–3327.
Hu, Y., Sun, D., and Liu, P. (2024). Adapting generative large language models for information extraction. Journal of Data Intelligence.
Klusty, M. A., Solie, E. C., Leach, C. N., and Logan, W. V. (2024). Leveraging LLMs for structured data extraction from unstructured patient notes. arXiv preprint arXiv:2512.13700.
Lin, D. (2024). Revolutionizing retrieval-augmented generation with enhanced pdf structure recognition. chatdoc.com Technical Report.
Machado, B., Lautert, D., Kapelinski, C., and Kreutz, D. (2025). Structured extraction of vulnerabilities in OpenVAS and Tenable WAS reports using LLMs. In Anais da XXII Escola Regional de Redes de Computadores (ERRC), pages 144–150, Porto Alegre. Sociedade Brasileira de Computação.
Mitchell, E., Are, E. B., Colijn, C., and Earn, D. J. D. (2025). Using artificial intelligence tools to automate data extraction for living evidence syntheses. PLoS One, 20(4):e0320151.
NIST (2026). NIST updates NVD operations to address record CVE growth. [link].
Park, Y. and Lee, T. (2022). Full-stack information extraction system for cybersecurity intelligence. In Proceedings of EMNLP 2022 Industry Track, pages 531–539.
Ponce, L. M., Ribeiro, I., Oliveira, E., Ítalo Cunha, Hoepers, C., Steding-Jessen, K., Chaves, M. H. P. C., Guedes, D., and Jr., W. M. (2024). Identificação de serviços e dispositivos em dados de motores de busca para o enriquecimento de análise de vulnerabilidades. In Anais do XXIV SBSeg, pages 367–382. SBC.
Rafati, R. (2025). Why relying solely on automated vulnerability scanners is risky for CVE detection. Online. Discusses limitations of automated scanners and the need for manual validation and expert review.
Sharmila, S. P. (2026). PDFInspect: A unified feature extraction framework for malicious document detection.
The Sequence (2025). The vulnerability data crisis: Why you can’t trust your security tools (and what to do about it). [link]. Accessed: 2025.
Wang, B. et al. (2024). Docllm: A lightweight extension for visually rich document understanding. Proceedings of the 62nd ACL 2024.
Zhang, X., Luo, S., et al. (2024). Tablellm: Enabling tabular data manipulation by llms in real office usage scenarios. arXiv preprint arXiv:2403.19318v3.
