MulitaMiner: An LLM-Based Tool for Structuring Vulnerability Scanner Report

Resumo


Vulnerability scanners produce heterogeneous, vendor-specific PDF reports that hinder automated analysis and vulnerability management. We present MulitaMiner, an open-source tool that uses LLMs to extract structured records from these reports into a canonical 18-field schema via a segment-prompt-validate pipeline. Evaluated on 129 OpenVAS reports (6,343 findings) across three pipeline versions and five LLMs, MulitaMiner raises aggregate correctness from 67.5% to 78.3% and cuts omission from 14.1% to 2.8%; DeepSeek offers the best balance, tying for the lowest omission (2.7%) and posting the lowest hallucination (4.1%). MulitaMiner turns archived scanner PDFs into queryable data for CSIRT triage, prioritization, and tracking.

Referências

Abdullah, B. F. et al. (2025). Using llms for security advisory investigations: How far are we? arXiv:2506.13161.

Aghaei, E. et al. (2025). Securebert 2.0: Advanced language model for cybersecurity intelligence. arXiv:2510.00240.

Alharbi, H., Hur, A., Alkahtani, H., et al. (2025). Enhancing cybersecurity through autonomous knowledge graph construction by integrating heterogeneous data sources. PeerJ Computer Science, 11.

Chen, T., Li, L., Zhu, L., et al. (2024). Vullibgen: Generating names of vulnerability-affected packages via a large language model. In ACL, pages 9767–9780.

Dabholkar, C. S. (2023). Integrating open-source vulnerability scanning tools reports with Open AI API for automated report generation. Master’s thesis, National College of Ireland.

Dagdelen, J., Dunn, A., Lee, S., et al. (2024). Structured information extraction from scientific text with large language models. Nature Communications, 15.

Dong, Y., Guo, W., Chen, Y., et al. (2019). Towards the detection of inconsistencies in public security vulnerability reports. In USENIX Security.

Elsharef, I., Zeng, Z., and Gu, Z. (2024). Facilitating threat modeling by leveraging large language models. In AISCC.

Fan, H. (2026). Capacity, not format: Rethinking structured reasoning failures. arXiv:2606.09410.

Ferrag, M. A., Alwahedi, F., Battah, A., et al. (2025). Generative AI in cybersecurity: A comprehensive review of LLM applications and vulnerabilities. Internet of Things and Cyber-Physical Systems, 5.

Foreman, P. (2019). Vulnerability Management. Auerbach Publications. Gao, P., Liu, X., Choi, E., et al. (2024). Threatkg: An ai-powered system for automated open-source cyber threat intelligence gathering and management. In 1st ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis (LAMPS).

Hashemi Chaleshtori, F. and Ray, I. (2023). Automation of vulnerability information extraction using transformer-based language models. In ESORICS Workshops.

Jiang, Y., Shang, F., You, F. T. W., et al. (2025). Vulcpe: Context-aware cybersecurity vulnerability retrieval and management. arXiv:2505.13895.

Kumar, V. (2025). Automating security advisory evaluation through large language models. Master’s thesis, Università degli Studi di Padova.

Liu, R., Xie, Y., Dang, Z., et al. (2025). Dynamic vulnerability knowledge graph construction via multi-source data fusion and large language model reasoning. Electronics, 14(12).

Machado, B., Lautert, D., Kapelinski, C., and Kreutz, D. (2025). Structured extraction of vulnerabilities in openvas and tenable was reports using llms. In Anais da XXII Escola Regional de Redes de Computadores (ERRC), pages 144–150. Sociedade Brasileira de Computação (SBC).

Mitchell, E., Are, E. B., Colijn, C., et al. (2025). Using artificial intelligence tools to automate data extraction for living evidence syntheses. PLoS One, 20(4).

Shimizu, N. and Hashimoto, M. (2026). Vulnerability management chaining: An integrated framework for efficient cybersecurity risk prioritization. IEEE Access, 14:31407–31424.

Tang, L. et al. (2025). Polar: Automating cyber threat prioritization through llm-powered assessment. arXiv:2510.01552.

Vijayan, A. (2023). A prompt engineering approach for structured data extraction from unstructured text using conversational llms. In ACAI.

Vithanage, D., Deng, C., Wang, L., et al. (2025). Adapting generative large language models for information extraction from unstructured electronic health records in residential aged care: A comparative analysis of training approaches. Journal of Healthcare Informatics Research, 9(2).

Wang, B., Xu, C., et al. (2024a). Mineru: An open-source solution for precise document content extraction. arXiv:2409.18839.

Wang, D., Raman, N., Sibue, M., et al. (2024b). DocLLM: A layout-aware generative language model for multimodal document understanding. In ACL.

Wang, L., Sun, J., Jiang, J., et al. (2025). Vulnerability aspects extraction and discrepancies detection across heterogeneous threat intelligence. In LM-SHIELD @ ACM AsiaCCS.

West-Brown, M. J., Stikvoort, D., Kossakowski, K.-P., et al. (2003). Handbook for computer security incident response teams (CSIRTs). Technical Report CMU/SEI-2003-HB-002, Software Engineering Institute, Carnegie Mellon University.

Yan, Z., Ye, Z., Ge, J., et al. (2025). DocExtractNet: A novel framework for enhanced information extraction from business documents. Information Processing & Management, 62(3).
Publicado
01/09/2026
LAUTERT, Douglas; MACHADO, Beatriz; KAPELINSKI, Cristhian; KREUTZ, Diego. MulitaMiner: An LLM-Based Tool for Structuring Vulnerability Scanner Report. In: SALÃO DE FERRAMENTAS - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 224-233. DOI: https://doi.org/10.5753/sbseg_estendido.2026.33714.

Artigos mais lidos do(s) mesmo(s) autor(es)

<< < 4 5 6 7 8 9 10 11 12 > >>