Robustez de LLMs na Detecção de Vulnerabilidades em Contratos Inteligentes com Transformações no Código
Resumo
Grandes Modelos de Linguagem (LLMs) têm apresentado resultados promissores na análise de segurança de código, mas sua avaliação pode ser influenciada por pistas textuais não executáveis. Este trabalho avalia a robustez de LLMs na detecção de vulnerabilidades em contratos inteligentes sob transformações de código que preservam a semântica. Na comparação de quatro modelos de LLM, sobre contratos originais e em variantes sem comentários, com identificadores renomeados e com ambas as transformações combinadas, constataram-se reduções de F1 de até 99,5%, mostrando que diferentes representações do mesmo código, mas semanticamente equivalentes, podem produzir diagnósticos substancialmente distintos.Referências
Achjian, R. and Junior, M. S. (2025). Building a labeled smart contract dataset for evaluating vulnerability detection tools’ effectiveness. In Anais Estendidos do XXV Simpósio Brasileiro de Cibersegurança, pages 1–10, Porto Alegre, RS, Brasil. SBC.
Boi, B., Esposito, C., and Lee, S. (2024). Vulnhunt-gpt: a smart contract vulnerabilities detector based on openai chatgpt. In Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing, pages 1517–1524.
Chen, C., Su, J., Chen, J., Wang, Y., Bi, T., Yu, J., Wang, Y., Lin, X., Chen, T., and Zheng, Z. (2025). When chatgpt meets smart contract vulnerability detection: How far are we? ACM Transactions on Software Engineering and Methodology, 34(4):1–30.
da Costa, E. V. A., Guzman, C. A. M., de Abreu, O. B. M., Sousa, J. C., and Braga, A. (2026). Análise comparativa do desempenho de llms e ferramentas de segurança na detecção de vulnerabilidades em contratos inteligentes. In Workshop em Blockchain: Teoria, Tecnologias e Aplicações (WBlockchain), pages 1–14. SBC.
Durieux, T., Ferreira, J. F., Abreu, R., and Cruz, P. (2020). Empirical review of automated analysis tools on 47,587 Ethereum smart contracts. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering, pages 530–541.
Fu, M. and Tantithamthavorn, C. (2022). Linevul: A transformer-based line-level vulnerability prediction. In Proceedings of the 19th international conference on mining software repositories, pages 608–620.
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. (2020). Shortcut learning in deep neural networks. Nature Machine Intelligence, 2:665–673.
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., and Smith, N. A. (2018). Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 107–112.
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., and Wang, H. (2024). Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology, 33(8):1–79.
Mandana, E., Vlahavas, G., and Vakali, A. (2025). Evullm: ethereum smart contract vulnerability detection using large language models. Electronics, 14(16):3226.
Manning, C. D. (2008). Introduction to information retrieval. Syngress Publishing,.
Moreno-Torres, J. G., Raeder, T., Alaiz-Rodríguez, R., Chawla, N. V., and Herrera, F. (2012). A unifying view on dataset shift in classification. Pattern recognition, 45(1):521–530.
Rabin, M. R. I., Bui, N. D. Q., Wang, K., Yu, Y., Jiang, L., and Alipour, M. A. (2021). On the generalizability of neural program models with respect to semantic-preserving program transformations. Information and Software Technology, 135:106552.
Sainz, O., Campos, J., García-Ferrero, I., Etxaniz, J., de Lacalle, O. L., and Agirre, E. (2023). Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10776–10787.
Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In Proceedings of the 12th International Conference on Learning Representations.
Sun, Y., Wu, D., Xue, Y., Liu, H., Wang, H., Xu, Z., Xie, X., and Liu, Y. (2024). Gptscan: Detecting logic vulnerabilities in smart contracts by combining gpt with program analysis. In Proceedings of the IEEE/ACM 46th international conference on software engineering, pages 1–13.
Yefet, N., Alon, U., and Yahav, E. (2020). Adversarial examples for models of code. Proceedings of the ACM on Programming Languages, 4(OOPSLA):1–30.
Zhou, X., Cao, S., Sun, X., and Lo, D. (2025). Large language model for vulnerability detection and repair: Literature review and the road ahead. ACM Transactions on Software Engineering and Methodology, 34(5):1–31.
Boi, B., Esposito, C., and Lee, S. (2024). Vulnhunt-gpt: a smart contract vulnerabilities detector based on openai chatgpt. In Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing, pages 1517–1524.
Chen, C., Su, J., Chen, J., Wang, Y., Bi, T., Yu, J., Wang, Y., Lin, X., Chen, T., and Zheng, Z. (2025). When chatgpt meets smart contract vulnerability detection: How far are we? ACM Transactions on Software Engineering and Methodology, 34(4):1–30.
da Costa, E. V. A., Guzman, C. A. M., de Abreu, O. B. M., Sousa, J. C., and Braga, A. (2026). Análise comparativa do desempenho de llms e ferramentas de segurança na detecção de vulnerabilidades em contratos inteligentes. In Workshop em Blockchain: Teoria, Tecnologias e Aplicações (WBlockchain), pages 1–14. SBC.
Durieux, T., Ferreira, J. F., Abreu, R., and Cruz, P. (2020). Empirical review of automated analysis tools on 47,587 Ethereum smart contracts. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering, pages 530–541.
Fu, M. and Tantithamthavorn, C. (2022). Linevul: A transformer-based line-level vulnerability prediction. In Proceedings of the 19th international conference on mining software repositories, pages 608–620.
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. (2020). Shortcut learning in deep neural networks. Nature Machine Intelligence, 2:665–673.
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., and Smith, N. A. (2018). Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 107–112.
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., and Wang, H. (2024). Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology, 33(8):1–79.
Mandana, E., Vlahavas, G., and Vakali, A. (2025). Evullm: ethereum smart contract vulnerability detection using large language models. Electronics, 14(16):3226.
Manning, C. D. (2008). Introduction to information retrieval. Syngress Publishing,.
Moreno-Torres, J. G., Raeder, T., Alaiz-Rodríguez, R., Chawla, N. V., and Herrera, F. (2012). A unifying view on dataset shift in classification. Pattern recognition, 45(1):521–530.
Rabin, M. R. I., Bui, N. D. Q., Wang, K., Yu, Y., Jiang, L., and Alipour, M. A. (2021). On the generalizability of neural program models with respect to semantic-preserving program transformations. Information and Software Technology, 135:106552.
Sainz, O., Campos, J., García-Ferrero, I., Etxaniz, J., de Lacalle, O. L., and Agirre, E. (2023). Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10776–10787.
Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In Proceedings of the 12th International Conference on Learning Representations.
Sun, Y., Wu, D., Xue, Y., Liu, H., Wang, H., Xu, Z., Xie, X., and Liu, Y. (2024). Gptscan: Detecting logic vulnerabilities in smart contracts by combining gpt with program analysis. In Proceedings of the IEEE/ACM 46th international conference on software engineering, pages 1–13.
Yefet, N., Alon, U., and Yahav, E. (2020). Adversarial examples for models of code. Proceedings of the ACM on Programming Languages, 4(OOPSLA):1–30.
Zhou, X., Cao, S., Sun, X., and Lo, D. (2025). Large language model for vulnerability detection and repair: Literature review and the road ahead. ACM Transactions on Software Engineering and Methodology, 34(5):1–31.
Publicado
01/09/2026
Como Citar
ALVES, Rafael Santa Rosa; HENRIQUES, Marco Amaral.
Robustez de LLMs na Detecção de Vulnerabilidades em Contratos Inteligentes com Transformações no Código. In: WORKSHOP DE CIBERSEGURANÇA EM IA - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 1065-1072.
DOI: https://doi.org/10.5753/sbseg_estendido.2026.33836.
