Uma Análise Sistemática da Estrutura Mecanística de Backdoors em LLMs

  • Gabriel Sousa Canto UFAM
  • Eduardo Luzeiro Feitosa UFAM

Resumo


Neste trabalho, realizamos uma análise mecanística sistemática de backdoors em LLMs utilizando técnicas de interpretabilidade, incluindo attribution patching e Sparse Autoencoders (SAEs). Avaliamos quatro modelos Llama-2-7B com gatilhos lexicais, contextuais e temporais. Os resultados mostram que, em três dos quatro casos, os mecanismos de backdoor concentram-se nas camadas finais do modelo e podem ser significativamente suprimidos pela ablação de poucas features SAE, reduzindo o ASR em até 91%. Também observamos que backdoors temporais apresentam comportamento mais distribuído e difícil de localizar. Os resultados indicam que backdoors podem deixar assinaturas internas interpretáveis, contribuindo para o desenvolvimento de métodos mais robustos e interpretáveis de detecção e mitigação em LLMs.

Referências

Abu Baker, M. and Babu-Saheer, L. (2025). Mechanistic exploration of backdoored large language model attention patterns. arXiv:2508.15847.

Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., and Nanda, N. (2024). Refusal in language models is mediated by a single direction. In Advances in Neural Information Processing Systems.

Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N. L., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C. (2023). Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits.

Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C. (2022). Toy models of superposition. Transformer Circuits.

Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. (2021). A mathematical framework for transformer circuits. Transformer Circuits.

Gu, T., Dolan-Gavitt, B., and Garg, S. (2017). BadNets: Identifying vulnerabilities in the machine learning model supply chain. In arXiv:1708.06733.

Hubinger, E., Denison, et al. (2024). Sleeper agents: Training deceptive LLMs that persist through safety training. arXiv:2401.05566.

Liu, Y., Ma, S., Aafer, Y., Lee, W.-C., Zhai, J., Wang, W., and Zhang, X. (2018). Trojaning attack on neural networks. In Proceedings of the 25th NDSS.

MacDiarmid, M., Maxwell, T., Schiefer, N., Mu, J., Kaplan, J., Duvenaud, D., Bowman, S., Tamkin, A., Perez, E., Sharma, M., Denison, C., and Hubinger, E. (2024). Simple probes can catch sleeper agents.

Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A. (2024). Sparse feature circuits: Discovering and editing interpretable causal graphs in language models. arXiv:2403.19647.

Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2022). Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems.

Min, N. M., Pham, L. H., Li, Y., and Sun, J. (2025). CROW: Eliminating backdoors from large language models via internal consistency regularization. Proceedings of the 42nd ICML.

Olsson, C., Elhage, et al. (2022). In-context learning and induction heads. Transformer Circuits.

OWASP (2025). OWASP top 10 for large language model applications, version 2025.

Ponkshe, K., Singhal, R., Gorbett, M., Kanwar, A., Brandfonbrener, D., Belinkov, Y., and Saxe, A. (2025). Safety subspaces are not linearly distinct: A fine-tuning case study.

Price, S., Panickssery, A., Bowman, S., and Stickland, A. C. (2024). Future events as backdoor triggers: Investigating temporal vulnerabilities in LLMs. arXiv:2407.04108.

Rando, J. and Tramèr, F. (2024). Universal jailbreak backdoors from poisoned human feedback. arXiv:2311.14455.

Syed, A., Rager, C., and Conmy, A. (2023). Attribution patching outperforms automated circuit discovery. arXiv:2310.10348.

Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T. (2024). Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet. Transformer Circuits.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems.

Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. (2023). Interpretability in the wild: A circuit for indirect object identification in GPT-2 small. arXiv:2211.00593.

Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, J. Z., and Hendrycks, D. (2023). Representation engineering: A top-down approach to AI transparency. arXiv:2310.01405.
Publicado
01/09/2026
CANTO, Gabriel Sousa; FEITOSA, Eduardo Luzeiro. Uma Análise Sistemática da Estrutura Mecanística de Backdoors em LLMs. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 1260-1275. DOI: https://doi.org/10.5753/sbseg.2026.27042.

Artigos mais lidos do(s) mesmo(s) autor(es)