Avaliação do uso de Explicabilidade de LLMs e Grafos de Ativação (NRAGs) em estudos bibliométricos
Resumo
Apesar de sua popularidade, os Grandes Modelos de Linguagem (LLMs) são modelos de ”caixa-preta”que apresentam desafios para a interpretabilidade e o alinhamento. A Explicabilidade em IA (XAI) foca em compreender esse tipo de modelo, permitindo que pesquisadores detectem padrões que auxiliem na compreensão de seu comportamento. Além de revelar informações sobre os modelos, esses padrões também podem revelar informações sobre os dados fornecidos aos algoritmos. Neste artigo, avaliamos como a Explicabilidade de LLMs poderia ser aplicada no campo da bibliometria. O objetivo é utilizar padrões de inferência de LLM para representar diferenças entre áreas de pesquisa (utilizando artigos publicados nessas áreas). Utilizamos a ferramenta LLM-MRI, que gera grafos (NRAGs) que representam padrões de ativação neural durante a inferência do LLM. Computamos métricas de redes complexas a partir dos grafos para avaliar três diferentes campos de pesquisa em Ciência da Computação.
Referências
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. (2003). A neural probabilistic language model. Journal of Machine Learning Research, 3(Feb):1137–1155.
Bricken, T. et al. (2023). Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread. [link].
Celestino, M. S., Belluzzo, R. C. B., Albino, J. P., and Valente, V. C. P. N. (2024). AnÁlise bibliomÉtrica: RevisÃo de literatura e proposta de framework metodolÓgico em 12 passos. ARACÊ, 6(4):13421–13446.
Chen, R. et al. (2025). Persona vectors: Monitoring and controlling character traits in language models.
Costa, L., Figenio, M., Santanchè, A., and Gomes-Jr, L. (2024). Llm-mri python module: a brain scanner for llms. In Companion Proceedings of the 39th Brazilian Symposium on Data Bases (SBBD), Florianópolis, SC, Brazil.
Cunningham, H. et al. (2023). Sparse autoencoders find highly interpretable features in language models. ArXiv, abs/2309.08600.
Doshi-Velez, F. and Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608.
Gomes-Jr, L., Santanchè, A., Figenio, M., and Costa, L. (2025). Explicabilidade de LLMs usando grafos de ativação de regiões neurais (NRAGs). In Proceedings of the XIX Brazilian e-Science Workshop (BreSci), pages 17–24, Fortaleza, CE, Brasil.
Horta, V. A., Tiddi, I., Little, S., and Mileo, A. (2021). Extracting knowledge from deep neural networks through graph analysis. Future Generation Computer Systems, 120:109–118.
Jurafsky, D. and Martin, J. H. (2026). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Stanford University, 3rd edition. Online manuscript released January 6, 2026.
Korbak, T. et al. (2025). Chain of thought monitorability: A new and fragile opportunity for ai safety.
Lindsey, J. et al. (2025). On the biology of a large language model. Transformer Circuits Thread.
Lundberg, S. M. and Lee, S. (2017). A unified approach to interpreting model predictions. CoRR, abs/1705.07874.
Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence, 267:1–38.
Passas, I. (2024). Bibliometric analysis: The main steps. Encyclopedia, 4:1014–1025.
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). ”why should I trust you?”: Explaining the predictions of any classifier. CoRR, abs/1602.04938.
Samek, W., Wiegand, T., and Müller, K. (2017). Explainable artificial intelligence: Understanding, visualizing and interpreting deep learning models. CoRR, abs/1708.08296.
Sharkey, L. et al. (2025). Open problems in mechanistic interpretability.
Sousa, M., ALMEIDA, E., and Dantas Bezerra, A. (2024). Bibliometria: o que é? para que serve? e como se faz? bibliometrics: what is it? what is it used for? and how to do it? bibliometría: ¿qué es? ¿para qué es? ¿y cómo lo haces? Cuadernos de Educación y Desarrollo, 16:01–35.
Vaswani, A. et al. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30.
Zhao, H. et al. (2024). Explainability for large language models: A survey. ACM Transactions on Intelligent Systems and Technology. Just Accepted.
