A Multi-Agent LLM Architecture with Dynamic Cascading for On-Premises, LGPD-Compliant Organizational Assistants
Resumo
A integração de modelos de linguagem de grande porte (LLMs) em ambientes organizacionais que tratam dados sensíveis enfrenta duas barreiras: restrições regulatórias (no Brasil, a LGPD) e a heterogeneidade das cargas de trabalho. Este trabalho apresenta uma arquitetura multi-agente de LLMs que combina execução totalmente local com um roteador dinâmico em cascata: um roteador heurístico, quatro agentes LLM especializados (Tier 1 a Tier 3), um agente de RAG isolado por workspace, um de auditoria e um de desempenho. O roteador despacha cada consulta ao agente mais barato capaz de respondê-la, com fallback transparente em caso de falha, e o sistema foi publicado como código aberto.
Referências
Brasil (2018). Lei n. 13.709, de 14 de agosto de 2018: Lei geral de proteção de dados pessoais (lgpd). [link].
Carlini, N., Ippolito, D., Jagielski, M., et al. (2021). Extracting training data from large language models. In Proceedings of the 30th USENIX Security Symposium, pages 2633–2650. USENIX Association.
Chen, L., Zaharia, M., and Zou, J. (2023). Frugalgpt: How to use large language models while reducing cost and improving performance. IEEE Transactions on Knowledge and Data Engineering, 36(7):3299–3313.
da Califórnia, E. (2018). California consumer privacy act (ccpa). [link].
da União Europeia, C. (2016). Regulamento (ue) 2016/679 relativo à proteção das pessoas singulares (gdpr). [link].
Ding, D., Mallick, A., Wang, C., et al. (2024). Hybrid llm: Cost-efficient and quality-aware query routing. In Proceedings of the 12th International Conference on Learning Representations. OpenReview.
Gao, Y., Xiong, Y., Gao, X., et al. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
Greshake, K., Abdelnabi, S., Mishra, S., et al. (2023). Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, pages 79–90. ACM.
LangChain (2024). Langchain documentation. [link].
Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, pages 9459–9474. Curran Associates.
Ollama (2024). Get up and running with large language models locally. [link].
Park, J. S., O’Brien, J. C., Cai, C. J., et al. (2023). Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pages 1–22. ACM.
Rao, A. S. and Georgeff, M. P. (1995). Bdi agents: From theory to practice. In Proceedings of the 1st International Conference on Multi-Agent Systems, pages 312–319. AAAI Press.
Rosa, R. d. O. (2026). Cerra.ai: Source code repository. [link].
Schick, T., Dwivedi-Yu, J., Dessı̀, R., et al. (2023). Toolformer: Language models can teach themselves to use tools. In Proceedings of the 37th Conference on Neural Information Processing Systems. Curran Associates.
Schuster, T., Fisch, A., Gupta, J., et al. (2021). Confident adaptive language modeling. In Advances in Neural Information Processing Systems, pages 17456–17472. Curran Associates.
Wei, J., Wang, X., Schuurmans, D., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
Wooldridge, M. (2009). An Introduction to Multiagent Systems. John Wiley & Sons, 2nd edition.
Yao, S., Zhao, J., Yu, D., et al. (2023). React: Synergizing reasoning and acting in language models. In Proceedings of the 11th International Conference on Learning Representations. OpenReview.
