PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents
Resumo
Connecting large language models (LLMs) to defensive enforcement requires more than asking a model whether an attack is happening. A defender must decide which model outputs may change the system state, which outputs must be rejected, and how failures should be recorded. We present PocketAgents, a manifest-driven library of autonomous defense agents. Each agent is installed as three data files: a manifest, a prompt, and a runtime context. The shared runtime gives the agent bounded telemetry access and accepts only typed reports whose requested action appears in the manifest. We implemented PocketAgents on top of Perry, a cyber-deception testbed, and evaluated two agents for the Command and Control and Exfiltration tactics in 18 closed-loop trials of a DarkSide-inspired attack on a small enterprise topology. Thirteen trials produced validated network-block actions and contained the attack; four failed schema validation; one produced a valid no-action decision. The experiments show that a typed boundary makes LLM-driven defense measurable, extensible, and attributable.
Referências
Ayzenshteyn, D., Weiss, R., and Mirsky, Y. (2025). Cloak, honey, trap: Proactive defenses against LLM agents. In Proceedings of the 34th USENIX Security Symposium (USENIX Security), pages 8095–8114. USENIX Association.
Casado, M., Garfinkel, T., Akella, A., Freedman, M. J., Boneh, D., McKeown, N., and Shenker, S. (2006). SANE: A protection architecture for enterprise networks. In Proceedings of the 15th USENIX Security Symposium.
CISA and FBI (2021). DarkSide ransomware: Best practices for preventing business disruption from ransomware attacks. Technical Report AA21-131A, U.S. Department of Homeland Security and U.S. Department of Justice.
Deng, G., Liu, Y., Mayoral-Vilches, V., Liu, P., Li, Y., Xu, Y., Zhang, T., Liu, Y., Pinzger, M., and Rass, S. (2024). PentestGPT: Evaluating and harnessing large language models for automated penetration testing. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security).
Han, X., Pasquier, T., Bates, A., Mickens, J., and Seltzer, M. (2020). UNICORN: Runtime provenance-based detector for advanced persistent threats. In Proceedings of the 27th Annual Network and Distributed System Security Symposium (NDSS). Internet Society.
Hassan, W. U., Guo, S., Li, D., Chen, Z., Jee, K., Li, Z., and Bates, A. (2019). NoDoze: Combatting threat alert fatigue with automated provenance triage. In Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS). Internet Society.
Kim, H., Reich, J., Gupta, A., Shahbaz, M., Feamster, N., and Clark, R. (2015). Kinetic: Verifiable dynamic network control. In Proceedings of the 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI).
Kim, H., Song, M., Na, S. H., Shin, S., and Lee, K. (2025). When LLMs go online: The emerging threat of web-enabled LLMs. In Proceedings of the 34th USENIX Security Symposium (USENIX Security), pages 1729–1748. USENIX Association.
Kokulu, F. B., Soneji, A., Bao, T., Shoshitaishvili, Y., Zhao, Z., Doupé, A., and Ahn, G.-J. (2019). Matched and mismatched SOCs: A qualitative study on security operations center issues. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security (CCS), pages 1955–1970. ACM.
Schlette, D., Empl, P., Caselli, M., Schreck, T., and Pernul, G. (2024). Do you play it by the books? a study on incident response playbooks and influencing factors. In Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP), pages 3625–3643. IEEE.
Sequeira, R., Damianakis, S., Iqbal, U., and Psounis, K. (2026). Agent-Sentry: Bounding LLM agents via execution provenance. arXiv preprint arXiv:2603.22868.
Singer, B., Saquib, Y., Bauer, L., and Sekar, V. (2025). Perry: A high-level framework for accelerating cyber deception experimentation. In Proceedings of the 28th International Symposium on Research in Attacks, Intrusions and Defenses (RAID), pages 158–173.
Strom, B. E., Applebaum, A., Miller, D. P., Nickels, K. C., Pennington, A. G., and Thomas, C. B. (2020). MITRE ATT&CK: Design and philosophy. Technical Report MP180360R1, The MITRE Corporation.
U.S. House Majority Staff (2018). The equifax data breach. Technical report, U.S. House of Representatives, Committee on Oversight and Government Reform.
Vermeer, M., Kadenko, N., Gañán, C., van Eeten, M., and Parkin, S. (2023). Alert alchemy: SOC workflows and decisions in the management of NIDS rules. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security (CCS). ACM.
Wang, H., Poskitt, C. M., and Sun, J. (2026). AgentSpec: Customizable runtime enforcement for safe and reliable LLM agents. In Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering (ICSE). ACM.
Wu, Y., Roesner, F., Kohno, T., Zhang, N., and Iqbal, U. (2025). IsolateGPT: An execution isolation architecture for LLM-based agentic systems. In Proceedings of the 32nd Annual Network and Distributed System Security Symposium (NDSS). Internet Society.
Xiao, Z., Sun, J., and Chen, J. (2026). AIR: Improving agent safety through incident response. arXiv preprint arXiv:2602.11749.
Yang, L., Chen, Z., Wang, C., Zhang, Z., Booma, S., Cao, P., Adam, C., Withers, A., Kalbarczyk, Z., Iyer, R. K., and Wang, G. (2024). True attacks, attack attempts, or benign triggers? an empirical measurement of network alerts in a security operations center. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security), pages 1525–1542. USENIX Association.
Yu, T., Fayaz, S. K., Collier, M. J., Sekar, V., and Seshan, S. (2017). PSI: Precise security instrumentation for enterprise networks. In Proceedings of the 24th Annual Network and Distributed System Security Symposium (NDSS).
Zhang, Z., Cui, S., Lu, Y., Zhou, J., Yang, J., Wang, H., and Huang, M. (2025). Agent-SafetyBench: Evaluating the safety of LLM agents. arXiv preprint arXiv:2412.14470.
