Beyond Guardrails: Tracing Cross-Layer Evidence in Local LLM-Agent Execution
Resumo
Agent-facing restrictions may reduce an LLM agent’s exposed tools without making corresponding tested effects unrealizable elsewhere in the runtime environment. We evaluated three NemoClaw/OpenShell profiles through mediated trials, direct profile-level probes, and independent effect oracles. The interface-restricted profile repeatedly acquired a protected marker despite its reduced tool surface, whereas the hardened profile prevented acquisition across all induced read attempts. The permissive profile reached the controlled destination and produced one confirmed exfiltration in a post-hoc replay. These results distinguish interface restriction, profile-level realizability, destination contact, and realized effects.
Referências
Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F. (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. In Advances in Neural Information Processing Systems, volume 37. Datasets and Benchmarks Track.
Hardy, N. (1988). The Confused Deputy: (or Why Capabilities Might Have Been Invented). ACM SIGOPS Operating Systems Review, 22(4):36–38.
Hou, X., Zhao, Y., Wang, S., and Wang, H. (2025). Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions.
Jia, F., Wu, T., Qin, X., and Squicciarini, A. (2024). The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents.
Jin, S., Guo, R., and Cheung, R. C. C. (2026). CapSeal: Capability-Sealed Secret Mediation for Secure Agent Execution.
Kravchenko, A., Liventsev, V., Konstantinov, I., Iskhakov, I., and Kukuy, M. (2026). Agentic permissions policy algebra for taint confinement in llm agents. arXiv preprint arXiv:2607.24625.
Lin, X., Lei, L., Wang, Y., Jing, J., Sun, K., and Zhou, Q. (2018). A Measurement Study on Linux Container Security: Attacks and Countermeasures. In Proceedings of the 34th Annual Computer Security Applications Conference, pages 418–429. Association for Computing Machinery.
NVIDIA (2026). NemoClaw Ecosystem: OpenClaw, OpenShell, and NemoClaw. [link]. Accessed: 2026-07-28.
Saltzer, J. H. and Schroeder, M. D. (1975). The Protection of Information in Computer Systems. Proceedings of the IEEE, 63(9):1278–1308.
Shi, T., He, J., Wang, Z., Wu, L., Li, H., Guo, W., and Song, D. (2025). Progent: Programmable Privilege Control for LLM Agents.
Souppaya, M., Morello, J., and Scarfone, K. (2017). Application Container Security Guide. Technical Report NIST Special Publication 800-190, National Institute of Standards and Technology.
Wu, B., Liu, Q., Bibi, A., King, I., and Lyu, S. (2026). The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents.
Yan, Z., Weng, J., Chen, C., Peng, D., Qin, E., Guan, J., Liu, J., Yu, Q., Yuan, Y., Meng, F., Che, C., and Hu, M. (2026). Do Coding Agents Understand Least-Privilege Authorization? Yang, Y., Gao, C., Wu, D., Chen, Y., Li, Y., and Wang, S. (2025). MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols.
Zhan, Q., Liang, Z., Ying, Z., and Kang, D. (2024). InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. In Findings of the Association for Computational Linguistics: ACL 2024, pages 10471–10506. Association for Computational Linguistics.
Zhao, W., Li, Z., Zhang, P., and Sun, J. (2026). ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection.
