A Comparative Study Between Qwen 3.6 Flash and Small Language Models for Email Prompt Injection in Android Edge Agents
Resumo
AI agents deployed on mobile and embedded devices are increasingly able to access private data and invoke external tools, making indirect prompt injection a critical security risk. In this paper, we evaluate e-mail-based prompt-injection attacks against an Android tool-using agent with access to mocked email, calendar, TV/device-control, and rule-management tools. The benchmark contains 50 adversarial runs per model, organized into 10 attack scenarios and five benign user-request phrasings. We compare Qwen 3.6 Flash as a cloud-based reference model with six small language models executed locally on a Samsung Galaxy S24 using MNN engine (Qwen3 0.6B, Qwen2.5 0.5B, Gemma 3 1B, Llama 3.2 1B, SmolLM2 135M, and SmolLM2 360M). The results show that Qwen 3.6 Flash was the most agentically capable model, triggering tools in 96.0% of the runs and detecting phishing in 52.0%, but still executing sensitive actions in 26.0%. Among local models, Qwen3 0.6B reached the highest sensitive-action rate, 32.0%, despite running fully on-device. Llama 3.2 1B reduced unsafe actions mainly through broad refusal, whereas the SmolLM2 models rarely emitted tools because they frequently failed to produce coherent agentic behavior. These findings indicate that low tool-trigger rates are not sufficient evidence of prompt-injection robustness and that secure Android edge agents require runtime guardrails, tool-level permission checks, and explicit separation between trusted user instructions and untrusted retrieved content.
Referências
T. Schick, J. Dwivedi-Yu, R. Dessı̀, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” in Advances in Neural Information Processing Systems (NeurIPS), 2023.
E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr, “AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents,” in Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2024.
Y. Liu, G. Deng, Y. Li, K. Wang, Z. Wang, X. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng, and Y. Liu, “Prompt injection attack against LLM-integrated applications,” arXiv preprint arXiv:2306.05499, 2023.
Y. Liu, Y. Jia, R. Geng, J. Jia, and N. Z. Gong, “Formalizing and benchmarking prompt injection attacks and defenses,” in Proc. 33rd USENIX Security Symposium, pp. 1831–1847, 2024.
Jia, Feiran, et al. ”The task shield: Enforcing task alignment to defend against indirect prompt injection in llm agents.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025.
Suo, Xuchen. ”Signed-prompt: A new approach to prevent prompt injection attacks against llm-integrated applications.” AIP Conference Proceedings. Vol. 3194. No. 1. AIP Publishing LLC, 2024.
JIANG, Xiaotang et al. ”MNN: A universal and efficient inference engine.” Proceedings of Machine Learning and Systems, v. 2, p. 1-13, 2020.
M. Zhang, J. Cao, X. Shen, and Z. Cui, “EdgeShard: Efficient LLM inference via collaborative edge computing,” arXiv preprint arXiv:2405.14371, 2024.
Y. Yan, J. Shen, X. Luo, and Y. Zhou, “EdgeFlow: Fast cold starts for LLMs on mobile devices,” arXiv preprint arXiv:2604.09083, 2026.
YANG, An et al. ”Qwen3 technical report.” arXiv preprint arXiv:2505.09388, 2025.
YANG, An et al. ”Qwen2. 5-math technical report: Toward mathematical expert model via self-improvement.” arXiv preprint arXiv:2409.12122, 2024. Gemma Team, “Gemma 3 technical report,” arXiv preprint arXiv:2503.19786, 2025.
GRATTAFIORI, Aaron et al. ”The llama 3 herd of models.” arXiv preprint arXiv:2407.21783, 2024. L. B. Allal, A. Lozhkov, E. Bakouch, G. Martín Blázquez, G. Penedo, L. Tunstall, A.
Marafioti, H. Kydlíček, A. Piqueres Lajarín, V. Srivastav, J. Lochner, C. Fahlgren, X.-S. Nguyen, C. Fourrier, B. Burtenshaw, H. Larcher, H. Zhao, C. Zakka, M. Morlon, C. Raffel, L. von Werra, and T. Wolf, “SmolLM2: When smol goes big—data-centric training of a small language model,” arXiv preprint arXiv:2502.02737, 2025.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS), 2017.
T. B. Brown et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
