In-Context Teacher–Student Guidance for Open-Weight Browser Agents
Resumo
Browser agents based on large language models can automate web tasks, but deployment remains constrained by repeated model calls, monetary cost, and privacy risks. In practical settings, agents often process page states that may expose credentials, personal data, internal records, business documents, or authenticated session information. Sending these states and interaction histories to external model APIs can increase the risk of data leakage, unauthorized retention, or unintended disclosure. This paper investigates whether reasoning traces from a stronger teacher model can improve open-weight browser agents without fine-tuning. We propose In-Context Learning Guidance (ICLG), a training-free teacher–student protocol in which the teacher solves representative browser tasks offline and its reasoning traces are converted into reusable prompt-level guidance for similar tasks. The student then uses this guidance during execution while reading the concrete target values from the current browser instance. We evaluate ICLG on MiniWoB++, using GPT-5.1 as the teacher and ten open-weight student variants. ICLG increases solved student-task pairs from 208 to 271 and improves 7 out of 10 students. Ablation results suggest that these gains are not due only to retrying or generic prompt refinement. Overall, the results show that reusable teacher guidance can improve lower-cost browser agents and reduce dependence on proprietary teacher APIs during deployment.
Referências
de Chezelles, T. L. S., Gasse, M., Lacoste, A., Caccia, M., Drouin, A., Boisvert, L., Thakkar, M., Marty, T., Assouel, R., Shayegan, S. O., Jang, L. K., Lù, X. H., Yoran, O., Kong, D., Xu, F. F., Reddy, S., Neubig, G., Cappart, Q., Salakhutdinov, R., and Chapados, N. The browsergym ecosystem for web agent research. Transactions on Machine Learning Research, 2025.
Dekoninck, J., Baader, M., and Vechev, M. A unified approach to routing and cascading for llms. In Proceedings of the 42nd International Conference on Machine Learning, 2025.
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y. Mind2web: Towards a generalist agent for the web. In Advances in Neural Information Processing Systems. Vol. 36, 2023.
Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Liu, T., Chang, B., Sun, X., and Sui, Z. A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1107–1128, 2024.
Drouin, A., Gasse, M., Caccia, M., Laradji, I. H., Del Verme, M., Marty, T., Vazquez, D., Chapados, N., and Lacoste, A. Workarena: How capable are web agents at solving common knowledge work tasks? In Proceedings of the 41st International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 235. PMLR, pp. 11642–11662, 2024.
Gu, Y., Zhang, K., Ning, Y., Zheng, B., Gou, B., Xue, T., Chang, C., Srivastava, S., Xie, Y., Qi, P., Sun, H., and Su, Y. Is your llm secretly a world model of the internet? model-based planning for web agents. Transactions on Machine Learning Research, 2025.
He, F., Zhu, T., Ye, D., Liu, B., Zhou, W., and Yu, P. S. The emerged security and privacy of llm agent: A survey with case studies. ACM Computing Surveys 58 (6): 1–36, 2025.
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D. Webvoyager: Building an end-to-end web agent with large multimodal models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L.-W. Ku, A. Martins, and V. Srikumar (Eds.). Association for Computational Linguistics, pp. 6864–6890, 2024.
Hinton, G., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network, 2015.
Kang, M., Jeong, J., Lee, S., Cho, J., and Hwang, S. J. Distilling llm agent into small models with retrieval and code tools, 2025.
Lai, H., Liu, X., Iong, I. L., Yao, S., Chen, Y., Shen, P., Yu, H., Zhang, H., Zhang, X., Dong, Y., and Tang, J. Autowebglm: A large language model-based web navigating agent, 2024.
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P. Reinforcement learning on web interfaces using workflow-guided exploration. In International Conference on Learning Representations, 2018.
Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J. Agentbench: Evaluating llms as agents. In International Conference on Learning Representations, 2024.
Ma, C., Zhang, J., Zhu, Z., Yang, C., Yang, Y., Jin, Y., Lan, Z., Kong, L., and He, J. Agentboard: An analytical evaluation board of multi-turn llm agents. In Advances in Neural Information Processing Systems. Vol. 37, 2024.
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P. Self-refine: Iterative refinement with self-feedback. In Advances in Neural Information Processing Systems. Vol. 36, 2023.
Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., and Severyn, A. Teaching small language models to reason. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, pp. 1773–1781, 2023.
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I. Routellm: Learning to route llms with preference data. In International Conference on Learning Representations, 2025.
Qi, Z., Liu, X., Iong, I. L., Lai, H., Sun, X., Zhao, W., Yang, Y., Yang, X., Sun, J., Yao, S., Zhang, T., Xu, W., Tang, J., and Dong, Y. Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2025.
Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems. Vol. 36. pp. 68539–68551, 2023.
Shen, J., Jain, A., Xiao, Z., Amlekar, I., Hadji, M., Podolny, A., and Talwalkar, A. Scribeagent: Towards specialized web agents using production-scale workflow data, 2024.
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., and Yao, S. Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems. Vol. 36, 2023.
Shridhar, K., Stolfo, A., and Sachan, M. Distilling reasoning capabilities into smaller language models. In Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics, pp. 7059–7073, 2023.
Sun, H., Zhuang, Y., Kong, L., Dai, B., and Zhang, C. Adaplanner: Adaptive planning from feedback with language models, 2023.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems. Vol. 35. pp. 24824–24837, 2022.
Xiao, Y.-A., Gao, P., Peng, C., and Xiong, Y. Reducing cost of llm agents with trajectory reduction. arXiv e-prints, 2025.
Yan, B., Li, K., Xu, M., Dong, Y., Zhang, Y., Ren, Z., and Cheng, X. On protecting the data privacy of large language models (llms) and llm agents: A literature review. High-Confidence Computing 5 (2): 100300, 2025.
Yang, K., Liu, Y., Chaudhary, S., Fakoor, R., Chaudhari, P., Karypis, G., and Rangwala, H. Agentoccam: A simple yet strong baseline for llm-based web agents. In International Conference on Learning Representations, 2025.
Yao, S., Chen, H., Yang, J., and Narasimhan, K. Webshop: Towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems. Vol. 35, 2022.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations, 2023.
Zhang, Z., Lyu, Z., Gong, J., Yi, H., Wang, X., Zhou, Y., Yang, J., Nie, P., Huang, Y., and Chen, W. Browseragent: Building web agents with human-inspired web browsing actions, 2025.
Zhao, A., Huang, D., Xu, Q., Lin, M., Liu, Y.-J., and Huang, G. Expel: Llm agents are experiential learners, 2023.
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y. Gpt-4v(ision) is a generalist web agent, if grounded. In Proceedings of the 41st International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 235. PMLR, pp. 61349–61385, 2024.
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., and Neubig, G. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, 2024.
