Applying Q-Learning to Herding Behavior: A Single-Agent Approach in NetLogo

  • Bernardo A. Santos UDESC
  • Fernando Santos UDESC

Resumo


This paper presents an agent-based simulation in NetLogo to evaluate Q-learning within a classic shepherding problem across two distinct operational configurations. The simulation evaluates a static baseline featuring fixed agent initialization and a stationary target, alongside a randomized dynamic scenario characterized by stochastic spawning and an actively moving sheep. Utilizing a low-complexity discrete state space based on relative distances and directional vectors, the dog agent learns to guide the target toward a designated gate via a reward mechanism, completely independent of pre-programmed heuristics. Rather than reporting isolated final metrics, the core analysis evaluates the behavioral stabilization and learning trajectories of the agent under varying environmental complexities. The experimental results demonstrate a policy optimization in both configurations, tracking how the system minimizes operational steps and stabilizes reward retention over extended training timelines.

Referências

Azuma, S., Tabuchi, A., and Sugie, T. (2012). Modeling of sheepdog control. Transactions of the Society of Instrument and Control Engineers, 48(12):882–888.

Hussein, A., Petraki, E., Elsawah, S., and Abbass, H. A. (2022). Autonomous swarm shepherding using curriculum-based reinforcement learning. In Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems, AAMAS ’22, pages 633–641.

Lien, J.-M., Bayazit, O. B., Sowell, R. T., Rodriguez, S., and Amato, N. M. (2004). Shepherding behaviors with multiple shepherds. In Proceedings of IEEE International Conference on Robotics and Automation (ICRA).

Macal, C. M. and North, M. J. (2010). Tutorial on agent-based modelling and simulation. Journal of Simulation, 4(3):151–162.

Railsback, S. F., Lytinen, S. L., and Jackson, S. K. (2006). Agent-based simulation platforms: Review and development recommendations. Simulation, 82(9):609–623.

Strömbom, D., Mann, R. P., Wilson, A. M., Hailes, S., Morton, A. J., Sumpter, D. J. T., and King, A. J. (2014). Solving the shepherding problem: heuristics for herding autonomous, interacting agents. Proceedings of the Royal Society B, 281(1787):20140719.

Van Havermaet, S., Khaluf, Y., and Simoens, P. (2024). Reactive shepherding along a dynamic path. Scientific Reports, 14(1):14915.

Watkins, C. J. C. H. and Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4):279–292.
Publicado
19/10/2026
SANTOS, Bernardo A.; SANTOS, Fernando. Applying Q-Learning to Herding Behavior: A Single-Agent Approach in NetLogo. In: WORKSHOP-ESCOLA DE SISTEMAS DE AGENTES, SEUS AMBIENTES E APLICAÇÕES (WESAAC), 20. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 181-192. ISSN 2326-5434. DOI: https://doi.org/10.5753/wesaac.2026.31724.