Sistema de Recomendação em Tempo Real via Processamento de Fluxo sobre Bancos de Dados Distribuídos

  • Lorenzo Ficher UNIPAMPA
  • Marcus Querol UNIPAMPA
  • Mirieli Oliveira UNIPAMPA
  • Natalia Gonçalves UNIPAMPA
  • Vinícius Gonçalves UNIPAMPA
  • Yuri Martins UNIPAMPA
  • Vitor Balsanello UNIPAMPA
  • Maicon Bernardino UNIPAMPA

Resumo


Sistemas de recomendação tradicionais baseados em lote falham em se adaptar às interações do usuário em tempo real. Este artigo propõe uma arquitetura integrada para sistemas de recomendação que combina processamento de fluxo e bancos de dados distribuídos para oferecer sugestões personalizadas de baixa latência. A solução se baseia em uma arquitetura híbrida que utiliza um algoritmo de filtragem colaborativa sensível ao contexto e uma camada de caching otimizada. O objetivo é resolver o gargalo na integração entre as camadas de processamento, algoritmo e persistência, garantindo a entrega de recomendações consistentes e atualizadas em milissegundos.

Referências

Adomavicius, G. and Tuzhilin, A. (2011). Context-aware recommender systems. In Recommender systems handbook, pages 217–253. Springer.

Akidau, T., Bradshaw, R., Chambers, C., Chernyak, S., Fernandez-Moctezuma, R. J., Lax, R., McVeety, S., Mills, D., Perry, F., Schmidt, E., and Whittle, S. (2015). The dataflow model: A practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing. Proceedings of the VLDB Endowment, 8(12):1792–1803.

Barajas, J. M. and Li, X. (2005). Collaborative filtering on data streams. In Jorge, A. M., Torgo, L., Brazdil, P., Camacho, R., and Gama, J., editors, Knowledge Discovery in Databases, pages 429–436. Springer Berlin Heidelberg.

Barré, A., Al-Ghossein, M., and Abdessalem, T. (2021). A survey on stream-based recommender systems. ACM Computing Surveys.

Blamey, B., Hellander, A., and Toor, S. (2018). Apache spark streaming, kafka and harmonicio: A performance benchmark and architecture comparison for enterprise and scientific computing. arXiv preprint arXiv:1807.07724.

Breck, E., Holt, G., Sculley, D., and Causmaecker, P. (2017). The ml test score: A rubric for ml production readiness and technical debt. In IEEE International Conference on Big Data, pages 1123–1132. IEEE.

Brewer, E. A. (2000). Towards robust distributed systems (abstract). In 19th Annual ACM Symposium on Principles of Distributed Computing, page 7, New York, USA. ACM.

Cai, C., Ren, Y., Chen, Z., Chen, K., Cui, Z., and Gao, Y. (2020). Distributed real-time recommender system based on shared-nothing architecture. Journal of Ambient Intelligence and Humanized Computing, 11(10):4065–4075.

Chandramouli, B., Levandoski, J. J., Eldawy, A., and Mokbel, M. F. (2011). Streamrec: A real-time recommender system. In ACM SIGMOD International Conference on Management of Data.

Chang, S., Zhang, Y., Tang, J., Yin, D., Chang, Y., and Hasegawa-Johnson, M. A. (2016). Streaming recommender systems. In Proc. of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.

Ezéchiel, K., Kant, S., and Agarwal, D. (2019). A systematic review on distributed databases systems and their techniques. Journal of Theoretical and Applied Information Technology, 96.

Gomez-Uribe, C. A. and Hunt, N. (2015). The netflix recommender system: Algorithms, business value, and innovation. ACM Transactions on Management Information Systems (TMIS), 6(4):1–19.

He, X., Zhang, H., Kan, M.-Y., and Chua, T.-S. (2017). Neural collaborative filtering. In Proc. of the 26th International Conference on World Wide Web (WWW), pages 173–182. ACM.

Jain, P., Kumar, R., and Gupta, G. (2014). A real-time recommender system based on distributed key-value store. In Proc. of the International Conference on Big Data, pages 48–53. IEEE.

Khandal, R. (2024). Distributed database systems: Issues in concurrency control, replication and fault tolerance. ShodhKosh: Journal of Visual and Performing Arts, 5(7).

Kleppmann, M. (2017). Designing data-intensive applications: The big ideas behind reliable, scalable, and maintainable systems. ”O’Reilly Media, Inc.”.

Koren, Y., Bell, R., and Volinsky, C. (2009). Matrix factorization techniques for recommender systems. Computer, 42(8):30–37.

Lian, X.-R., Liu, Y., Sun, Y., Wang, R., and Xu, J.-L. (2014). Parallelized stochastic gradient descent with mini-batching for large-scale recommender systems. In Proc. of the International Conference on Machine Learning.

Liang, J., Zhao, J., Chen, Z., Zhou, L., Li, Y., and Wang, H. (2024). Ensure timeliness and accuracy: A novel sliding window data stream paradigm for live streaming recommendation. IEEE Transactions on Knowledge and Data Engineering, 36(5):1556–1570.

Lops, P., de Gemmis, M., and Semeraro, G. (2011). Content-based recommender systems: State of the art and trends. In Ricci, F., Rokach, L., Shapira, B., and Kantor, P. B., editors, Recommender Systems Handbook, pages 73–105. Springer.

Lu, P., Yue, Y., Yuan, L., and Zhang, Y. (2022). Autoflow: Hotspot-aware, dynamic load balancing for distributed stream processing. In Lai, Y., Wang, T., Jiang, M., Xu, G., Liang, W., and Castiglione, A., editors, Algorithms and Architectures for Parallel Processing, pages 133–151, Cham. Springer International Publishing.

Marz, N. and Warren, J. (2015). Big Data: Principles and best practices of scalable realtime data systems. Manning Publications.

Mi, F., Du, X., and Schuller, B. (2020). ADER: Adaptively distilled exemplar replay towards continual learning for session-based recommendation. In Proc. of the 29th ACM International Conference on Information and Knowledge Management.

Ricci, F., Rokach, L., and Shapira, B., editors (2015). Recommender Systems Handbook. Springer, 2 edition.

Roy, S., Sarker, S., Ghosh, A., Sarkar, S., and Kim, K. M. (2022). A systematic review and research perspective on recommender systems. Journal of Big Data, 9(1):1–47.

Sarwar, B. M., Karypis, G., Konstan, J. A., and Riedl, J. T. (2001). Item-based collaborative filtering recommendation algorithms. In Proceedings of the 10th international conference on World Wide Web, pages 285–295. ACM.

Wang, Y., Zhang, Y., Yin, Y., Yi, D., and Wei, B. (2013). A cluster-based incremental recommendation algorithm on stream processing architecture. In Proc. of the 15th International Conference on Asian Digital Libraries.

Zaharia, M., Das, T., Li, H., Shenker, S., and Stoica, I. (2013). Discretized streams: Fault-tolerant streaming computation at scale. In Proc. of the 24th Symposium on Operating Systems Principles.

Zaharia, M., Xin, R., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J., Shenker, S., and Stoica, I. (2016). Apache spark: A unified engine for big data processing. Communications of the ACM, 59(11):56–65.

Zhang, S., Yao, L., Sun, A., and Tay, Y. (2019). Deep learning based recommender system: A survey. ACM Computing Surveys, 52(1):1–38.

Zhang, X., Chen, Y., Ma, C., Fang, Y., and King, I. (2024). Influential exemplar replay for incremental learning in recommender systems. In Proc. of the 38th AAAI Conference on Artificial Intelligence. AAAI Press.
Publicado
22/04/2026
FICHER, Lorenzo; QUEROL, Marcus; OLIVEIRA, Mirieli; GONÇALVES, Natalia; GONÇALVES, Vinícius; MARTINS, Yuri; BALSANELLO, Vitor; BERNARDINO, Maicon. Sistema de Recomendação em Tempo Real via Processamento de Fluxo sobre Bancos de Dados Distribuídos. In: ESCOLA REGIONAL DE BANCO DE DADOS (ERBD), 21. , 2026, Dois Vizinhos/PR. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 109-118. ISSN 2595-413X. DOI: https://doi.org/10.5753/erbd.2026.21358.