Expeditus: Database Optimization Based on Workload Analysis with Structural Query Clustering
Resumo
Workload analysis is fundamental to database performance management, as heterogeneous and complex workloads require automated identification of similar query patterns. In this context, this work investigates whether SQL query clustering based on structural embeddings extracted from the workload can support more efficient database tuning strategies. As an experimental foundation, we employ the TPC-H benchmark, whose queries were transformed into vector representations using the Qwen3-Embedding-0.6B model. After dimensionality reduction with UMAP, we applied the HDBSCAN algorithm to identify structurally coherent query clusters. We then evaluated whether indexes derived from a representative query within each cluster could improve the performance of the remaining queries in the same group. The experiments resulted in seven structurally cohesive clusters, demonstrating the model’s ability to capture structural similarities across distinct queries. The results indicate that optimizing a single query can benefit other queries within the same cluster, supporting the use of structure-driven tuning strategies for complex database workloads.
Palavras-chave:
Database, Workloads, Embeddings, Query Clusters
Referências
Abdi, H. and Williams, L. J. (2010). Principal component analysis. Wiley interdisciplinary reviews: computational statistics, 2(4):433–459.
Almeida, J. A. O. d. S. and Moura, R. S. (2024). Investigação de métodos de similaridade textual no contexto da avaliação automática de questões discursivas. In Escola Regional de Computação do Ceará, Maranhão e Piauí(ERCEMAPI), pages 110–118. SBC.
BRASIL, R. M., de Freitas BRITO, D., Homero, J., et al. (2025). Fundamentos e aplicação de redes neurais artificiais à saúde e humanidades: Tutorial–parte 1. Revista Presença, 11(26):405–419.
Cer, D., Yang, Y., Kong, S.-y., Hua, N., Limtiaco, N., John, R. S., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., et al. (2018). Universal sentence encoder. arXiv preprint arXiv:1803.11175.
do Amaral, K. P. N., de Andrade, L. P., and de Campos, E. A. V. (2025). O uso da inteligência artificial na otimização de consultas em bancos de dados. Revista Camalotes, 4(1).
Guabtni, A., Ranjan, R., and Rabhi, F. A. (2013). A workload-driven approach to database query processing in the cloud. The Journal of Supercomputing, 63(3):722–736.
Koßmann, J. (2023). Unsupervised database optimization: efficient index selection & data dependency-driven query optimization. PhD thesis, Universität Potsdam.
McInnes, L., Healy, J., Astels, S., et al. (2017). hdbscan: Hierarchical density based clustering. J. Open Source Softw., 2(11):205.
McInnes, L., Healy, J., and Melville, J. (2018). Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426.
Oloruntoba, N. (2025). Ai-driven autonomous database management: Self-tuning, predictive query optimization, and intelligent indexing in enterprise it environments. World Journal of Advanced Research and Reviews, 25(2):1558–1580.
Petrola, L., Brayner, A., and Franco, W. (2025). Heuristic-guided text-to-sql translation with llms: Optimizing natural language interfaces for relational databases. In Simpósio Brasileiro de Banco de Dados (SBBD), pages 126–139. SBC.
Shaheen, N., Raza, B., and Malik, A. K. (2018). A cbr model for workload characterization in autonomic database management system. In 2018 14th international conference on emerging technologies (ICET), pages 1–6. IEEE.
Van der Maaten, L. and Hinton, G. (2008). Visualizing data using t-sne. Journal of machine learning research, 9(11).
Wang, J., Li, T., Wang, A., Liu, X., Chen, L., Chen, J., Liu, J., Wu, J., Li, F., and Gao, Y. (2023). Real-time workload pattern analysis for large-scale cloud databases. arXiv preprint arXiv:2307.02626.
Almeida, J. A. O. d. S. and Moura, R. S. (2024). Investigação de métodos de similaridade textual no contexto da avaliação automática de questões discursivas. In Escola Regional de Computação do Ceará, Maranhão e Piauí(ERCEMAPI), pages 110–118. SBC.
BRASIL, R. M., de Freitas BRITO, D., Homero, J., et al. (2025). Fundamentos e aplicação de redes neurais artificiais à saúde e humanidades: Tutorial–parte 1. Revista Presença, 11(26):405–419.
Cer, D., Yang, Y., Kong, S.-y., Hua, N., Limtiaco, N., John, R. S., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., et al. (2018). Universal sentence encoder. arXiv preprint arXiv:1803.11175.
do Amaral, K. P. N., de Andrade, L. P., and de Campos, E. A. V. (2025). O uso da inteligência artificial na otimização de consultas em bancos de dados. Revista Camalotes, 4(1).
Guabtni, A., Ranjan, R., and Rabhi, F. A. (2013). A workload-driven approach to database query processing in the cloud. The Journal of Supercomputing, 63(3):722–736.
Koßmann, J. (2023). Unsupervised database optimization: efficient index selection & data dependency-driven query optimization. PhD thesis, Universität Potsdam.
McInnes, L., Healy, J., Astels, S., et al. (2017). hdbscan: Hierarchical density based clustering. J. Open Source Softw., 2(11):205.
McInnes, L., Healy, J., and Melville, J. (2018). Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426.
Oloruntoba, N. (2025). Ai-driven autonomous database management: Self-tuning, predictive query optimization, and intelligent indexing in enterprise it environments. World Journal of Advanced Research and Reviews, 25(2):1558–1580.
Petrola, L., Brayner, A., and Franco, W. (2025). Heuristic-guided text-to-sql translation with llms: Optimizing natural language interfaces for relational databases. In Simpósio Brasileiro de Banco de Dados (SBBD), pages 126–139. SBC.
Shaheen, N., Raza, B., and Malik, A. K. (2018). A cbr model for workload characterization in autonomic database management system. In 2018 14th international conference on emerging technologies (ICET), pages 1–6. IEEE.
Van der Maaten, L. and Hinton, G. (2008). Visualizing data using t-sne. Journal of machine learning research, 9(11).
Wang, J., Li, T., Wang, A., Liu, X., Chen, L., Chen, J., Liu, J., Wu, J., Li, F., and Gao, Y. (2023). Real-time workload pattern analysis for large-scale cloud databases. arXiv preprint arXiv:2307.02626.
Publicado
08/09/2026
Como Citar
DUARTE, Marlon Gonçalves; MARINHO, Júlio César de Lira; BRAYNER, Ângelo Rocalli Alencar; FRANCO, Wellington.
Expeditus: Database Optimization Based on Workload Analysis with Structural Query Clustering. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 523-535.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249252.
