Understanding Storage Trade-offs for Containerized Dataflows Across the Computing Continuum

  • Wesley Ferreira Universidade Federal Fluminense (UFF)
  • Luiz Viana Universidade Federal Fluminense (UFF)
  • Lucas Natan Universidade Federal Fluminense (UFF)
  • Liliane Kunstmann Universidade Federal do Rio de Janeiro (UFRJ)
  • Daniel de Oliveira Universidade Federal Fluminense (UFF)

Resumo


Dataflows model and execute workloads in scientific and data-intensive domains by representing applications as tasks linked by data dependencies. This structure supports scheduling, execution, and reproducibility. Many dataflows process large datasets in distributed environments such as HPC clusters and cloud platforms. However, tasks often have heterogeneous computational and data requirements, motivating the use of the computing continuum, where workloads migrate across environments. Containers enable portable execution across heterogeneous infrastructures, but storage strategies for data sharing remain insufficiently explored. This paper evaluates multiple storage approaches, analyzing how data size, granularity, and role influence performance and efficiency.
Palavras-chave: dataflows, computing continuum, containers

Referências

Al-Dulaimy, A. et al. (2024). The computing continuum: From iot to the cloud. Internet of Things, 27:101272.

Bachiega, N. G. et al. (2020). Performance evaluation of container’s shared volumes. In 2020 IEEE ICSTW, pages 114–123.

Chazapis, A. et al. (2021). A unified storage layer for supporting distributed workflows in kubernetes. In Proc. of the CHEOPS ’21, New York, NY, USA. ACM.

Crosas, M. (2011). The dataverse network®: An open-source application for sharing, discovering and preserving data. D Lib Mag., 17(1/2).

de Oliveira, D. et al. (2019). Data-Intensive Workflow Management: For Clouds and Data-Intensive and Scalable Computing Environments. Morgan & Claypool Publishers.

Di Tommaso, P. et al. (2017). Nextflow enables reproducible computational workflows. Nature Biotechnology, 35(4):316–319.

Köster, J. and Rahmann, S. (2012). Snakemake—a scalable bioinformatics workflow engine. Bioinformatics, 28(19):2520–2522.

Liu, F. et al. (2019). Kubestorage: A cloud native storage engine for massive small files. In 2019 6th BESC, pages 1–4.

Mercl, L. and Pavlik, J. (2019). Public cloud kubernetes storage performance analysis. In ICCCI 2019, page 649–660, Berlin, Heidelberg. Springer-Verlag.

Merenstein, A. et al. (2021). CNSBench: A cloud native storage benchmark. In 19th USENIX FAST, pages 263–276. USENIX Association.

Plale, B. A. et al. (2021). Reproducibility practice in high-performance computing: Community survey results. Comput. Sci. Eng., 23(5):55–60.

Ponte Ahón, S. A. et al. (2025). Retroh-unlp: Conservation of the historical observational work of the astronomical observatory of la plata with computer vision. J. Comput. Cult. Herit., 18(4).

Sakellariou, R. et al. (2009). Mapping workflows on grid resources: Experiments with the montage workflow. In ERCIM W. Group on Grids, pages 119–132.

Shah, S. T., Lahaye, R. J. W. E., Kazmi, S. A. A., Chung, M. Y., and Hasan, S. F. (2014). Htcondor system for running extensive simulations related to D2D communication. In ICTC, pages 283–284. IEEE.

Struhár, V., Behnam, M., Ashjaei, M., and Papadopoulos, A. V. (2020). Real-time containers: A survey. In Fog-IoT, volume 80 of OASIcs, pages 7:1–7:9.

Suter, F. et al. (2026). A terminology for scientific workflow systems. Future Gener. Comput. Syst., 174:107974.

Tarasov, V. et al. (2019). Evaluating docker storage performance: from workloads to graph drivers. Cluster Computing, 22(4):1159–1172.

Wang, L. et al. (2018). Sciapps: a bioinformatics workflow platform powered by xsede and cyverse. In Proc. of the PEARC ’18, PEARC ’18, New York, NY, USA. ACM.

Woyames, P. et al. (2025). Avaliação da capacidade de llms para especificar workflows. In Anais do XIX Brazilian e-Science Workshop, pages 81–88, Porto Alegre, RS, Brasil. SBC.

Zhao, N. et al. (2021). Large-scale analysis of docker images and performance implications for container storage systems. IEEE Trans. Parallel Distrib. Syst., 32(4):918–930.
Publicado
08/09/2026
FERREIRA, Wesley; VIANA, Luiz; NATAN, Lucas; KUNSTMANN, Liliane; DE OLIVEIRA, Daniel. Understanding Storage Trade-offs for Containerized Dataflows Across the Computing Continuum. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 854-860. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249438.