Rampart: Reproducible Benchmarking of Data Processing Paradigms with Automated Anti-Leakage Verification

Resumo


The growing adoption of data-processing engines for machine-learning pipelines has increased the need for rigorous benchmarks that ensure validity and reproducibility. However, many studies still lack mechanisms to prevent data leakage. Rampart is presented as an open-source framework executing the same walk-forward validation pipeline across three data-processing paradigms while enforcing anti-leakage invariants. Rampart treats prediction equivalence as a validity clause for cross-paradigm latency comparisons and produces bitwise-reproducible executions. Evaluated on two public time-series panels, the framework reveals a scale-dependent latency crossover and supports inclusion of new paradigms with low engineering overhead.

Palavras-chave: software frameworks, cross-paradigm benchmarking, data leakage, reproducibility, experimental methodology

Referências

Brücke, C., Härtling, P., Palacios, R. D. E., et al. (2023). Tpcx-ai - an industry standard benchmark for artificial intelligence and machine learning systems. Proc. VLDB Endow., 16(12):3649–3661.

de Oliveira, R., Lifschitz, S., Kalinowski, M., et al. (2020). Evaluating database self-tuning strategies in a common extensible framework. In Anais do XXXV Simpósio Brasileiro de Bancos de Dados, pages 97–108, Porto Alegre, RS, Brasil. SBC.

Gundersen, O. E. and Kjensmo, S. (2018). State of the art: Reproducibility in artificial intelligence. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1).

Harby, A. A. and Zulkernine, F. (2025). Data lakehouse: A survey and experimental study. Information Systems, 127:102460.

Hoerl, A. E. and Kennard, R. W. (1970). Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1):55–67.

Hyndman, R. J. and Koehler, A. B. (2006). Another look at measures of forecast accuracy. International Journal of Forecasting, 22(4):679–688.

Kapoor, S. and Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9):100804.

Kiehn, F., Schmidt, M., Glake, D., et al. (2022). Polyglot data management: State of the art & open challenges. Proceedings of the VLDB Endowment, 15(12):3750–3753.

Lakens, D., Scheel, A. M., and Isager, P. M. (2018). Equivalence testing for psychological research: A tutorial. Advances in Methods and Practices in Psychological Science, 1(2):259–269.

McSherry, F., Isard, M., and Murray, D. G. (2015). Scalability! but at what cost? In Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems, HOTOS’15, page 14, USA. USENIX Association.

Nambiar, R. O. and Poess, M. (2006). The making of tpc-ds. In Proceedings of the 32nd International Conference on Very Large Data Bases, VLDB ’06, pages 1049–1058. VLDB Endowment.

Pineau, J., Vincent-Lamarre, P., Sinha, K., et al. (2021). Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). Journal of Machine Learning Research, 22(164):1–20.

Roberts, D. R., Bahn, V., Ciuti, S., et al. (2017). Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8):913–929.

Rupprecht, L., Davis, J., Arnold, K., et al. (2020). Improving reproducibility of data science pipelines through transparent provenance capture. Proceedings of the VLDB Endowment, 13(12):3354–3368.

Sculley, D., Holt, G., Golovin, D., et al. (2015). Hidden technical debt in machine learning systems. In Cortes, C., Lawrence, N., Lee, D., Sugiyama, M., and Garnett, R., editors, Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.

Semmelrock, H., Ross-Hellauer, T., Kopeinik, S., et al. (2025). Reproducibility in machine-learning-based research: Overview, barriers, and drivers. AI Magazine, 46(2):e70002.

Soares, E., Souza, R., Thiago, R., et al. (2021). A recommender for choosing data systems based on application profiling and benchmarking. In Anais do XXXVI Simpósio Brasileiro de Bancos de Dados, pages 265–270, Porto Alegre, RS, Brasil. SBC.

Zaharia, M. A., Chen, A., Davidson, A., et al. (2018). Accelerating the machine learning lifecycle with mlflow. IEEE Data Eng. Bull., 41:39–45.
Publicado
08/09/2026
XAVIER, Eos; RIPARDO, Lívia R. Duarte; MOTA, Fernando M. da; NOGUEIRA, Bruno M.; BORGES, Vanessa A.. Rampart: Reproducible Benchmarking of Data Processing Paradigms with Automated Anti-Leakage Verification. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 742-748. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249291.