Eidolon: An Architecture for the Reproducible Publication of Scientific Data Collections
Resumo
Scientific data collections used in data analytics workflows must reconcile four aspects: (i) data curation, (ii) reusable documentation, (iii) direct integration with a computational environment, and (iv) lightweight distribution. Although repositories and portals support storage and discovery, they do not always provide a runtime usage contract. This work proposes Eidolon, a publication architecture based on curation, structural standardization, compact local miniatures, and on-demand remote expansion. The architecture is instantiated in the integration between tspredbench and tspredit. This case study indicates that Eidolon preserves local discovery, documentation, and reproducibility without embedding the full dataset in the package.
Referências
CRAN Repository Maintainers (2025). CRAN Repository Policy. Technical report, [link].
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Iii, H. D., and Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64:86 – 92.
Hugging Face (2026). Big data? Datasets to the rescue!
Kane, M. J., Emerson, J. W., Haverty, P., and Determan, C. (2025). bigmemory: Manage Massive Matrices with Shared Memory and Memory-Mapped Files. CRAN.
Lhoest, Q., del Moral, A. V., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, L., Davison, J., Šaško, M., et al. (2021). Datasets: A Community Library for Natural Language Processing. In EMNLP 2021 - 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175 – 184.
Ogasawara, E., Alexandrino, F., Gea, C., Santos, D., Salles, R., Birindiba, V., Pacheco, C., Bezerra, E., Pacitti, E., Porto, F., and Carvalho, D. (2025). A Benchmark Time Series Dataset Collection for Prediction Models.
Peng, R. D. (2011). Reproducible research in computational science. Science, 334:1226 – 1227.
Perry, D. E. and Wolf, A. L. (1992). Foundations for the study of software architecture. SIGSOFT Softw. Eng. Notes, 17:40–52.
Pushkarna, M., Zaldivar, A., and Kjartansson, O. (2022). Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI. In ACM International Conference Proceeding Series, FAccT ’22, pages 1776 – 1826, New York, NY, USA. Association for Computing Machinery.
PyPI Docs (2026). Storage Limits. Technical report, [link].
Salles, R., Pacitti, E., Bezerra, E., Marques, C., Pacheco, C., Oliveira, C., Porto, F., and Ogasawara, E. (2023). TSPredIT: Integrated Tuning of Data Preprocessing and Time Series Prediction Models. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 14160 LNCS:41 – 55.
Sandve, G. K., Nekrutenko, A., Taylor, J., and Hovig, E. (2013). Ten Simple Rules for Reproducible Computational Research. PLoS Computational Biology, 9.
Vanschoren, J., van Rijn, J. N., Bischl, B., and Torgo, L. (2014). OpenML: networked science in machine learning. SIGKDD Explor. Newsl., 15:49–60.
Wickham, H. and Bryan, J. (2023). R Packages: Organize, Test, Document, and Share Your Code. "O’Reilly Media, Inc.".
Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., et al. (2016). Comment: The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3.
