Cross-Provider Consistency of LLM-Based Semantic Enrichment for e-Science Workflows

Resumo


Integrating large language models (LLMs) into data-intensive scientific workflows requires understanding their robustness, efficiency, and cross-provider consistency. This paper evaluates five LLMs (ChatGPT, Gemini, DeepSeek, Mistral, and Claude) as interchangeable components of a semantic enrichment pipeline, using 2,387 Billboard Year-End Hot 100 songs (2000–2023) as a case study. All models had near-perfect success rates (99.7–100%) but over 20-fold cost differences ($0.46–$10.70). Thematic agreement was substantial (Cohen’s Kappa = 0.70–0.81), and temporal trajectories were consistent across providers (Pearson r = 0.65–0.82). Thus, provider choice minimally affects aggregate analytical conclusions, allowing workflow designs that prioritize cost or latency without sacrificing reproducibility.

Palavras-chave: Large Language Models, Semantic Enrichment, e-Science Workflows

Referências

Askin, N. and Mauskapf, M. (2017). What makes popular culture popular? product features and optimal differentiation in music. American Sociological Review, 82(5):910–944.

Choi, K., Lee, J. H., and Downie, J. S. (2014). What is this song about anyway?: Automatic classification of subject using user interpretations and lyrics. In IEEE/ACM joint conference on digital libraries, pages 453–454. IEEE.

Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and psychological measurement, 20(1):37–46.

Foramitti, M., Nater, U. M., Lamm, C., and Martins, M. (2025). Societal crises disrupt long-term increases in stress, negativity, and simplicity in us billboard song lyrics from 1973 to 2023. Scientific Reports, 15(1):41733.

Gilardi, F., Alizadeh, M., and Kubli, M. (2023). Chatgpt outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30):e2305016120.

Guo, Y., Ovadje, A., Al-Garadi, M. A., and Sarker, A. (2024). Evaluating large language models for health-related text classification tasks with public social media data. Journal of the American Medical Informatics Association, 31(10):2181–2189.

Hu, Q. Q., Azmi-Murad, M. A. M., Azman, A. B., and Nasharuddin, N. A. (2026). Revisiting the role of lyrics in music emotion recognition: A critical analysis of the semantic–perceived emotion gap and methodological challenges. Information Fusion, page 104303.

Hunke, T., Huber, F., and Steffens, J. (2025). The evolution of song lyrics: An nlp-based analysis of popular music in germany from 1954 to 2022. Music & Science, 8:20592043251331155.

Mihalcea, R. and Strapparava, C. (2012). Lyrics, music, and emotions. In Proceedings of the 2012 joint conference on empirical methods in natural language processing and computational natural language learning, pages 590–599.

Papazoglou, V. and Gaizauskas, R. (2021). Using listeners’ interpretations in topic classification of song lyrics. In Proceedings of the 2nd Workshop on NLP for Music and Spoken Audio (NLP4MusA), pages 22–26.

Wang, J. J. and Wang, V. X. (2025). Assessing consistency and reproducibility in the outputs of large language models: Evidence across diverse finance and accounting tasks. arXiv preprint arXiv:2503.16974.

Watanabe, K. and Goto, M. (2020). Lyrics information processing: Analysis, generation, and applications. In Proceedings of the 1st Workshop on NLP for Music and Audio (NLP4MusA), pages 6–12.
Publicado
08/09/2026
DE MELO, Tiago. Cross-Provider Consistency of LLM-Based Semantic Enrichment for e-Science Workflows. In: BRAZILIAN E-SCIENCE WORKSHOP (BRESCI), 20. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 1-8. ISSN 2763-8774. DOI: https://doi.org/10.5753/bresci.2026.249331.