Data-Faithful Natural Language Generation for Explaining Fitness-Market Anomalies
Resumo
This study presents an integrated pipeline that transforms anomaly-detection outputs from global fitness-market data into data-faithful narratives. Using a 132-country panel and a 2019 baseline, we analyze observed post-COVID-19 recovery through 2025, while treating 2026 only as a modelled extension. Four complementary detectors—regional-median deviations, linear-regression residuals, Random Forest residuals, and Isolation Forest—are combined through method agreement and mapped to a rule-based typology. A deterministic, template-based Natural Language Generation (NLG) component then produces short, medium, and complete explanations from tabular evidence. In 2025, 14 countries were consensus anomalies; Guyana exemplified market growth without participatory expansion, with revenue recovery of +401.43%, participation recovery of -2.06%, and a 403.50-point gap. All 132 narratives passed automatic factual-consistency checks. These checks validate field preservation rather than fluency, usefulness, or explanatory quality; human evaluation remains future work.
Referências
Alonso, I. and Agirre, E. (2024). Automatic logical forms improve fidelity in table-to-text generation. Expert Systems with Applications, 238:121869.
Cecere, R. and Bernardi, P. (2023). From crisis to opportunity: exploring the interplay between entrepreneurial processes and technological changes in italian fitness industry. International Journal of Entrepreneurship and Innovation Management, 27(3/4):237–268.
Dibia, V. (2023). LIDA: A tool for automatic generation of grammar-agnostic visualizations and infographics using large language models. In Bollegala, D., Huang, R., and Ritter, A., editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 113–126, Toronto, Canada. Association for Computational Linguistics.
Duijvestijn, M., de Wit, G. A., van Gils, P. F., and Wendel-Vos, G. C. W. (2023). Impact of physical activity on healthcare costs: a systematic review. BMC Health Services Research, 23(1):572.
Eime, R., Harvey, J., and Charity, M. (2024). Australian sport and physical activity behaviours pre, during and post-covid-19. BMC Public Health, 24(1):834.
Hasson, R., Sallis, J. F., Coleman, N., Kaushal, N., Nocera, V. G., and Keith, N. (2022). Covid-19: Implications for physical activity, health disparities, and health equity. American Journal of Lifestyle Medicine, 16(4):420–433.
Li, Z. and van Leeuwen, M. (2023). Explainable contextual anomaly detection using quantile regression forests. Data Mining and Knowledge Discovery, 37(6):2517–2563.
Lin, Y., Ruan, T., Liu, J., and Wang, H. (2024). A survey on neural data-to-text generation. IEEE Transactions on Knowledge and Data Engineering, 36(4):1431–1449.
Liu, R., Menhas, R., Dai, J., Saqib, Z. A., and Peng, X. (2022). Fitness apps, live streaming workout classes, and virtual reality fitness for physical activity during the covid-19 lockdown: An empirical study. Frontiers in Public Health, Volume 10 - 2022.
Mersha, M., Lam, K., Wood, J., AlShami, A. K., and Kalita, J. (2024). Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction. Neurocomputing, 599:128111.
Panjei, E., Gruenwald, L., Leal, E., Nguyen, C., and Silvia, S. (2022). A survey on outlier explanations. The VLDB Journal, 31(5):977–1008.
Ren, P., Wang, Y., and Zhao, F. (2023). Re-understanding of data storytelling tools from a narrative perspective. Visual Intelligence, 1(1):11.
Strain, T., Flaxman, S., Guthold, R., Semenova, E., Cowan, M., Riley, L. M., et al. (2024). National, regional, and global trends in insufficient physical activity among adults from 2000 to 2022: a pooled analysis of 507 population-based surveys with 5.7 million participants. The Lancet Global Health, 12(8):e1232–e1243.
Thomson, C., Reiter, E., and Belz, A. (2024). Common flaws in running human evaluation experiments in nlp. Computational Linguistics, 50(2):795–805.
Thomson, C., Reiter, E., and Sundararajan, B. (2023). Evaluating factual accuracy in complex data-to-text. Computer Speech & Language, 80:101482.
Tritscher, J., Krause, A., and Hotho, A. (2023). Feature relevance xai in anomaly detection: Reviewing approaches and challenges. Frontiers in Artificial Intelligence, Volume 6 - 2023.
