Monitoring Explanation Drift in Longitudinal NLP Pipelines: A Cross-Domain Study under Temporal Distribution Shift
Resumo
This paper studies explanation drift : the temporal instability of ranked human-readable terms produced by fixed explanation mechanisms in longitudinal NLP pipelines. We compare corpus-level TF–IDF explanations with model-level explanations derived from logistic-regression coefficients in arXiv and PubMed abstracts from 2014 to 2024. The vocabulary is learned only from the 2014–2016 reference window and then frozen, approximating a deployed monitoring scenario. arXiv is used as the primary supervised setting, while PubMed provides a cross-domain robustness check with query-defined weak labels. Results show that model-level explanations drift more than corpus-level TF–IDF summaries. On arXiv, mean DriftScore is 0.50 for model-level explanations and 0.15 for TF–IDF, a 3.33-fold increase with strong statistical support (padj = 0.0040). On PubMed, the same direction appears with smaller magnitude, 0.20 versus 0.13, but is not significant after Bonferroni correction (padj = 0.0938). The proposed audit rule produces no automatic alerts under the strict configuration because observed drift does not strictly exceed the null 95th percentile, illustrating the distinction between measurable explanation turnover and drift that is sufficiently unusual to trigger operational review.
Referências
Ribeiro, M. T., Singh, S., Guestrin, C.: “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. In: KDD (2016)
Miller, T.: Explanation in Artificial Intelligence: Insights from the Social Sciences. Artificial Intelligence 267 (2019) Alvarez-Melis, D., Jaakkola, T. S.: On the Robustness of Interpretability Methods. ICML Workshop on Human Interpretability in Machine Learning (2018)
Slack, D., Hilgard, S., Jia, E., Singh, S., Lakkaraju, H.: Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods. In: AIES, pp. 180–186 (2020)
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J. F., Dennison, D.: Hidden Technical Debt in Machine Learning Systems. In: NeurIPS (2015)
Rabanser, S., Günnemann, S., Lipton, Z.: Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift. In: NeurIPS (2019)
Hamilton, W., Leskovec, J., Jurafsky, D.: Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change. In: ACL (2016)
Gama, J., Zliobaite, I., Bifet, A., Pechenizkiy, M., Bouchachia, A.: A Survey on Concept Drift Adaptation. ACM Computing Surveys 46(4) (2014)
Hinder, F., Vaquet, V., Hammer, B.: One or Two Things We Know About Concept Drift. Part B: Locating and Explaining Concept Drift. Frontiers in Artificial Intelligence 7 (2024)
Fumagalli, F., Muschalik, M., Hüllermeier, E., Hammer, B.: Incremental Permutation Feature Importance. Machine Learning 112(12), 4863–4903 (2023)
Hinder, F., Vaquet, V., Brinkrolf, J., Hammer, B.: Model-based Explanations of Concept Drift. Neurocomputing 555 (2023)
Hinder, F., Vaquet, V., Hammer, B.: Feature-based Analyses of Concept Drift. Neurocomputing 600 (2024)
Lundberg, S. M., Lee, S.-I.: A Unified Approach to Interpreting Model Predictions. In: Advances in Neural Information Processing Systems (NeurIPS), pp. 4765–4774 (2017)
Sundararajan, M., Taly, A., Yan, Q.: Axiomatic Attribution for Deep Networks. In: International Conference on Machine Learning (ICML), pp. 3319–3328 (2017)
