Where Does Hardness Come From? An Instance Hardness Analysis of Substance-Use Prediction in Brazilian Adolescents
Resumo
The intrinsic difficulty of a classification dataset is routinely measured, yet its origin is rarely examined: is an instance hard because of the features that describe it, or because of the target to be predicted? We study this through Instance Hardness (IH) in a real, highly imbalanced problem — predicting frequent substance use among 165,838 Brazilian adolescents from the 2019 PeNSE survey, a setting that is hard by construction: the data are self-reported and the positive class has a prevalence of only about 4%. Hardness is quantified per instance in two independent ways — empirically, as the misclassification propensity across a committee of four learning architectures, and geometrically, from local class overlap (kDN), which uses no model. Three findings emerge across ten data partitions. First, IH is largely but not uniformly model-invariant (Spearman ρ > 0.94 among the four committee members; ≈ 0.79–0.97 against a fifth, neural architecture), which justifies treating hardness as a property of the data rather than of any single learner. Second, local overlap explains only part of the difficulty: kDN predicts true errors with an AUC of ≈ 0.73, against ≈ 0.95 for empirical IH, leaving a substantial residual attributable to non-local structure such as label noise or feature interactions. Third, through a controlled ablation, hardness responds to the target, not the features — removing an entire block of strong features leaves the hardness distribution essentially unchanged, whereas enlarging the target raises mean hardness roughly 2.5×. In this domain, difficulty is governed by what is predicted, not by which variables are used to predict it; past a reasonable feature set, the performance ceiling is more productively addressed by reconsidering the target than by further feature engineering.
Palavras-chave:
class overlap, data complexity, imbalanced classification, instance hardness
Referências
Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M. Optuna: A Next-Generation Hyperparameter Optimization Framework. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Anchorage, USA, pp. 2623–2631, 2019.
Chen, T. and Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, USA, pp. 785–794, 2016.
Cruz, R. M. O., Sabourin, R., and Cavalcanti, G. D. C. Prototype selection for dynamic classifier and ensemble selection. Neural Computing and Applications 29 (2): 447–457, 2018.
IBGE. Pesquisa Nacional de Saúde do Escolar: 2019. IBGE, Coordenação de População e Indicadores Sociais, Rio de Janeiro, 2021.
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the International Conference on Neural Information Processing Systems. Long Beach, USA, pp. 3149–3157, 2017.
Lorena, A. C., Garcia, L. P. F., Lehmann, J., de Souto, M. C. P., and Ho, T. K. How complex is your classification problem? a survey on measuring classification complexity. ACM Computing Surveys 52 (5): 1–34, 2019.
Nunes, G. H., Martins, G. O., Forster, C. H. Q., and Lorena, A. C. Using instance hardness measures in curriculum learning. In Encontro Nacional de Inteligência Artificial e Computacional (ENIAC). Online, pp. 177–188, 2021.
Paiva, P. Y. A., Moreno, C. C., Smith-Miles, K., Valeriano, M. G., and Lorena, A. C. Relating instance hardness to classification performance in a dataset: a visual approach. Machine Learning 111 (8): 3085–3123, 2022.
Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., and Gulin, A. CatBoost: Unbiased Boosting with Categorical Features. In Proceedings of the International Conference on Neural Information Processing Systems. Montréal, Canada, pp. 6639–6649, 2018.
Smith, M. R., Martinez, T., and Giraud-Carrier, C. An instance level analysis of data complexity. Machine Learning 95 (2): 225–256, 2014.
Torquette, G. P., Nunes, V. S., Paiva, P. Y. A., and Lorena, A. C. Instance hardness measures for classification and regression problems. Journal of Information and Data Management 15 (1): 152–174, 2024.
Chen, T. and Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, USA, pp. 785–794, 2016.
Cruz, R. M. O., Sabourin, R., and Cavalcanti, G. D. C. Prototype selection for dynamic classifier and ensemble selection. Neural Computing and Applications 29 (2): 447–457, 2018.
IBGE. Pesquisa Nacional de Saúde do Escolar: 2019. IBGE, Coordenação de População e Indicadores Sociais, Rio de Janeiro, 2021.
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the International Conference on Neural Information Processing Systems. Long Beach, USA, pp. 3149–3157, 2017.
Lorena, A. C., Garcia, L. P. F., Lehmann, J., de Souto, M. C. P., and Ho, T. K. How complex is your classification problem? a survey on measuring classification complexity. ACM Computing Surveys 52 (5): 1–34, 2019.
Nunes, G. H., Martins, G. O., Forster, C. H. Q., and Lorena, A. C. Using instance hardness measures in curriculum learning. In Encontro Nacional de Inteligência Artificial e Computacional (ENIAC). Online, pp. 177–188, 2021.
Paiva, P. Y. A., Moreno, C. C., Smith-Miles, K., Valeriano, M. G., and Lorena, A. C. Relating instance hardness to classification performance in a dataset: a visual approach. Machine Learning 111 (8): 3085–3123, 2022.
Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., and Gulin, A. CatBoost: Unbiased Boosting with Categorical Features. In Proceedings of the International Conference on Neural Information Processing Systems. Montréal, Canada, pp. 6639–6649, 2018.
Smith, M. R., Martinez, T., and Giraud-Carrier, C. An instance level analysis of data complexity. Machine Learning 95 (2): 225–256, 2014.
Torquette, G. P., Nunes, V. S., Paiva, P. Y. A., and Lorena, A. C. Instance hardness measures for classification and regression problems. Journal of Information and Data Management 15 (1): 152–174, 2024.
Publicado
19/10/2026
Como Citar
MOTA, Thalles; GARCIA, Leonardo Arruda Vilela; MARTINS, Claudia.
Where Does Hardness Come From? An Instance Hardness Analysis of Substance-Use Prediction in Brazilian Adolescents. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 265-272.
ISSN 2763-8944.
DOI: https://doi.org/10.5753/kdmile.2026.31988.
