Lições de Implementação de um Projeto de Ciência de Dados para Investigação Financeira: Diretrizes, Evidências e Resultados Preliminares
Resumo
Este artigo apresenta diretrizes, lições aprendidas e evidências preliminares de um projeto real de ciência de dados voltado ao apoio à investigação financeira pela Polícia Civil de Pernambuco. O estudo discute curadoria data-centric, mitigação de vieses, fluxos analíticos flexíveis, IA centrada no humano, incorporação de conhecimento especializado, processamento escalável e melhoria contínua por meio de feedback humano. Os resultados iniciais indicam potencial prático para organizar evidências, priorizar alertas e apoiar a tomada de decisão investigativa, preservando a supervisão humana.
Referências
Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S. T., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., and Horvitz, E. (2019). Guidelines for human-ai interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–13. ACM.
Barocas, S. and Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104(3):671–732.
Bolton, R. J. and Hand, D. J. (2002). Statistical fraud detection: A review. Statistical Science, 17(3):235–255.
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., and Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12):86–92.
Hernandez Aros, L., Bustamante Molano, L. X., Gutierrez-Portela, F., Moreno Hernandez, J. J., and Rodríguez Barrero, M. S. (2024). Financial fraud detection through the application of machine learning techniques: a literature review. Humanities and Social Sciences Communications, 11:1130.
Holzinger, A. (2016). Interactive machine learning for health informatics: when do we need the human-in-the-loop? Brain Informatics, 3(2):119–131.
Jakubik, J., Vössing, M., Kühl, N., Walk, J., and Satzger, G. (2024). Data-centric artificial intelligence. Business & Information Systems Engineering, 66(4):507–515.
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1–35.
Pang, G., Shen, C., Cao, L., and van den Hengel, A. (2021). Deep learning for anomaly detection: A review. ACM Computing Surveys, 54(2):1–38.
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., and Dennison, D. (2015). Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems, volume 28, pages 2503–2511.
Settles, B. (2012). Active Learning. Synthesis Lectures on Artificial Intelligence and Machine Learning. Morgan & Claypool Publishers.
Shneiderman, B. (2020). Human-centered artificial intelligence: Reliable, safe & trustworthy. International Journal of Human–Computer Interaction, 36(6):495–504.
