Text segmentation with mixed data: evaluating symbolic and neural approaches

  • Amanda Sammer do N. Brandão UFBA
  • Felipe Bomfim Sanches UFBA
  • Alfredo Sena UFBA
  • Laila Mota UFBA
  • Daniela Barreiro Claro UFBA
  • Rerisson Cavalcante UFBA

Resumo


A segmentação de texto vem sendo amplamente difundida e utilizada para extrair informação relevante de documentos. Essas informações estão distribuídas em diversos tipos de estruturas, que incluem dados heterogêneos. A extração desses dados tem sido ultimamente, utilizada por LLMs (Large Language Models). Porém, os treinamentos dos modelos de linguagem não detém características peculiares de dados mistos, tais como dados lexicais e dados fonéticos. Assim, torna-se um desafio analisar a corretude das transcrições de maneira automatizada por LLMs. Neste estudo, a segmentação textual e extração de informação foi paralelizada em duas principais abordagens com o intuito de avaliar a comparação de uma abordagem simbólica e uma abordagem neural. Os resultados demonstram as vantagens e desvantagens em ambas as abordagens para os tipos de dados mistos.

Referências

Cabral, B., Claro, D., and Souza, M. (2024). Exploring open information extraction for Portuguese using large language models. In Gamallo, P., Claro, D., Teixeira, A., Real, L., Garcia, M., Oliveira, H. G., and Amaro, R., editors, Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1, pages 127–136, Santiago de Compostela, Galicia/Spain. Association for Computational Lingustics.

Candido Junior, A., Casanova, E., Soares, A. d. S., Oliveira, F. S. d., Oliveira, L., Fernandes Junior, R. C., Silva, D. P. P. d., Fayet, F. G., Carlotto, B. B., Gris, L. R. S., and Aluísio, S. M. (2023). CORAA ASR: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese. Language Resources and Evaluation, 57(3):1139–1171.

Cardoso, S. A. M. d. S., Mota, J. A., Aguilera, V. d. A., Aragão, M. d. S. S. d., Isquerdo, A. N., Razky, A., Margotti, F. W., and Altenhofen, C. V. (2014). Atlas Linguístico do Brasil, volume 1. EDUEL, Londrina.

Gazzola, M., Souto, H. G., Silva, S., Peixoto, J. S., Siqueira, F., Morais, A. L. P. d., and Gomes, C. (2025). AI-PAVE-Br: Leveraging large language models for enhanced product attribute value extraction through a golden set approach. In Proceedings of the Symposium in Information and Human Language Technology (STIL).

Ghinassi, I., Wang, L., Newell, C., and Purver, M. (2024). Recent trends in linear text segmentation: A survey. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 3084–3095, Miami, Florida, USA. Association for Computational Linguistics.

Godoi, G. A. d., Rivolli, A., Freire, D. L., Pereira, F. S. F., Ventura, N. R., Almeida, A. M. G. d., Garcia, L. P. F., Dias, M. d. S., and Carvalho, A. C. P. d. L. F. d. (2025). Assessing rule-based document segmentation and word normalization for legal ruling classification. In Conference on Digital Government Research, volume 26.

International Phonetic Association (1999). Handbook of the International Phonetic Association: A Guide to the Use of the International Phonetic Alphabet. Cambridge University Press, Cambridge.

Jegan, R. and Henrich, A. (2025). Contrasting traditional models and LLMs: An evaluation based on text segmentation. In Wartena, C. and Heid, U., editors, Proceedings of the 21st Conference on Natural Language Processing (KONVENS 2025): Workshops, pages 274–281, Hannover, Germany. HsH Applied Academics.

Krassovitskiy, A., Mussabayev, R., and Yakunin, K. (2025). Llm-enhanced semantic text segmentation. Applied Sciences, 15(19).

Li, Y., Krishnamurthy, R., Raghavan, S., Vaithyanathan, S., and Jagadish, H. V. (2008). Regular expression learning for information extraction. In Lapata, M. and Ng, H. T., editors, Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, pages 21–30, Honolulu, Hawaii. Association for Computational Linguistics.

Lu, Y., Liu, Q., Dai, D., Xiao, X., Lin, H., Han, X., Sun, L., and Wu, H. (2022). Unified structure generation for universal information extraction. In Muresan, S., Nakov, P., and Villavicencio, A., editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5755–5772, Dublin, Ireland. Association for Computational Linguistics.

Ollama (2024). Ollama: Run large language models locally. [link]. Accessed: 25 fev. 2026.

Ollama Library (2024). Deepseek-r1 8b. [link]. Large Language Model. Acesso em: 25 fev. 2026.

Queiroz, B., Cavalcante, R., and Claro, D. (2023). Desafios da tarefa de extração de informação aberta: uma abordagem metodológica de um corpus automatizado até o corpus manual. In Anais do XIV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 388–392, Porto Alegre, RS, Brasil. SBC.

Tang, X., Zong, Y., Phang, J., Zhao, Y., Zhou, W., Cohan, A., and Gerstein, M. (2024). Struc-bench: Are large language models good at generating complex structured tabular data? In Duh, K., Gomez, H., and Bethard, S., editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 12–34, Mexico City, Mexico. Association for Computational Linguistics.

Xu, D., Chen, W., Peng, W., Zhang, C., Xu, T., Zhao, X., Wu, X., Zheng, Y., Wang, Y., and Chen, E. (2024). Large language models for generative information extraction: a survey. Frontiers of Computer Science, 18(6):186357.
Publicado
19/10/2026
BRANDÃO, Amanda Sammer do N.; SANCHES, Felipe Bomfim; SENA, Alfredo; MOTA, Laila; CLARO, Daniela Barreiro; CAVALCANTE, Rerisson. Text segmentation with mixed data: evaluating symbolic and neural approaches. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 50-61. DOI: https://doi.org/10.5753/stil.2026.26577.