Traditional Machine Learning versus Large Language Models for Predicting Ship Turnaround Times at Brazilian Ports

  • Eduardo C. B. Pacheco USP
  • Jonnathan A. Ramos USP

Resumo


Predicting ship turnaround time at ports is an operationally valuable regression problem on heterogeneous, text-rich tabular data. We ask whether Large Language Models (LLMs) can match purpose-built Machine Learning (ML) on this task, using a leakage-controlled pipeline over seven years of Brazilian (ANTAQ) data. On a fixed 200-berthing test set, with 20 seeds and seed-paired statistics (Wilcoxon, Cohen’s d), we compare five families that read the same cases: classical ML on engineered tabular features plus multilingual sentence embeddings; frontier LLMs (Qwen3-32B, DeepSeek-V3) prompted zero-, 5and 20-shot; a small generative LLM (Qwen2.5-1.5B) fine-tuned on 100–6000 serialized records; an encoder with a regression head fine-tuned on 100–12000 records; and a tabular foundation model (TabFM) reading the raw table in-context. Tuned XGBoost attains R2 ≈ 0.73. Every prompted configuration stays at R2 ≤ 0 and in-context examples do not help; generative fine-tuning barely clears zero, reaching only R2 ≈ 0.10 at 6000 examples. The encoder is the only text regime to rise substantially above R2 = 0, reaching ≈ 0.43 at 12000 examples yet still trailing all ML. In contrast, TabFM, with no task training, matches or beats a data-matched XGBoost at every budget (100–6000; R2 up to ≈ 0.60), leading the low-data regime, though full-data XGBoost still wins and TabFM is far costlier at inference. The predictive signal is tabular, not textual: text-only LLMs miss it, whereas a tabular foundation model recovers it in-context and engineered boosting wins at scale.

Palavras-chave: large language models, machine learning, maritime transportation, regression, ship turnaround time, tabular data

Referências

Abreu, L. R., Maciel, I. S. F., Alves, J. S., Braga, L. C., and Pontes, H. L. J. A decision tree model for the prediction of the stay time of ships in Brazilian ports. Engineering Applications of Artificial Intelligence 117 (1): 105634, 2023.

ANTAQ. Estatístico Aquaviário. [link], 2024.

Chen, T. and Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, USA, pp. 785–794, 2016.

Cohen, J. Statistical Power Analysis for the Behavioral Sciences. Lawrence Erlbaum Associates, Hillsdale, USA, 1988. DeepSeek-AI. DeepSeek-V3 Technical Report. arXiv:2412.19437, 2024.

Dinh, T., Zeng, Y., Zhang, R., Lin, Z., Gira, M., Rajput, S., yong Sohn, J., Papailiopoulos, D., and Lee, K. LIFT: Language-Interfaced Fine-Tuning for Non-Language Machine Learning Tasks. In Advances in Neural Information Processing Systems. New Orleans, USA, pp. 11763–11784, 2022.

Google Research. TabFM: A Tabular Foundation Model (v1.0.0). [link], 2025.

Grinsztajn, L., Oyallon, E., and Varoquaux, G. Why do tree-based models still outperform deep learning on typical tabular data? In Advances in Neural Information Processing Systems. New Orleans, USA, pp. 507–520, 2022.

Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large Language Models Are Zero-Shot Time Series Forecasters. In Advances in Neural Information Processing Systems. New Orleans, USA, pp. 19622–19635, 2023.

Hegselmann, S., Buendia, A., Lang, H., Agrawal, M., Jiang, X., and Sontag, D. TabLLM: Few-shot Classification of Tabular Data with Large Language Models. In Proceedings of the International Conference on Artificial Intelligence and Statistics. Valencia, Spain, pp. 5549–5581, 2023.

Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S. B., Schirrmeister, R. T., and Hutter, F. Accurate predictions on small data with a tabular foundation model. Nature 637 (8045): 319–326, 2025.

Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations. Online, pp. 1–13, 2022.

Pacheco, E. C. B. and Louzada, F. ANTAQ Ship-Turnaround Dataset and Reproduction Artifacts. Zenodo, DOI: 10.5281/zenodo.20549160, 2026.

Pacheco, E. C. B., Martins, R. P., Gualberto, A. S., Germano, J. V. S., Neto, M. M., and Louzada, F. Methodological Pitfalls in Predicting Ship Turnaround Time at Brazilian Ports: An Empirical Audit and a Reproducible Pipeline. Manuscript under review, 2026.

Pacheco, E. C. B. and Ramos, J. A. antaq-ml-vs-llm: code, metrics and predictions. [link], 2026.

Qwen Team. Qwen2.5 Technical Report. arXiv:2412.15115, 2025.

Rao, A. R., Wang, H., and Gupta, C. Predictive Analysis for Optimizing Port Operations. Applied Sciences 15 (6): 2877, 2025.

Reimers, N. and Gurevych, I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Hong Kong, China, pp. 3982–3992, 2019.

Wilcoxon, F. Individual Comparisons by Ranking Methods. Biometrics Bulletin 1 (6): 80–83, 1945.
Publicado
19/10/2026
PACHECO, Eduardo C. B.; RAMOS, Jonnathan A.. Traditional Machine Learning versus Large Language Models for Predicting Ship Turnaround Times at Brazilian Ports. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 249-256. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.32084.