Prompt Engineering for Small Language Models: Evaluating ICL for Portuguese Sentiment Analysis
Resumo
The In-Context Learning (ICL) paradigm enables adapting LLMs without parameter tuning. This work evaluates the impact of prompt engineering on 9B models for Brazilian Portuguese sentiment classification, comparing six prompting formats (fewand zero-shot) across three linguistically diverse models (Gemma2-9B-it, Boto-9B-IT, Qwen3.5-9B) and four public datasets, anchored by a fine-tuned encoder (MSA-DistilBERT, ∼0.1B) and a 685B MoE model (DeepSeek-V3.2) under a unified protocol. Results show that all 9B models outperform the encoder and recover 71–101% of the weak-to-strong gap, with Qwen3.5-9B achieving the highest mean coverage (94.5%). Statistical tests indicate that well-defined instructions are the main driver of ICL performance, capturing most of the achievable accuracy even in zero-shot settings, while elaborate formats (roleplay, JSON schema, meta-prompting) add no consistent gain over a minimal instruction-plus-demonstrations prompt, offering practical guidance for deploying SLMs in resource-constrained PT-BR NLP scenarios.
Referências
Alves, D. M., Guerreiro, N. M., Alves, J., et al. (2023). Steering large language models for machine translation with finetuning and in-context learning. [link].
Assis, G., Amorim, A., Carvalho, J., et al. (2024). Exploring Portuguese hate speech detection in low-resource settings: Lightly tuning encoder models or in-context learning of large models? In Gamallo, P., Claro, D., Teixeira, A., Real, L., Garcia, M., Oliveira, H. G., and Amaro, R., editors, Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1, pages 301–311, Santiago de Compostela, Galicia/Spain. Association for Computational Lingustics.
Berg-Kirkpatrick, T., Burkett, D., and Klein, D. (2012). An empirical investigation of statistical significance in NLP. In Tsujii, J., Henderson, J., and Paşca, M., editors, Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, pages 995–1005, Jeju Island, Korea. Association for Computational Linguistics.
Biderman, D., Portes, J., Ortiz, J. J. G., Paul, M., Greengard, P., Jennings, C., King, D., Havens, S., Chiley, V., Frankle, J., Blakeney, C., and Cunningham, J. P. (2024). Lora learns less and forgets less.
Birjali, M., Kasri, M., and Beni-Hssane, A. (2021). A comprehensive survey on sentiment analysis: Approaches, challenges and trends. Knowledge-Based Systems, 226:107134.
Brígida, L. S. (2024). Boto 9b it. [link].
Brown, C. E. (1998). Coefficient of variation. In Applied multivariate statistics in geohydrology and related sciences, pages 155–157. Springer.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020). Language models are few-shot learners.
DeepSeek-AI (2025). Deepseek-v3.2: Pushing the frontier of open large language models.
Devlin, J., Chang, M.-W., Lee, K., et al. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. [link].
Dietterich, T. G. (2000). Ensemble methods in machine learning. In Multiple Classifier Systems, pages 1–15, Berlin, Heidelberg. Springer Berlin Heidelberg.
Dong, Q., Li, L., Dai, D., et al. (2024). A survey on in-context learning. [link].
Dror, R., Baumer, G., Shlomov, S., et al. (2018). The hitchhiker’s guide to testing statistical significance in natural language processing. In Gurevych, I. and Miyao, Y., editors, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1383–1392, Melbourne, Australia. Association for Computational Linguistics.
Fu, J., Ng, S.-K., and Liu, P. (2022). Polyglot prompt: Multilingual multitask promptraining. [link].
Gao, S., Wen, X.-C., Gao, C., et al. (2023). What makes good in-context demonstrations for code intelligence tasks with llms? In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE.
Gemma Team, Riviere, M., Pathak, S., et al. (2024). Gemma 2: Improving open language models at a practical size. DOI: 10.48550/arXiv.2408.00118.
Geng, M., Wang, S., Dong, D., et al. (2024). Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, ICSE ’24, New York, NY, USA. Association for Computing Machinery.
Grandini, M., Bagli, E., and Visani, G. (2020). Metrics for multi-class classification: an overview. [link].
Koehn, P. (2004). Statistical significance tests for machine translation evaluation. In Lin, D. and Wu, D., editors, Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 388–395, Barcelona, Spain. Association for Computational Linguistics.
Lima, J. V. R. J., Pinheiro, V., and Caminha, C. (2026). Portuguese sentiment analysis with open-source LLMs: Models, prompts, and efficient deployment. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 212–221, Salvador, Brazil. Association for Computational Linguistics.
Liu, B. (2012). Sentiment Analysis and Opinion Mining. Springer International Publishing.
Liu, J., Shen, D., Zhang, Y., et al. (2021a). What makes good in-context examples for gpt-3? [link].
Liu, P., Yuan, W., Fu, J., et al. (2021b). Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. DOI: 10.48550/arXiv.2107.13586.
Lu, Y., Bartolo, M., Moore, A., et al. (2022). Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. [link].
Lu, Z., Li, X., Cai, D., et al. (2025). Small language models: Survey, measurements, and insights. [link].
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y. (2025). An empirical study of catastrophic forgetting in large language models during continual fine-tuning.
Maas, A. L., Daly, R. E., Pham, P. T., et al. (2011). Learning word vectors for sentiment analysis. [link].
Marreira, E., De Melo, T., De Oliveira, M., et al. (2025). Rating prediction in brazilian portuguese: A benchmark of large language models. Journal of the Brazilian Computer Society, 31:828–839.
Min, S., Lyu, X., Holtzman, A., et al. (2022). Rethinking the role of demonstrations: What makes in-context learning work? [link].
Nguyen, C. V., Shen, X., Aponte, R., et al. (2025). A survey on small language models. In Angelova, G., Kunilovskaya, M., Escribe, M., and Mitkov, R., editors, Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, pages 807–821, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Oliveira, A., Nascimento, E., Pinheiro, J., et al. (2025). Small, medium, and large language models for text-to-sql. In Maass, W., Han, H., Yasar, H., and Multari, N., editors, Conceptual Modeling, pages 276–294, Cham. Springer Nature Switzerland.
Pang, B. and Lee, L. (2008). Opinion mining and sentiment analysis. Foundations and Trends in Information Retrieval, 2:1–135.
Pereira, D. A. (2021). A survey of sentiment analysis in the portuguese language. Artif. Intell. Rev., 54(2):1087–1115.
Peres, R. S. (2023). Grandes modelos de linguagem na resolução de questões de vestibular: o caso dos institutos militares brasileiros. Master’s thesis, Universidade Federal do Estado do Rio de Janeiro.
Piorino, G., Moreira, V., Lima, L. H. Q., et al. (2025). Sentiment analysis of shared content in brazilian reddit communities. Journal on Interactive Systems, 16:666–686.
Qwen Team (2026). Qwen3.5: Towards native multimodal agents. [link].
Reynolds, L. and McDonell, K. (2021). Prompt programming for large language models: Beyond the few-shot paradigm. [link].
Sanh, V., Debut, L., Chaumond, J., et al. (2020). Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. [link].
Schoch, S. and Ji, Y. (2025). The good, the bad, and the debatable: A survey on the impacts of data for in-context learning. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 29798–29812, Suzhou, China. Association for Computational Linguistics.
Schuck, A. d. F., Garcia, G. L., Manesco, J. R. R., et al. (2025). Evaluating large language models for brazilian portuguese sentiment analysis: A comparative study of multilingual state-of-the-art vs. brazilian portuguese fine-tuned llms. Journal of the Brazilian Computer Society, 31(1):885–917.
Schulhoff, S., Ilie, M., Balepur, N., et al. (2024). The prompt report: A systematic survey of prompting techniques. [link].
Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.
Shin, T., Razeghi, Y., au2, R. L. L. I., et al. (2020). Autoprompt: Eliciting knowledge from language models with automatically generated prompts. [link].
Simmering, P. F. and Huoviala, P. (2023). Large language models for aspect-based sentiment analysis. [link].
Sturua, S., Mohr, I., Akram, M. K., et al. (2024). jina-embeddings-v3: Multilingual embeddings with task lora. [link].
TabularisAI, Gyamfi, S., Borisov, V., et al. (2025). multilingual-sentiment-analysis (revision 69afb83). [link].
Wang, F., Lin, M., Ma, Y., et al. (2025). A survey on small language models in the era of large language models: Architecture, capabilities, and trustworthiness. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, KDD ’25, pages 6173–6183, New York, NY, USA. Association for Computing Machinery.
Wang, X., Wei, J., Schuurmans, D., et al. (2023). Self-consistency improves chain of thought reasoning in language models. [link].
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. (2022). Finetuned language models are zero-shot learners.
Wei, J., Wei, J., Tay, Y., et al. (2023). Larger language models do in-context learning differently. [link].
White, J., Fu, Q., Hays, S., et al. (2023). A prompt pattern catalog to enhance prompt engineering with chatgpt. [link].
Xu, H., Wang, Q., Zhang, Y., Yang, M., Zeng, X., Qin, B., and Xu, R. (2024). Improving in-context learning with prediction feedback for sentiment analysis. In Findings of the Association for Computational Linguistics: ACL 2024, pages 3879–3890.
Zhao, W. X., Zhou, K., Li, J., et al. (2024). A survey of large language models. [link].
Zheng, M., Pei, J., Logeswaran, L., et al. (2024). When ”a helpful assistant” is not really helpful: Personas in system prompts do not improve performances of large language models. [link].
Zhou, Y., Muresanu, A. I., Han, Z., et al. (2023). Large language models are human-level prompt engineers. [link].
