The types of annotations, encoding, and interfaces of the Lácio-Web Project: How far are we from international standards for corpora?

  • Sandra Maria Aluísio USP
  • Leandro H. M. de Oliveira USP
  • Gisele Montilha Pinheiro USP

Abstract


Neste artigo discutimos questões relacionadas aos tipos de anotação, codificação (no sentido de “forma de representação”), ferramentas e arquiteturas para dados, considerados padrões em ambientes de desenvolvimento e disponibilização de córpus. É discutido, também, quão próximas estão as decisões do Projeto Lácio-Web desses padrões e da construção de um Corpus Nacional do Português Brasileiro.

References

Aluísio, S. M., Pinheiro, G. M., Finger, M., Nunes, M.G.V. and Tagnin, S. E. O. (2003a) “The Lacio-Web Project: overview and issues in Brazilian Portuguese corpora creation”, Corpus Linguistics 2003, Lancaster, UK, Proceedings of Corpus Linguistics 2003. Lancaster: 2003. v. 16, 14-21.

Aluísio, S. M., Pelizzoni, J. M., Marchi, A. R., Oliveira, L. H., Manenti, R. and Maquiafável, V. (2003b) “An account of the challenge of tagging a reference corpus of Brazilian Portuguese”, Lecture Notes on Artificial Intelligence 2721, 110-117.

Aluísio, S. M., Pinheiro, G. M., Manfrim, A. M. P., Oliveira, L. H. M. de, Genovês Jr. L. C. e Tagnin, S. E. O. (2004) “The Lácio-Web: Corpora and Tools to advance Brazilian Portuguese Language Investigations and Computational Linguistic Tools”, LREC 2004. Proceedings of LREC, 2004, Lisboa, Portugal.

Andersen, M. S., Asmussen, H. e Asmussen, J. (2002) “The project of Korpus 2000 going public”, Proceedings of Euralex 2002, 291-299.

Ide, N. e Romary, L. (2003). “Outline of the International Standard Linguistic Annotation Framework.”, Proceedings of ACL'03 Workshop on Linguistic Annotation: Getting the Model Right, Sapporo, 1-5.

Ide, N., Romary, L., de la Clergerie, E. (2003). “International Standard for a Linguistic Annotation Framework”, Proceedings of HLT-NAACL'03 Workshop on The Software Engineering and Architecture of Language Technology, Edmunton.

Ide, N. e Macleod, C. (2001). “The American National Corpus: A Standardized Resource of American English”, Proceedings of Corpus Linguistics 2001, Lancaster UK.

Ide, N. (2000). “The XML Framework and Its Implications for the Development of Natural Language Processing Tools”, Proceedings of the COLING Workshop on Using Toolsets and Architectures to Build NLP Systems, Luxembourg, 5 August 2000.

Ide, N. e Brew, C. (2000). “Requirements, Tools, and Architectures for Annotated Corpora”, Proceedings of Data Architectures and Software Support for Large Corpora. Paris: European Language Resources Association, 1-5.

Ide, N., Bonhomme, P. e Romary, L. (2000). “XCES: An XML-based Standard for Linguistic Corpora”, Proceedings of the Second Language Resources and Evaluation Conference (LREC), Athens, Greece, 825-830.

Ide, N. (1998). “Corpus Encoding Standard: SGML Guidelines for Encoding Linguistic Corpora”, Proceedings of the First International Language Resources and Evaluation Conference, Granada, Spain, 463-470.

Santos D. e Sarmento, L. (2003). "O projecto AC/DC: acesso a corpora / disponibilização de corpora", Amália Mendes & Tiago Freitas (orgs.), Anais do XVIII Encontro da Associação Portuguesa de Linguística (Porto, 2-4 de Outubro de 2002), APL, 2003, 705-717.

Santos, D. e Bick, E. (2000) “Providing Internet Acces to Portuguese Corpora: the AC/DC Project”, Proceedings of the Second International Conference on Language Resources and Evaluation (LREC 2000), 205-210. Atenas, 31 May-2 June 2000.
Published
2004-07-31
ALUÍSIO, Sandra Maria; OLIVEIRA, Leandro H. M. de; PINHEIRO, Gisele Montilha. The types of annotations, encoding, and interfaces of the Lácio-Web Project: How far are we from international standards for corpora?. In: BRAZILIAN SYMPOSIUM IN INFORMATION AND HUMAN LANGUAGE TECHNOLOGY (STIL), 2. , 2004, Salvador/BA. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2004 . p. 84-93.