Proposal for a Platform for Automatic Information Extraction and Summarization in a Web Environment

  • Carlos N. Silla Jr. PUCPR
  • Andre G. Hochuli PUCPR
  • Celso A. A. Kaestner PUCPR

Abstract


In this work we present the architecture of an automatic framework for information extraction and summarization of Brazilian Portuguese news, obtained from the web. This framework was applied to a case study for information extraction of news about soccer games. We performed three experiments with the goal of establishing baseline for future experiments and also to observe the behavior of the existing summarizers acting in a complementary way in the task of information extraction.

References

Baeza-Yates, R. and Ribeiro-Neto, B. (1999). Modern Information Retrieval. Addison-Wesley.

Brin, S., Motwani, R., Page, L., and Winograd, T. (1998). What can you do with a web in your pocket? Data Engineering Bulletin, 21(2):37–47.

Califf, M. E. (1998). Relational learning techniques for natural language extraction. Technical Report AI98-276, Univ. of Texas at Austion.

Ciravegna, F. (2001). Adaptive information extraction form text by rule induction and generalization. In Proceedings of the 17 th. International Joint Conference on Artificial Intelligence, IJCAI’01.

Freitag, D. (1998). Information extraction from HTML: Application of a general learning approach. In Proceedings of the 15th. Conference on Artificial Intelligence, AAAI-98, pages 517–523.

Grishman, R. (1997). Information extraction: Techniques and challenges. In Information Extraction: A Multidisciplinary Approach to an Emerging Information Technology, pages 10–27.

Larocca Neto, J., Freitas, A. A., and Kaestner, C. A. A. (2002). Automatic text summarization using a machine learning approach. In XVI Brazilian Symposium on Artificial Intelligence, number 2057 in Lecture Notes in Computer Science, pages 205–215, Porto de Galinhas, PE, Brazil.

Larocca Neto, J., Santos, A. D., Kaestner, C. A. A., and Freitas, A. A. (2000). Document clustering and text summarization. In Proc. 4th Int. Conf. Practical Applications of Knowledge Discovery and Data Mining (PADD-2000), pages 41–55, London: The Practical Application Company.

Luhn, H. (1958). The automatic creation of literature abstracts. IBM Journal of Research and Development, 2(92):159–165.

Lyman, P. and H.R., V. (2003). How much information. Retrieved from [link] [Acesso em: 01/19/04].

Mani, I. (2001). Automatic Summarization. John Benjamins Publishing Company.

Manning, C. D. and Schutze, H. (2001). Foundations of Statistical Natural Language Processing. The MIT Press.

McKeown, K., Barzilay, R., Evans, D., Hatzivassiloglou, V., Klavans, J. L., Nenkova, A., Sable, C., Schiffman, B., and Sigelman, S. (2003). Projeto columbia newsblaster.

Mitchell, T. M. (1997). Machine Learning. McGraw-Hill.

Pardo, T. A. S., Rino, L. H. M., and Nunes, M. G. V. (2003). Gistsumm: A summarization tool based on a new extractive method. In 6th Workshop on Computational Processing of the Portuguese Language - Written and Spoken, number 2721 in Lecture Notes in Artificial Intelligence, pages 210–218, Germany.

Salton, G., Allan, J., and Singhal, A. (1996). Automatic text decomposition and structuring. Information Processing and Management, 32(2):127–138.

Silla Jr., C. N., Kaestner, C. A. A., and Freitas, A. A. (2003). A non-linear topic detection method for text summarization using wordnet. In 1o Workshop em Tecnologia da Informação e Linguagem Humana (TIL), São Carlos, SP, Brazil.

Sparck-Jones, K. (1999). Advances in Automatic Text Summarization, chapter Automatic Summarizing: factors and directions, pages 1 – 12. MIT Press.
Published
2004-07-31
SILLA JR., Carlos N.; HOCHULI, Andre G.; KAESTNER, Celso A. A.. Proposal for a Platform for Automatic Information Extraction and Summarization in a Web Environment. In: BRAZILIAN SYMPOSIUM IN INFORMATION AND HUMAN LANGUAGE TECHNOLOGY (STIL), 2. , 2004, Salvador/BA. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2004 . p. 40-46.