N-gram Language Models for Large-Vocabulary Speech Recognition

  • Ênio Silva UFPA
  • Marcus Pantoja UFPA
  • Jackline Celidônio UFPA
  • Aldebaro Klautau UFPA

Abstract


This work describes preliminary results on N-gram language models applied to Brazilian Portuguese. The project is part of an effort to develop a large vocabulary continuous speech recognition system, where language modeling plays a fundamental role. We present a brief summary of state-of-art techniques, including the recently proposed interpolated additive (AI) model. We also describe simulation results, which show that the AI model is competitive with some well-established techniques.

References

Brown, P. F., Pietra, S. A. D., Pietra, V. J. D., Lai, J. C., and Mercer, R. L. (1992). An estimate of an upper bound for the entropy of english. Computational Linguistics, 18:31–40.

Chen, S. F. (1996). Building Probabilistic Models for Natural Language. PhD Thesis.

Chen, S. F. and Goodman, J. (1999). An empirical study of smoothing techniques for language modeling. Computer Speech and Language, 13:359–394.

de Laplace, P. S. (1816). Essay Philosophique sur la Probabilités. Courcier Imprimeur, Paris.

Fagundes, R. and Sanches, I. (2003). Uma nova abordagem fonético-fonológica em sistemas de reconhecimento de fala espontânea. Revista da Sociedade Brasileira de Telecomunicações, 95.

Gale, W. A. and Church, K. W. (1994). What’s wrong with adding one. Corpus-Based Research Into Language (Oosdijk, N. and de Haan, P., eds).

Huang, X., Acero, A., and Hon, H.-W. (2001). Spoken language processing. Prentice-Hall.

Jeffreys, H. (1939). Theory of Probability. Clarendon, Oxford.

Jelinek, F. and Mercer, R. L. (1980). Interpolated estimation of markov source parameters from sparse data. Proceedings of the Workshop on Pattern Recognition in Practice, pages 381–397.

Jevtic, N. and Orlitsky, A. (2003). On the relation between additive smoothing and universal coding. IEEE ASRU.

Kneser, R. and Ney, H. (1995). Improved backing-off for m-gram language modeling. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, 1:181–184.

Ney, H. and Essen, U. (1991). On smoothing techniques for bigram-based natural language modeling. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, 2:825–829.

Ney, H., Essen, U., and Kneser, R. (1994). On structuring probabilistic dependences in stochastic language modeling. Computer Speech and Language, 8:1–38.

Ney, H., Martin, S., and Wessel, F. (1997). Statistical language modeling using leaving-one-out. In Corpus Based Methods in Language and Speech Processing, pages 174–207.

Pessoa, L., Violaro, F., and Barbosa, P. (1999a). Modelo de língua baseado em gramática gerativa aplicado ao reconhecimento de fala contínua. In XVII Simpósio Brasileiro de Telecomunicações, pages 455–458.

Pessoa, L., Violaro, F., and Barbosa, P. (1999b). Modelos da língua baseados em classes de palavras para sistema de reconhecimento de fala contínua. Revista da Sociedade Brasileira de Telecomunicações, 14(2):75–84.

Santos, S. and Alcaim, A. (2002). Um sistema de reconhecimento de voz contínua dependente da tarefa em língua portuguesa. Revista da Sociedade Brasileira de Telecomunicações, 17(2):135–147.

Seara et al, I. (2003). Geração automática de variantes de léxicos do português brasileiro para sistemas de reconhecimento de fala. In XX Simpósio Brasileiro de Telecomunicações, pages v.1. p.1–6.

Witten, I. H. and Bell, T. C. (1991). The zero frequency problem: Estimating the probabilities of novel events in adaptive text compression. IEEE Transactions on Information Theory, 37(4):1085–94.

Ynoguti, C. A. and Violaro, F. (1999). Influência da transcrição fonética no desempenho de sistemas de reconhecimento de fala contínua. In XVII Simpósio Brasileiro de Telecomunicações, pages 449–454.
Published
2004-07-31
SILVA, Ênio; PANTOJA, Marcus; CELIDÔNIO, Jackline; KLAUTAU, Aldebaro. N-gram Language Models for Large-Vocabulary Speech Recognition. In: BRAZILIAN SYMPOSIUM IN INFORMATION AND HUMAN LANGUAGE TECHNOLOGY (STIL), 2. , 2004, Salvador/BA. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2004 . p. 75-83.