Text Processing for Developing Unrestricted Tamil Text to Speech Synthesis System

Vaibhavi Rajendran   and G  Bharadwaja Kumar

doi:10.17485/ijst/2015/v8i29/72294

Article

Text Processing for Developing Unrestricted Tamil Text to Speech Synthesis System

VIEWS 1332
PDF 1392

Abstract
Full-Text HTML
Full-Text PDF
How to Cite

Indian Journal of Science and Technology

DOI: 10.17485/ijst/2015/v8i29/72294

Year: 2015, Volume: 8, Issue: 29, Pages: 1-10

Original Article

Text Processing for Developing Unrestricted Tamil Text to Speech Synthesis System

Vaibhavi Rajendran^* and G. Bharadwaja Kumar

School of Computing Science and Engineering, VIT University, Chennai Campus, Tamil Nadu, India; [email protected]

This work is licensed under a Creative Commons Attribution 4.0 International License.

Abstract

In this Information and communication technology era, designing interactive computer systems that are effective, efficient, easy, and enjoyable to use is becoming increasingly important. Of the numerous ways explored by researchers to enhance Human-Computer Interaction, Text to Speech or Speech Synthesis affirms to be one such modality for developing better interfaces. The focal point here is to enhance the text processing module of Tamil speech synthesizer with an efficient and robust text normalizer and loan word identifier. Text normalization is performed on unrestricted Tamil text to convert nonstandard words into standard words for the reduction of ambiguous utterances along the interim processing of the words. Loan words in Tamil text are identified in order to improve the pronunciation model of the Tamil speech synthesizer system. In this paper, we describe a ‘semiotic classifier’ based on decision list approach with which we are able to tackle many varieties of non-standard words. We also describe a ‘loan/native word classifier’ based on multiple linear regression which works efficiently even on shorter words of 3 syllables in length. In today’s predominant Digital, InformationCommunication Technology and Human-Computer Interaction era such profound text processors is imperative.
Keywords: Natural Language Processing, Tamil, Text Processing, Text-to-Speech (TTS), Unrestricted Text