Journal:Informatica
Volume 24, Issue 3 (2013), pp. 435–446
Abstract
The performance of an automatic speech recognition system heavily depends on the used feature set. Quality of speech recognition features is estimated by classification error, but then the recognition experiments must be performed, including both front-end and back-end implementations. We propose a method for features quality estimation that does not require recognition experiments and accelerate automatic speech recognition system development. The key component of our method is usage of metrics right after front-end features computation. The experimental results show that our method is suitable for recognition systems with back-end Euclidean space classifiers.
Journal:Informatica
Volume 18, Issue 3 (2007), pp. 395–406
Abstract
This paper describes a framework for making up a set of syllables and phonemes that subsequently is used in the creation of acoustic models for continuous speech recognition of Lithuanian. The target is to discover a set of syllables and phonemes that is of utmost importance in speech recognition. This framework includes operations with lexicon, and transcriptions of records. To facilitate this work, additional programs have been developed that perform word syllabification, lexicon adjustment, etc. Series of experiments were done in order to establish the framework and model syllable- and phoneme-based speech recognition. Dominance of a syllable in lexicon has improved speech recognition results and encouraged us to move away from a strict definition of syllable, i.e., a syllable becomes a simple sub-word unit derived from a syllable. Two sets of syllables and phonemes and two types of lexicons have been developed and tested. The best recognition accuracy achieved 56.67% ±0.33. The speech recognition system is based on Hidden Markov Models (HMM). The continuous speech corpus LRN0 was used for the speech recognition experiments.
Journal:Informatica
Volume 17, Issue 4 (2006), pp. 587–600
Abstract
There is presented a technique of transcribing Lithuanian text into phonemes for speech recognition. Text-phoneme transformation has been made by formal rules and the dictionary. Formal rules were designed to set the relationship between segments of the text and units of formalized speech sounds – phonemes, dictionary – to correct transcription and specify stress mark and position. Proposed the automatic transcription technique was tested by comparing its results with manually obtained ones. The experiment has shown that less than 6% of transcribed words have not matched.
Journal:Informatica
Volume 10, Issue 4 (1999), pp. 377–388
Abstract
The problem of text-independent speaker recognition based on the use of vocal tract and residue signal LPC parameters is investigated. Pseudostationary segments of voiced sounds are used for feature selection. Parameters of the linear prediction model (LPC) of vocal tract and residue signal or LPC derived cepstral parameters are used as features for speaker recognition. Speaker identification is performed by applying nearest neighbour rule to average distance between speakers. Comparison of distributions of intraindividual and interindividual distortions is used for speaker verification. Speaker recognition performance is investigated. Results of experiments demonstrate speaker recognition performance.
Journal:Informatica
Volume 6, Issue 2 (1995), pp. 167–180
Abstract
The use of vector quantization for speaker identification is investigated. This method differs from the known methods in that the number of centroids is not doubled but increases by 1 at every step. This enables us to obtain identification results at any number of centroids. This method is compared experimentally with the method (Lipeika and Lipeikienė, 1993a, 1993b), where feature vectors of investigative and comparative speakers are compared directly.