INFORMATICA

Informatica

0868-49520868-4952

inf22204

10.15388/Informatica.2011.323

Research article

Semi-Automatic Bilingual Corpus Creation with Zero Entropy Alignments

Laukaitis

Algirdas

algirdas.laukaitis@fm.vgtu.ltVasilecas

Olegas

olegas@fm.vgtu.ltLaukaitis

Ricardas

ricardas.laukaitis@vva.ltPlikynas

Darius

darius.plikynas@vva.ltFundamental Sciences Faculty, Vilnius Gediminas Technical University, Saulėtekio al. 11, LT-10223 Vilnius, LithuaniaAcademy of Business and Management, Research Centre, Basanavičiaus 29A, LT-03109 Vilnius, Lithuania

01012011

2222032240110201001042011

In this paper, we describe a model for aligning books and documents from bilingual corpus with a goal to create “perfectly” aligned bilingual corpus on word-to-word level. Presented algorithms differ from existing algorithms in consideration of the presence of human translator which usage we are trying to minimize. We treat human translator as an oracle who knows exact alignments and the goal of the system is to optimize (minimize) the use of this oracle. The effectiveness of the oracle is measured by the speed at which he can create “perfectly” aligned bilingual corpus. By “Perfectly” aligned corpus we mean zero entropy corpus because oracle can make alignments without any probabilistic interpretation, i.e., with 100% confidence. Sentence level alignments and word-to-word alignments, although treated separately in this paper, are integrated in a single framework. For sentence level alignments we provide a dynamic programming algorithm which achieves low precision and recall error rate. For word-to-word level alignments Expectation Maximization algorithm that integrates linguistic dictionaries is suggested as the main tool for the oracle to build “perfectly” aligned bilingual corpus. We show empirically that suggested pre-aligned corpus requires little interaction from the oracle and that creation of perfectly aligned corpus can be achieved almost with the speed of human reading. Presented algorithms are language independent but in this paper we verify them with English–Lithuanian language pair on two types of text: law documents and fiction literature.

KeywordsViterbi alignmentsdynamic programmingstring alignmentsmachine translationnatural language processingrapid developmentlow-density languages