ECAI 2008
18th European
Conference on Artificial Intelligence
Edited by Malik Ghallab, Constantine D. Spyropoulos,
Nikos Fakotakis, Nikos Avouris
pp.338-342
Improved Statistical
Machine Translation Using Monolingual Paraphrases
We propose a novel monolingual sentence paraphrasing method for augmenting the training data for statistical machine translation systems “for free” – by creating it from data that is already available rather than having to create more aligned data. Starting with a syntactic tree, we recursively generate new sentence variants where noun compounds are paraphrased using suitable prepositions, and vice-versa – preposition-containing noun phrases are turned into noun compounds. The evaluation shows an improvement equivalent to 33%–50% of that of doubling the amount of training data.