Logo of Science Foundation Ireland  Logo of the Higher Education Authority, Ireland7 Capacities
Ireland's High-Performance Computing Centre | ICHEC
Home | News | Infrastructure | Outreach | Services | Research | Support | Education & Training | Consultancy | About Us | Login


Title:Bilingually Motivated Domain-Adapted Word Segmentation for Statistical Machine Translation
Authors:Ma Y. and Way A., 2009
Abstract: We introduce a word segmentation approach to languages where word boundaries are not orthographically marked, with application to Phrase-Based Statistical Machine Translation (PB-SMT). Instead of using manually segmented monolingual domain-specific corpora to train segmenters, we make use of bilingual corpora and statistical word alignment techniques. First of all, our approach is adapted for the specific translation task at hand by taking the corresponding source (target) language into account. Secondly, this approach does not rely on manually segmented training data so that it can be automatically adapted for different domains. We evaluate the performance of our segmentation approach on PB-SMT tasks from two domains and demonstrate that our approach scores consistently among the best results across different data conditions.
ICHEC Project:Optimising Statistical Machine Translation
Publication:Proceedings of the 12th Conference of the European Chapter of the ACL, Athens, Greece (2009) pp. 549-557
URL: http://www.aclweb.org/anthology/E09-1063
Status: Published

return to publications list