On the use of BERT for neural machine translation

Published by Stéphane Clinchant at 25 September 2019

Stéphane Clinchant, Kweon Woo Jung, Vassilina Nikoulina

3rd Workshop on Neural Generation and Translation (WNGT 2019), EMNLP, Hong Kong, China, 3-7 November, 2019

@article{clinchant2019use,
  title={On the use of BERT for Neural Machine Translation},
  author={Clinchant, St{\'e}phane and Jung, Kweon Woo and Nikoulina, Vassilina},
  journal={arXiv preprint arXiv:1909.12744},
  year={2019}
}

Careers home

Abstract

Exploiting large pretrained models for various NMT tasks have gained a lot of visibility recently. In this work we study how BERT pretrained models could be exploited for supervised Neural Machine Translation. We compare various ways to integrate pretrained BERT model with NMT model and study the impact of the monolingual data used for BERT training on the final translation quality. We use WMT-14 English-German, IWSLT15 English-German and IWSLT14 English-Russian datasets for these experiments. In addition to standard task test set evaluation, we perform evaluation on out-of-domain test sets and noise injected test sets, in order to assess how BERT pretrained representations affect model robustness.

All

Publications

Blog

News

Code & Data

Careers

People

NAVER FRANCE Gender Equality 2024

NAVER FRANCE Gender Equality 2023

VISION

Perception to help robots understand and interact with the environment.

INTERACTION

Equip robots to interact safely with humans, other robots and systems.

ACTION

Providing embodied agents with sequential decision-making capabilities to safely execute complex tasks in dynamic environments.

Action

On the use of BERT for neural machine translation

All

Publications

Blog

News

Code & Data

Careers

People

Cookie settings