NAVER LABS Europe submission to the Instruction-following Track

Published by Ioan Calapodescu at 31 July 2025

Beomseok Lee, Marcely Zanon Boito, Laurent Besacier, Ioan Calapodescu

The International Conference on Spoken Language Translation (IWSLT), Vienna, Austria, 31 July - 1 August, 2025

In this paper we describe NAVER LABS Europe submission to the instruction-following speech processing short track at IWSLT 2025. We participate in the constrained settings, developing systems that can simultaneously perform ASR, ST, and SQA tasks from English speech input into the following target languages: Chinese, Italian, and German. Our solution leverages two pretrained modules: (1) a speech-to-LLM embedding projector trained using representations from the SeamlessM4T-v2-large speech encoder; and (2) LoRA adapters trained on text data on top of Llama-3.1-8B-Instruct. These modules are jointly loaded and further instruction-tuned for 1K steps on multilingual and multimodal data to form our final system submitted for evaluation.

INTERACTION

Equip robots to interact safely with humans, other robots and systems.

VISION

Perception to help robots understand and interact with the environment.

ACTION

Providing embodied agents with sequential decision-making capabilities to safely execute complex tasks in dynamic environments.

NAVER FRANCE Gender Equality 2025

All

Publications

Blog

News

Code & Data

Careers

People

NAVER LABS Europe submission to the Instruction-following Track

All

Publications

Blog

News

Code & Data

Careers

People

Cookie settings