Code-Switching DPO

Interspeech 2026

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

  • Nguyen Quang Trung1,2
  • Cheng Yi Lewis Won1
  • Minh Duc Pham1
  • Yingxu He1
  • Shuo Sun1
  • Ai Ti Aw1
  • 1Institute for Infocomm Research (I2R), A*STAR, Singapore
  • 2Nanyang Technological University, Singapore

Corresponding author

TL;DR

Multilingual Audio LLMs often fail on English-Mandarin code-switching speech: they drop one language, translate instead of transcribe, or hallucinate. We train three Audio LLMs with Direct Preference Optimization (DPO) on pairs in which the chosen response keeps the mixed-language transcription and the rejected response imitates a failure. DPO shifts the models toward verbatim mixed-language output, with relative MER reductions of up to 89.6% in-distribution and 20.0% out-of-distribution. We read this as DPO eliciting a capability that the models already have.

preference pairs
100,766
hours of audio
566.8
Audio LLMs
3
benchmarks
4

Abstract

Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language omission, translation-instead-of-transcription, and hallucination. We apply Direct Preference Optimization (DPO) to align models, constructing preference pairs in which chosen responses preserve mixed-language content while rejected responses mimic failure patterns. Training three Audio LLMs on 100K pairs (570 hours), we observe consistent behavioral shifts: models learn to preserve language composition rather than translating when prompted for transcription. This alignment yields MER reductions up to 89.6% (in-distribution) and 20.0% (out-of-distribution). Our findings suggest DPO can effectively elicit correct code-switching transcription behavior from multilingual Audio LLMs.

The problem

Three failure modes

When we ask for a transcription of code-switching audio, even models trained with code-switching data fail in three systematic ways. The examples below are the ones in the paper.

1

Language omission

The model outputs only one language and drops the other.

GT我住 temasek poly 那边 Out我住那边

2

Translation instead of transcription

The model translates the mixed-language content into one language.

GT我们都应该 pursue a healthy lifestyle Out我们都应该追求健康的生活方式

3

Hallucination

The model generates repeated or fabricated content.

GT我们二月多有 valentine's day Outah month... (×250)

Method

Preference pairs from ground truth

The chosen response is the ground-truth code-switching transcription. Qwen3-32B rewrites the same transcription into a rejected response that imitates a failure. Vanilla DPO then raises the likelihood of the chosen response and lowers that of the rejected one, relative to the frozen base model.

Pipeline: code-switching audio and its ground-truth transcription. The ground truth becomes the chosen response; Qwen3-32B turns it into a rejected response by global or partial translation. The audio, chosen and rejected responses form a preference pair, which DPO training uses to produce an aligned Audio LLM.
Overview of DPO training for code-switching alignment. Ground-truth transcriptions serve as chosen responses, while an LLM generates rejected responses that mimic failure modes via Global Translation (full) and Partial Translation (spans only).
1

Global translation (80%)

All content of one language is translated into the other, which imitates translation-instead-of-transcription.
What grade are you? 真的很好哎… → 你几年级?真的很好哎…

2

Partial translation (20%)

Only short spans are translated, which imitates partial language omission.
It's so boring and dull → It's so 无聊 and dull

3

DPO training

One epoch on 8 H100 GPUs with 20 English and 20 Chinese transcription prompts. Evaluation uses a held-out prompt, “Please transcribe this speech.”

Both strategies target translation failures only. We do not generate rejected samples for omission of content or for hallucination, but DPO still reduces all three failure modes.

Results

Consistent gains on all models and benchmarks

We train MERaLiON-2-3B, Phi-4-multimodal-instruct and Qwen2-Audio-7B-Instruct. SEAME dev_sge and dev_man are out-of-distribution, because no SEAME data is in training. EMILIA-test and CS-Dialogue-test are held-out portions of the training sources.

modelbenchmarkBaseDPOΔRel
MERaLiON-2-3BSEAME dev_sge32.3831.75−2.0%
SEAME dev_man25.7925.61−0.7%
EMILIA32.0130.41−5.0%
CS-Dialogue25.4122.58−11.1%
Phi-4-multimodal-instructSEAME dev_sge69.9761.09−12.7%
SEAME dev_man51.9746.63−10.3%
EMILIA70.987.38−89.6%
CS-Dialogue49.6110.65−78.5%
Qwen2-Audio-7B-InstructSEAME dev_sge95.1185.52−10.1%
SEAME dev_man72.8958.30−20.0%
EMILIA44.7042.08−5.9%
CS-Dialogue38.9131.40−19.3%

Mixed Error Rate (MER, %, lower is better): character level for Chinese, word level for English. ΔRel is the relative change from Base to DPO. Each number comes from a single training run.

MERaLiON-2-3B

Small SEAME gains (0.7–2.0%), because the model already saw extensive code-switching data in supervised fine-tuning. CS-Dialogue still improves by 11.1%.

Phi-4-multimodal-instruct

The base model often translates or repeats itself. One epoch of DPO lowers MER on EMILIA from 70.98% to 7.38%.

Qwen2-Audio-7B-Instruct

A 20.0% relative gain on SEAME dev_man, the largest out-of-distribution gain, and 5.9–19.3% in-distribution. Trained with LoRA.

Qualitative examples

What DPO changes

After DPO, the models keep the mixed-language pattern and produce more stable transcriptions. One example for each failure mode, from the paper.

Translation · Qwen2-Audio-7B-Instruct

GT我们都应该 pursue a healthy lifestyle Base我们都应该追求健康的生活方式 DPO我们都应该 pursue a healthy lifestyle MER 100% → 0%

Hallucination · Phi-4-multimodal-instruct

GT我们二月多有 valentine's day Baseah month ah month ah month... (×250) DPO二月多有 Valentine's Day MER 56.89% → 0.33%

Language omission · MERaLiON-2-3B

GT我住 temasek poly 那边 Base我住达马士科波利那边 (transliterated) DPO我住 tamasek poly 那边 MER 100% → 17%

Data

Released preference datasets

The paper trains on 100,766 pairs (566.8 hours): 13,795 pairs (77.3 hours) from CS-Dialogue and 86,971 pairs (489.5 hours) from EMILIA. CS-Dialogue gives natural intra-sentential switches and concatenated EN and CN utterances from the same conversation. EMILIA gives synthetic inter-sentential switches from concatenated English and Chinese clips.

datasettraintest
CS-Dialogue-DPO13,795359
Emilia-DPO86,9711,000

Hugging Face splits. Each row has four columns: id, audio, chosen (the ground-truth transcription) and rejected (the failure-style rewrite).

from datasets import load_dataset

ds = load_dataset("ngqtrung/cs-dialogue-dpo")
print(ds["train"][0]["chosen"])

The audio derives from CS-Dialogue (Zhou et al., 2025) and Emilia (He et al., SLT 2024). Users must follow the licence terms of the original datasets.

Citation

BibTeX

bibtex
@inproceedings{nguyen2026codeswitchdpo,
  title     = {Direct Preference Optimization for {English-Mandarin} Code-Switching Speech Recognition in Audio {LLMs}},
  author    = {Nguyen Quang, Trung and Won, Cheng Yi Lewis and Pham, Minh Duc and He, Yingxu and Sun, Shuo and Aw, Ai Ti},
  booktitle = {Proc. Interspeech 2026},
  year      = {2026},
  doi       = {10.21437/Interspeech.2026-110}
}