01 TL;DR
The model knows the premise is false.
It almost never says so.
We ask omnimodal LLMs about movie scenes and change one detail in the question. A linear probe on their hidden states finds the false premise up to 86% of the time. In their answers, seven of eight open models reject it at most 16.2% of the time for vision and 6.6% for audio.
Behaviour: 7 of 8 open-source omnimodal LLMs, fixed option order. Probe: best model, Baichuan-Omni-1.5 excluded.
The bottleneck is acting on the encoded mismatch signal, not encoding it.
02 Overview video
03 Abstract
When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sensory input. We introduce IMAVB, a curated benchmark of 500 multi-minute movie clips with a 2×2 design crossing target modality (vision, audio) and premise condition (standard, misleading), which lets us measure conflict detection separately from ordinary multimodal comprehension. Across eight open-source omnimodal LLMs and Gemini 3.1 Pro, we document a Representation–Action Gap: hidden states encode premise–perception mismatches reliably for vision and more weakly for audio, even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions, sacrificing ordinary comprehension accuracy. The gap is modality-asymmetric (models reject false audio premises far less often than false visual ones) and prompt-resistant across seven variants. As an initial diagnostic intervention, a probe-guided logit adjustment (PGLA) re-injects the encoded mismatch signal into decoding and consistently improves rejection behavior at a bounded cost to standard accuracy. Together, these results suggest the bottleneck for omnimodal grounding lies in acting on the encoded mismatch signal rather than in encoding it.
04 Try it yourself
One clip, four questions.
One word changes.
Watch the excerpt with sound, then pick a cell of the 2×2. Misleading questions copy the standard question and swap exactly one premise detail. The correct answer then becomes E or F.
Excerpts from IMAVB clips (dataset: ngqtrung/IMAVB). Each excerpt starts at the question's 10-second answer window. Click an option to guess.
05 The benchmark
IMAVB: long movie scenes,
with every sense left on.
Video and audio stay unedited. Only the textual question varies. This lets us measure conflict detection separately from ordinary comprehension, while both modalities stay present and compete.
- Movie clips
- 500
- Hours of video
- 20.7
- Clip length (median 140 s)
- 60–300 s
- Questions
- 2,000
- Question categories
- 8
- Misleading subcategories (9 V · 9 A)
- 18
- Items verified by hand
- 100%
- Independent agreement on misleading premises
- 89%
The 2×2 design
Every clip gets four questions. The rows pick the target modality. The columns pick the premise. A misleading question is the standard question with exactly one premise detail swapped, for example “maroon shirt” to “blue shirt”.
Options E (“The visual detail in the question is incorrect”) and F (“The audio detail in the question is incorrect”) appear in every prompt. Standard answers are A–D.
- existence
- time order
- emotional
- scene description
- cross modality
- plot
- causal
- temporal
“…the younger man wearing a maroon polo shirt…”
answer A–D“…the younger man wearing a blue polo shirt…”
answer E“…while a delicate string melody swells…”
answer A–D“…while a loud drum beat swells…”
answer F
We verified all 2,000 items by hand, watching each clip with audio. An independent annotator re-rated 200 items from 50 videos: 89% agreement on misleading-premise judgments, and within one scale point on 94–96% of items per criterion.
06 Results
Two ways to fail:
under- and over-rejection.
Under-rejection · 7 of 8 open models
They answer 40–75% of standard questions, yet reject misleading ones in only ≤16.2% (vision) and ≤6.6% (audio) of cases. Four models score 0% on misleading audio.
Over-rejection · Qwen3-Omni, Gemini 3.1 Pro
They reject more often (72.8% and 94.0% on vision), but they trade 15–25 pp of standard accuracy.
Baseline accuracy, fixed option order (%)
Standard · vision Standard · audio Misleading · vision Misleading · audio
Bal = ½((std_v + std_a)/2 + (mis_v + mis_a)/2). Human: 1 volunteer, 50-video subset, unaware of the design. Hover or tap a bar for its value.
- Modality-asymmetric. Models reject false audio premises far less often than false visual ones.
- Prompt-resistant. The gap survives seven prompt variants (A1–A7), option shuffling, and stratification by video length.
- Solvable. A volunteer who did not know the design detected 78% of the false premises on a 50-video subset.
07 The Representation–Action Gap
Inside, the signal is there.
Knows vs. says, model by model.
We train linear probes on each model's hidden states to tell a misleading question from its standard twin. The violet dot is what the probe recovers. The red dot is how often the model rejects the false premise in its answer.
HS probe Residualized probe Behaviour (mis) Text-only baselines
† Baichuan-Omni-1.5 is excluded from aggregates because its layer-2 signal is near ceiling. Dashed lines: text-only TF-IDF and SBERT classifiers on the question text alone.
Text confound control
The two questions differ by one swapped detail, so text alone is partly predictive. TF-IDF reaches 73.4% (vision) and 71.4% (audio). SBERT reaches 66.8% and 57.3%.
Residualization
We project out text-predictive features from the hidden states (Ridge regression with orthogonal projection) and retrain the probes. They stay above SBERT for every model and modality, and above TF-IDF for vision. The audio evidence is weaker.
Runner-up analysis
For the 7 under-rejecting models, the correct reject option is often the runner-up (25.8% vision, 20.2% audio). A residue of the signal reaches the output logits.
08 A diagnostic intervention
Feed the signal back,
and the behaviour moves.
If the encoded signal is only an artifact of probing, feeding it back to the output should do nothing. Probe-guided logit adjustment (PGLA) tests this. A two-layer MLP probe reads the hidden state at the peak layer. Its confidence g = P_{\text{mis}}^{\,p} gates a boost on the reject options E and F.
Δ is the gap between the best content logit (A–D) and the best reject logit (E, F). β corrects the E/F positional bias. PGLA needs no input perturbation and no second forward pass. Results use 5-fold cross-validation.
ΔBal per model (pp)
Baseline Bal → Bal after PGLA. PGLA improves all 8 models.
09 Cross-modal interference
Does the other sense get in the way?
We run each open model with one modality removed and measure the change in misleading accuracy. A→V removes audio and reads mis_v. V→A removes video and reads mis_a. Positive values mean the removed modality was interfering.
The effect depends on the architecture. Audio interferes with visual detection for three models. MiniCPM-o 2.6 is the only model where joint audio-visual input helps. For Qwen3-Omni, video interferes with audio detection.
pp change in misleading accuracy when the other modality is removed.
10 Citation
@inproceedings{nguyen2026senses,
title = {Senses Wide Shut: A Representation--Action Gap in Omnimodal {LLM}s},
author = {Nguyen Quang, Trung and Gao, Yiming and Pu, Fanyi and Zhang, Kaichen and Sun, Shuo and Liu, Ziwei},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}
Acknowledgements
This study is supported by the Ministry of Education, Singapore, under its MOE AcRF Tier 2 (MOE-T2EP20223-0002). This research is also supported by cash and in-kind funding from NTU S-Lab and industry partner(s). We wish to acknowledge the support of Nanyang Technological University through the URECA Undergraduate Research Programme.