IMAVB · 500 CLIPS · 2 × 2 NeurIPS 2026
NeurIPS 2026

Senses Wide Shut

A Representation–Action Gap in Omnimodal LLMs

Nguyen Quang Trung1,2* Yiming Gao1,2* Fanyi Pu1,2 Kaichen Zhang1,2 Shuo Sun3 Ziwei Liu1,2†

1Nanyang Technological University 2LMMs-Lab Team 3Johns Hopkins University

* Equal contribution    † Corresponding author

01 TL;DR

The model knows the premise is false.
It almost never says so.

We ask omnimodal LLMs about movie scenes and change one detail in the question. A linear probe on their hidden states finds the false premise up to 86% of the time. In their answers, seven of eight open models reject it at most 16.2% of the time for vision and 6.6% for audio.

The bottleneck is acting on the encoded mismatch signal, not encoding it.

02 Overview video

03 Abstract

When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sensory input. We introduce IMAVB, a curated benchmark of 500 multi-minute movie clips with a 2×2 design crossing target modality (vision, audio) and premise condition (standard, misleading), which lets us measure conflict detection separately from ordinary multimodal comprehension. Across eight open-source omnimodal LLMs and Gemini 3.1 Pro, we document a Representation–Action Gap: hidden states encode premise–perception mismatches reliably for vision and more weakly for audio, even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions, sacrificing ordinary comprehension accuracy. The gap is modality-asymmetric (models reject false audio premises far less often than false visual ones) and prompt-resistant across seven variants. As an initial diagnostic intervention, a probe-guided logit adjustment (PGLA) re-injects the encoded mismatch signal into decoding and consistently improves rejection behavior at a bounded cost to standard accuracy. Together, these results suggest the bottleneck for omnimodal grounding lies in acting on the encoded mismatch signal rather than in encoding it.

Overview figure. Left: a video of a cat chasing a red ball with soft guitar music. Middle: the model's internal representation correctly encodes the cat and the guitar music. Right: text prompts with a misleading premise (a dog, or a loud drum beat). The correct action is to flag the wrong detail, but the model answers the question as if the premise were true.
Figure 1. Overview of the Representation–Action Gap on IMAVB. The model encodes what it saw and heard, yet it answers a question built on a false detail as if the detail were true.

04 Try it yourself

One clip, four questions.
One word changes.

Watch the excerpt with sound, then pick a cell of the 2×2. Misleading questions copy the standard question and swap exactly one premise detail. The correct answer then becomes E or F.

Standard Misleading Vision Audio

    Excerpts from IMAVB clips (dataset: ngqtrung/IMAVB). Each excerpt starts at the question's 10-second answer window. Click an option to guess.

    05 The benchmark

    IMAVB: long movie scenes,
    with every sense left on.

    Video and audio stay unedited. Only the textual question varies. This lets us measure conflict detection separately from ordinary comprehension, while both modalities stay present and compete.

    Movie clips
    500
    Hours of video
    20.7
    Clip length (median 140 s)
    60–300 s
    Questions
    2,000
    Question categories
    8
    Misleading subcategories (9 V · 9 A)
    18
    Items verified by hand
    100%
    Independent agreement on misleading premises
    89%

    The 2×2 design

    Every clip gets four questions. The rows pick the target modality. The columns pick the premise. A misleading question is the standard question with exactly one premise detail swapped, for example “maroon shirt” to “blue shirt”.

    Options E (“The visual detail in the question is incorrect”) and F (“The audio detail in the question is incorrect”) appear in every prompt. Standard answers are A–D.

    • existence
    • time order
    • emotional
    • scene description
    • cross modality
    • plot
    • causal
    • temporal

    We verified all 2,000 items by hand, watching each clip with audio. An independent annotator re-rated 200 items from 50 videos: 89% agreement on misleading-premise judgments, and within one scale point on 94–96% of items per criterion.

    06 Results

    Two ways to fail:
    under- and over-rejection.

    Under-rejection · 7 of 8 open models

    They answer 40–75% of standard questions, yet reject misleading ones in only ≤16.2% (vision) and ≤6.6% (audio) of cases. Four models score 0% on misleading audio.

    Over-rejection · Qwen3-Omni, Gemini 3.1 Pro

    They reject more often (72.8% and 94.0% on vision), but they trade 15–25 pp of standard accuracy.

    Baseline accuracy, fixed option order (%)

    Standard · vision Standard · audio Misleading · vision Misleading · audio

    Bal = ½((std_v + std_a)/2 + (mis_v + mis_a)/2). Human: 1 volunteer, 50-video subset, unaware of the design. Hover or tap a bar for its value.

    • Modality-asymmetric. Models reject false audio premises far less often than false visual ones.
    • Prompt-resistant. The gap survives seven prompt variants (A1–A7), option shuffling, and stratification by video length.
    • Solvable. A volunteer who did not know the design detected 78% of the false premises on a 50-video subset.

    07 The Representation–Action Gap

    Inside, the signal is there.
    Knows vs. says, model by model.

    We train linear probes on each model's hidden states to tell a misleading question from its standard twin. The violet dot is what the probe recovers. The red dot is how often the model rejects the false premise in its answer.

    HS probe Residualized probe Behaviour (mis) Text-only baselines

    † Baichuan-Omni-1.5 is excluded from aggregates because its layer-2 signal is near ceiling. Dashed lines: text-only TF-IDF and SBERT classifiers on the question text alone.

    Text confound control

    The two questions differ by one swapped detail, so text alone is partly predictive. TF-IDF reaches 73.4% (vision) and 71.4% (audio). SBERT reaches 66.8% and 57.3%.

    Residualization

    We project out text-predictive features from the hidden states (Ridge regression with orthogonal projection) and retrain the probes. They stay above SBERT for every model and modality, and above TF-IDF for vision. The audio evidence is weaker.

    Runner-up analysis

    For the 7 under-rejecting models, the correct reject option is often the runner-up (25.8% vision, 20.2% audio). A residue of the signal reaches the output logits.

    08 A diagnostic intervention

    Feed the signal back,
    and the behaviour moves.

    If the encoded signal is only an artifact of probing, feeding it back to the output should do nothing. Probe-guided logit adjustment (PGLA) tests this. A two-layer MLP probe reads the hidden state at the peak layer. Its confidence g = P_{\text{mis}}^{\,p} gates a boost on the reject options E and F.

    \begin{aligned} L'_E &= L_E + \sigma\!\bigl(\gamma(g-\alpha)\bigr)\cdot(s\cdot\Delta+\delta) - \tfrac{\beta}{2} \\[2pt] L'_F &= L_F + \sigma\!\bigl(\gamma(g-\alpha)\bigr)\cdot(s\cdot\Delta+\delta) + \tfrac{\beta}{2} \end{aligned}

    Δ is the gap between the best content logit (A–D) and the best reject logit (E, F). β corrects the E/F positional bias. PGLA needs no input perturbation and no second forward pass. Results use 5-fold cross-validation.

    +15.0pp mean ΔBal, unconstrained
    +9.9pp at a 3 pp standard-accuracy budget
    +5.4pp at a 1 pp budget

    ΔBal per model (pp)

    Baseline Bal → Bal after PGLA. PGLA improves all 8 models.

    09 Cross-modal interference

    Does the other sense get in the way?

    We run each open model with one modality removed and measure the change in misleading accuracy. A→V removes audio and reads mis_v. V→A removes video and reads mis_a. Positive values mean the removed modality was interfering.

    The effect depends on the architecture. Audio interferes with visual detection for three models. MiniCPM-o 2.6 is the only model where joint audio-visual input helps. For Qwen3-Omni, video interferes with audio detection.

    pp change in misleading accuracy when the other modality is removed.

    10 Citation

    BibTeX
    @inproceedings{nguyen2026senses,
      title     = {Senses Wide Shut: A Representation--Action Gap in Omnimodal {LLM}s},
      author    = {Nguyen Quang, Trung and Gao, Yiming and Pu, Fanyi and Zhang, Kaichen and Sun, Shuo and Liu, Ziwei},
      booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
      year      = {2026}
    }

    Acknowledgements

    This study is supported by the Ministry of Education, Singapore, under its MOE AcRF Tier 2 (MOE-T2EP20223-0002). This research is also supported by cash and in-kind funding from NTU S-Lab and industry partner(s). We wish to acknowledge the support of Nanyang Technological University through the URECA Undergraduate Research Programme.