Projected Audio Tokens Gain Retrieval and Lose Grounded Generation in Multimodal LLMs

Preprint; submitted to ICASSP 2027
Computer Science Department, University of Rochester, NY, USA

Abstract

We freeze Large Language and Vision Assistant 1.6 (LLaVA-1.6), a multimodal large language model (LLM) with a Mistral-7B language backbone, together with five frozen audio encoders. Each encoder supplies either a raw token, formed by padding or truncating to 1,024 dimensions, or a projected token learned in contrastive language-image pretraining (CLIP) space, and the token budget k is the number of retained visual patch tokens. On a 1,006-clip evaluation pool, projection raises audio-to-video recall at one (R@1) for every encoder, including ImageBind from 0.10% to 15.7%, although contrastive language-audio pretraining (CLAP) and Whisper R@1 values are lower bounds for their original configurations. At k=150, the CLAP-based reference-free grounding score (B1) decreases from 0.113–0.132 to 0.062–0.088, a 23–53% relative reduction, and the CLAP result is read with source recall (B3) because the encoder and scorer share a checkpoint family. B3 decreases from 0.161–0.185 to 0.090–0.112, a 31–48% relative reduction, while caption-to-frame CLIP similarity (B2) decreases for four encoders and is unchanged at printed precision for Wav2CLIP. Across 3,568 disjoint paired clips, projection increases centered kernel alignment with visual embeddings in all five encoders, while the norm outside the visual principal subspace and nearest-neighbor overlap with the raw audio graph decrease. The noncausal comparisons show that raw-token bilingual evaluation understudy (BLEU) exceeds projected-token BLEU in nine of ten audio-aware-reference cells and that raw-over-projected generation scores predominate across five frozen backbones from 7B to 34B wherever raw-token generation is stable.

Controlled Audio-Token Substitution

Projected and raw audio-token paths for an off-screen sound source. off-screen sound projected token generic caption video match raw token source caption
The projected-token path retrieves the matching video but yields a generic caption, whereas the raw-token path misses the video but names the off-screen sound source.

SoundCLIP compares raw and projected tokens from AudioCLIP, Wav2CLIP, ImageBind, CLAP, and Whisper while keeping the audio encoders and LLaVA-1.6-Mistral-7B frozen. The raw token is padded or truncated to 1,024 dimensions. The projected token uses a three-layer MLP with layer normalization and GELU activations, trained to align the audio token with the paired CLIP visual embedding.

Frozen audio encoder, raw or projected token substitution, visual patch selection, and frozen LLaVA. audio encoder projected token raw token visual patches frozen LLaVA video retrieval
A frozen audio encoder yields a projected token or a 1,024-dimensional raw token. The token replaces the visual class token. At token budget k ∈ {15, 150}, the k most similar of 576 visual patch tokens condition frozen LLaVA-1.6. The token also supports audio-to-video retrieval.

AVE-2 Dataset

The evaluation pool is drawn from Audio-Visual Event 2 (AVE-2). The corrected local rebuild contains 570,138 three-second audio-video segments sourced from AudioSet. The paper uses AVE-2 for evaluation and geometry analysis. Its active-source fields support source recall (B3); machine-generated alignment scores are not used as ground truth.

570,138
Three-second Segments
1,006
Evaluation Clips
1,010
Generation Clip-segments
3,568
Disjoint Geometry Clips

Getting Started with AVE-2

Use the dataset documentation to inspect the metadata and media.

Load Metadata

Inspect clip metadata and source annotations.

Read Media Instructions

Follow the dataset documentation for the corresponding media files.

Inspect Source Fields

Explore visible and invisible active sound sources.

from datasets import load_dataset
dataset = load_dataset("ali-vosoughi/ave-2")

Dataset Links:

Interactive Demo

These archived examples illustrate raw and projected captions from the earlier project page. They are not results from the current evaluation pool. Click an encoder mode to inspect an earlier generated caption; bracketed ellipses mark omitted passages.

Example 1
Video ID: rvbmYs4Kl3Y (Segment 01)

Earlier Visual Description

The video takes place in a workshop setting, specifically at Finney's. The background features various tools and materials typically found in woodworking or similar crafts, such as wooden planks, a workbench, and shelves with additional items. There is also signage indicating "Finney's Expert Advice Simple Solutions Beautiful Work" on the wall. In the first scene, an individual wearing glasses and a black shirt with a white apron stands behind a workbench. They are holding a can of paint and appear to be preparing it for use. On the workbench, there is a rectangular piece of wood that looks like a cutting board, along with other tools and materials. A bottle labeled "Paint" is placed nearby, suggesting that painting might be part of their task. The second scene shows the same person continuing to prepare the paint can. This time, they are using their hands to apply the paint onto the cutting board. The environment remains consistent throughout, reinforcing the idea that this is a dedicated workspace for painting tasks. The third scene continues with the individual still engaged in painting. The focus remains on the process of applying the paint, which appears to be a light color given its sheen. The surrounding area includes more tools and materials, maintaining the workshop ambiance. The fourth scene captures another moment where the individual is seen working on the painting process. The lighting seems slightly dimmer than in previous scenes, possibly due to the nature of the work being done. The overall mood remains focused and industrious. The fifth scene returns to the individual standing behind the workbench, now holding a small container that likely contains additional paint or a different type of material. The arrangement of objects and characters suggests a continuation of the painting activity. The final scene shifts to a different angle, showing the individual from a side perspective. The lighting is brighter compared to earlier scenes, emphasizing the ongoing painting activity. The surroundings remain unchanged, reinforcing the continuity of the workshop environment. Overall, the video provides a detailed look into a typical day in a woodworking workshop, focusing on the activities involved in painting.

Earlier Audio Description

Labels with acoustic features: Deep, resonant, authoritative, and assertive -> Male speech, man speaking; Characterized by its loudness, pitch, and timbre -> Speech; Generally high frequency, short duration -> Animal

Sound Source Visibility Analysis

Earlier machine-generated source annotations:
👁️ Visible Active: Man, Cutting board, Paint canister
AUDIOCLIP
CLAP
WAV2CLIP
WHISPER
IMAGEBIND
Select an encoder mode to view generated caption
Click any button above to see the generated caption for this audiovisual example.
Example 2
Video ID: IwqD859w2_E (Segment 02)

Earlier Visual Description

The video begins with a close-up of a person playing a red acoustic guitar on a wooden floor. The lighting is dim, creating shadows and highlighting the guitar's details. The background is dark, emphasizing the subject in focus. As the video progresses, the scene transitions to a wider shot showing the same person from behind, still holding the guitar. The lighting remains consistent, casting soft shadows around the person and the guitar. The setting appears to be indoors, possibly a room or studio, given the presence of what looks like a microphone stand in the background. In the final part of the video, the camera shifts slightly to show more of the person's hands as they play the guitar. The lighting continues to be dim, maintaining the mood of the earlier scenes. The person's movements are deliberate and focused, suggesting a practice session or performance. Throughout the video, there are no significant changes in the environment or actions apart from the progression of the shots. The overall atmosphere is one of concentration and musical expression.

Earlier Audio Description

Labels with acoustic features: Rich and varied -> Musical instrument; Bright and percussive -> Plucked string instrument; Full, bright, and complex -> Guitar; Rich and dynamic -> Music; Strong and resonant -> Piano; Bright and percussive -> Percussion; Warm and full-bodied -> Acoustic guitar; Bright and percussive -> Strum

Sound Source Visibility Analysis

Earlier machine-generated source annotations:
👁️ Visible Active: Guitar
👁️ Visible Silent: Microphone stand
AUDIOCLIP
CLAP
WAV2CLIP
WHISPER
IMAGEBIND
Select an encoder mode to view generated caption
Click any button above to see the generated caption for this audiovisual example.
Example 3
Video ID: DDer7K8WG4I (Segment 02)

Earlier Visual Description

The video begins with a view of an old, tall brick tower with pointed arches and decorative elements. […] In the foreground, there's a building with a sign that reads "PITI" in red letters against a white background. To the left side of the frame, part of another building can be seen, which seems to have a more modern appearance compared to the one in the foreground. The scene transitions smoothly into a festive atmosphere, indicated by garlands draped across the lower portion of the image. The sky remains overcast, suggesting it might be a cloudy day. As the video progresses, the focus shifts back to the tower, now showing its full height and detailed architecture. The bells remain visible, and the overall setting remains consistent throughout the video. Towards the end of the video, the camera angle changes slightly, revealing more of the surroundings, including additional buildings and trees. The lighting dims slightly, indicating either early morning or late afternoon light. Throughout the video, the visual elements include the brick tower, the bell towers, the greenery, and the architectural style of the surrounding structures. There are no significant movements apart from the slight change in perspective.

Earlier Audio Description

Audio caption: A church bell is ringing in a town square.

Sound Source Visibility Analysis

Earlier machine-generated source annotations:
👁️ Visible Active: Church bells
👁️ Visible Silent: Tower, surrounding structures
AUDIOCLIP
CLAP
WAV2CLIP
WHISPER
IMAGEBIND
Select an encoder mode to view generated caption
Click any button above to see the generated caption for this audiovisual example.
Example 4
Video ID: UorSpZVnX_M (Segment 02)

Earlier Visual Description

The video begins with a close-up view of a red stand mixer on a kitchen countertop. The mixer is in the process of mixing dough, which appears to be light yellow and slightly sticky. The background shows a granite countertop and some kitchen utensils, indicating that this scene takes place in a home kitchen. As the video progresses, the focus shifts to the dough being mixed by the mixer. The dough is thick and smooth, with visible air bubbles indicating it has been kneaded. The mixer's motor is spinning rapidly as the dough is being incorporated into the bowl. The lighting is bright, illuminating the dough and making its texture more visible. Towards the end of the video, the dough is fully mixed and ready for the next step in the recipe. The mixer continues to spin at high speed, ensuring the dough is well-mixed and evenly spread out. The background remains consistent throughout, showing the same kitchen countertop and utensils. Throughout the video, there are no significant changes or transitions between scenes; the setting and actions remain constant. The overall mood conveyed by the visual elements is one of preparation and cooking, suggesting that the person preparing the dough is likely following a recipe or instructions provided in the video.

Earlier Audio Description

Labels with acoustic features: High-pitched and whirring -> Power tool; Rich in frequency modulation -> Speech

Sound Source Visibility Analysis

Earlier machine-generated source annotations:
👁️ Visible Active: mixer, speaker
👁️ Visible Silent: kitchen utensils, granite countertop
AUDIOCLIP
CLAP
WAV2CLIP
WHISPER
IMAGEBIND
Select an encoder mode to view generated caption
Click any button above to see the generated caption for this audiovisual example.
Example 5
Video ID: EQHrQIaQNv8 (Segment 03)

Earlier Visual Description

The video begins with a close-up of a person's hands playing an electric guitar. The individual is wearing a black shirt and brown pants, seated on what appears to be a bench or chair. The background is out of focus, but it seems to be an indoor setting with artificial lighting. The scene transitions smoothly into the next frame, where the same person is seen holding the same electric guitar. This time, the guitar has a unique design that resembles a long-necked instrument, possibly a custom-built model known as a "Guitar Hero." The guitar features a prominent logo at the bottom that reads "Guitar Hero," indicating its brand identity. The person continues to play the guitar, adjusting their fingers on the fretboard and strumming the strings. In the final part of the video, the person is still engaged in playing the Guitar Hero guitar. They are shown from a slightly different angle, maintaining the same posture and focus on the guitar. The background remains consistent throughout this segment, reinforcing the continuity between scenes. Throughout the video, there are no significant changes in the environment or actions; the setting remains unchanged, and the main action revolves around the guitar player and their interaction with the instrument.

Earlier Audio Description

Labels with acoustic features: Rich in frequency modulation -> Speech; Bright and clear -> Harmonica; Smooth and resonant -> Guitar

Sound Source Visibility Analysis

Earlier machine-generated source annotations:
👁️ Visible Active: Guitar
👂 Invisible Active: Harmonica
AUDIOCLIP
CLAP
WAV2CLIP
WHISPER
IMAGEBIND
Select an encoder mode to view generated caption
Click any button above to see the generated caption for this audiovisual example.
Example 6
Video ID: GJhYkfI7jpU (Segment 03)

Earlier Visual Description

The video begins with a close-up of a person's hand holding a black speaker. The background shows a wooden block and a green plastic container on a table, suggesting an indoor setting possibly in a kitchen or workshop. The scene transitions to the speaker being turned off, indicating that it is not currently playing music. Next, the focus shifts to a small circuit board placed on the same table as the speaker. A blue LED light is illuminated, which could be part of an electronic project or experiment. The circuit board has several wires connected to it, some of which are visible and appear to be part of the electronics setup. The video then moves forward to show the same circuit board now powered up by a battery, with the blue LED light still on. This suggests that the circuitry is functioning correctly and is likely part of a larger electronic device or system. In the final segment, the video continues with the circuit board now powered up, showing the blue LED light glowing brightly. The surrounding environment remains consistent with the previous scenes, maintaining continuity between the different stages of the project. Throughout the video, there are no significant changes in the actions or events occurring; instead, the sequence of images captures the progression from turning off the speaker to powering up the circuit board and finally illuminating the LED light.

Earlier Audio Description

Labels with acoustic features: Rich in frequency modulation -> Speech; Bright and sharp -> Alarm clock

Sound Source Visibility Analysis

Earlier machine-generated source annotations:
👁️ Visible Active: Speaker
👁️ Visible Silent: Green plastic container
AUDIOCLIP
CLAP
WAV2CLIP
WHISPER
IMAGEBIND
Select an encoder mode to view generated caption
Click any button above to see the generated caption for this audiovisual example.
Example 7
Video ID: AVL7Kbpw13U (Segment 01)

Earlier Visual Description

The video begins with a person sitting on a chair, wrapped in a maroon blanket or towel. The background features a wooden floor and various plants, creating an indoor setting that appears cozy and lived-in. As the scene develops, the person starts to move around, bending down and adjusting their position slightly. They eventually sit up straight again, still wrapped in the same maroon item. In the next few frames, the person continues to adjust their posture while still wrapped in the blanket. This suggests they are either getting ready for bed or preparing to get out of bed. The room's decor remains consistent throughout, maintaining a warm and homely atmosphere. Towards the end of this segment, the person is seen standing up from the chair, holding the blanket over their head as if it were a hat. This action indicates they might be about to put it back on or have just taken it off. The final frame shows the person walking away from the chair, leaving the scene empty once more. Throughout the video, there are no significant changes in the environment or actions apart from the person's movements. The lighting is soft and natural, suggesting daytime or well-lit indoor conditions. There are no visible texts or subtitles within the video itself.

Earlier Audio Description

Labels with acoustic features: Rich in frequency modulation -> Speech; High-pitched and short -> Sneeze; Generally high pitched, sweet, and giggly -> Child speech, kid speaking; Loud, harsh, and discordant -> Noise; Chaotic and non-structured -> Background noise; Full range and varied -> Human voice; Richly complex -> Human voice

Sound Source Visibility Analysis

Earlier machine-generated source annotations:
👁️ Visible Active: Person
👂 Invisible Active: Sneeze sound
AUDIOCLIP
CLAP
WAV2CLIP
WHISPER
IMAGEBIND
Select an encoder mode to view generated caption
Click any button above to see the generated caption for this audiovisual example.

Key Findings

Within the tested configurations, projection raises audio-to-video retrieval and lowers audio grounding and source recall. The main comparison uses frozen LLaVA-1.6-Mistral-7B.

Retrieval and Grounded Generation

At k=150, B1 decreases by 23–53% and B3 decreases by 31–48% relative to raw tokens across all five encoders. B2 decreases for four encoders and is unchanged at printed precision for Wav2CLIP.

Raw → projected tokens at k=150. R@1 (%) uses 1,006 evaluation clips; B1, B2, and B3 use 1,010 clip-segments.
EncoderR@1 (%)B1B2B3
AudioCLIP0.00 → 1.190.125 → 0.0880.219 → 0.2040.161 → 0.111
Wav2CLIP0.00 → 3.880.113 → 0.0860.211 → 0.2110.163 → 0.112
ImageBind0.10 → 15.710.132 → 0.0780.226 → 0.1980.169 → 0.094
CLAP0.10 → 1.990.129 → 0.0710.225 → 0.1990.185 → 0.107
Whisper0.10 → 0.700.131 → 0.0620.223 → 0.1920.174 → 0.090

B1 is caption-to-audio CLAP similarity, B2 is caption-to-middle-frame CLIP similarity, and B3 is semantic recall of active sound sources. CLAP and Whisper R@1 values are lower bounds for their original encoder configurations. Interpret CLAP B1 with B3 because the encoder and scorer share a checkpoint family.

Representation Geometry

Across 3,568 disjoint paired clips, projection raises centered kernel alignment with visual embeddings and effective rank for every encoder. Orthogonal residual and nearest-neighbor overlap with the raw audio graph decrease. These comparisons do not identify a causal mechanism.

Scope and Backbone Stability

Raw-over-projected generation scores predominate across five frozen backbones from 7B to 34B wherever raw-token generation is stable. Raw-token BLEU exceeds projected-token BLEU in nine of ten audio-aware-reference cells. The study is limited to the tested encoders, backbones, and evaluation pool.

Code and Resources

Read the current paper and explore the project repository and AVE-2 documentation.

Citation

If you use SoundCLIP or the AVE-2 dataset in your research, please cite our paper:

@unpublished{vosoughi2026soundclip,
  title={{Projected Audio Tokens Gain Retrieval and Lose Grounded Generation in Multimodal LLMs}},
  author={Vosoughi, Ali and Bi, Jing and Liu, Pinxin and Tang, Yolo Y. and Xu, Chenliang},
  note={Preprint; submitted to ICASSP 2027},
  year={2026},
  url={https://ali-vosoughi.github.io/SoundCLIP/paper/main.pdf}
}