We freeze Large Language and Vision Assistant 1.6 (LLaVA-1.6), a multimodal large language model (LLM) with a Mistral-7B language backbone, together with five frozen audio encoders. Each encoder supplies either a raw token, formed by padding or truncating to 1,024 dimensions, or a projected token learned in contrastive language-image pretraining (CLIP) space, and the token budget k is the number of retained visual patch tokens. On a 1,006-clip evaluation pool, projection raises audio-to-video recall at one (R@1) for every encoder, including ImageBind from 0.10% to 15.7%, although contrastive language-audio pretraining (CLAP) and Whisper R@1 values are lower bounds for their original configurations. At k=150, the CLAP-based reference-free grounding score (B1) decreases from 0.113–0.132 to 0.062–0.088, a 23–53% relative reduction, and the CLAP result is read with source recall (B3) because the encoder and scorer share a checkpoint family. B3 decreases from 0.161–0.185 to 0.090–0.112, a 31–48% relative reduction, while caption-to-frame CLIP similarity (B2) decreases for four encoders and is unchanged at printed precision for Wav2CLIP. Across 3,568 disjoint paired clips, projection increases centered kernel alignment with visual embeddings in all five encoders, while the norm outside the visual principal subspace and nearest-neighbor overlap with the raw audio graph decrease. The noncausal comparisons show that raw-token bilingual evaluation understudy (BLEU) exceeds projected-token BLEU in nine of ten audio-aware-reference cells and that raw-over-projected generation scores predominate across five frozen backbones from 7B to 34B wherever raw-token generation is stable.
off-screen sound
projected token
generic caption
video match
raw token
source caption
SoundCLIP compares raw and projected tokens from AudioCLIP, Wav2CLIP, ImageBind, CLAP, and Whisper while keeping the audio encoders and LLaVA-1.6-Mistral-7B frozen. The raw token is padded or truncated to 1,024 dimensions. The projected token uses a three-layer MLP with layer normalization and GELU activations, trained to align the audio token with the paired CLIP visual embedding.
audio encoder
projected token
raw token
visual patches
frozen LLaVA
video retrieval
The evaluation pool is drawn from Audio-Visual Event 2 (AVE-2). The corrected local rebuild contains 570,138 three-second audio-video segments sourced from AudioSet. The paper uses AVE-2 for evaluation and geometry analysis. Its active-source fields support source recall (B3); machine-generated alignment scores are not used as ground truth.
Use the dataset documentation to inspect the metadata and media.
Inspect clip metadata and source annotations.
Follow the dataset documentation for the corresponding media files.
Explore visible and invisible active sound sources.
from datasets import load_dataset dataset = load_dataset("ali-vosoughi/ave-2")
Dataset Links:
These archived examples illustrate raw and projected captions from the earlier project page. They are not results from the current evaluation pool. Click an encoder mode to inspect an earlier generated caption; bracketed ellipses mark omitted passages.
The video takes place in a workshop setting, specifically at Finney's. The background features various tools and materials typically found in woodworking or similar crafts, such as wooden planks, a workbench, and shelves with additional items. There is also signage indicating "Finney's Expert Advice Simple Solutions Beautiful Work" on the wall. In the first scene, an individual wearing glasses and a black shirt with a white apron stands behind a workbench. They are holding a can of paint and appear to be preparing it for use. On the workbench, there is a rectangular piece of wood that looks like a cutting board, along with other tools and materials. A bottle labeled "Paint" is placed nearby, suggesting that painting might be part of their task. The second scene shows the same person continuing to prepare the paint can. This time, they are using their hands to apply the paint onto the cutting board. The environment remains consistent throughout, reinforcing the idea that this is a dedicated workspace for painting tasks. The third scene continues with the individual still engaged in painting. The focus remains on the process of applying the paint, which appears to be a light color given its sheen. The surrounding area includes more tools and materials, maintaining the workshop ambiance. The fourth scene captures another moment where the individual is seen working on the painting process. The lighting seems slightly dimmer than in previous scenes, possibly due to the nature of the work being done. The overall mood remains focused and industrious. The fifth scene returns to the individual standing behind the workbench, now holding a small container that likely contains additional paint or a different type of material. The arrangement of objects and characters suggests a continuation of the painting activity. The final scene shifts to a different angle, showing the individual from a side perspective. The lighting is brighter compared to earlier scenes, emphasizing the ongoing painting activity. The surroundings remain unchanged, reinforcing the continuity of the workshop environment. Overall, the video provides a detailed look into a typical day in a woodworking workshop, focusing on the activities involved in painting.
Labels with acoustic features: Deep, resonant, authoritative, and assertive -> Male speech, man speaking; Characterized by its loudness, pitch, and timbre -> Speech; Generally high frequency, short duration -> Animal
The video begins with a close-up of a person playing a red acoustic guitar on a wooden floor. The lighting is dim, creating shadows and highlighting the guitar's details. The background is dark, emphasizing the subject in focus. As the video progresses, the scene transitions to a wider shot showing the same person from behind, still holding the guitar. The lighting remains consistent, casting soft shadows around the person and the guitar. The setting appears to be indoors, possibly a room or studio, given the presence of what looks like a microphone stand in the background. In the final part of the video, the camera shifts slightly to show more of the person's hands as they play the guitar. The lighting continues to be dim, maintaining the mood of the earlier scenes. The person's movements are deliberate and focused, suggesting a practice session or performance. Throughout the video, there are no significant changes in the environment or actions apart from the progression of the shots. The overall atmosphere is one of concentration and musical expression.
Labels with acoustic features: Rich and varied -> Musical instrument; Bright and percussive -> Plucked string instrument; Full, bright, and complex -> Guitar; Rich and dynamic -> Music; Strong and resonant -> Piano; Bright and percussive -> Percussion; Warm and full-bodied -> Acoustic guitar; Bright and percussive -> Strum
The video begins with a view of an old, tall brick tower with pointed arches and decorative elements. […] In the foreground, there's a building with a sign that reads "PITI" in red letters against a white background. To the left side of the frame, part of another building can be seen, which seems to have a more modern appearance compared to the one in the foreground. The scene transitions smoothly into a festive atmosphere, indicated by garlands draped across the lower portion of the image. The sky remains overcast, suggesting it might be a cloudy day. As the video progresses, the focus shifts back to the tower, now showing its full height and detailed architecture. The bells remain visible, and the overall setting remains consistent throughout the video. Towards the end of the video, the camera angle changes slightly, revealing more of the surroundings, including additional buildings and trees. The lighting dims slightly, indicating either early morning or late afternoon light. Throughout the video, the visual elements include the brick tower, the bell towers, the greenery, and the architectural style of the surrounding structures. There are no significant movements apart from the slight change in perspective.
Audio caption: A church bell is ringing in a town square.
The video begins with a close-up view of a red stand mixer on a kitchen countertop. The mixer is in the process of mixing dough, which appears to be light yellow and slightly sticky. The background shows a granite countertop and some kitchen utensils, indicating that this scene takes place in a home kitchen. As the video progresses, the focus shifts to the dough being mixed by the mixer. The dough is thick and smooth, with visible air bubbles indicating it has been kneaded. The mixer's motor is spinning rapidly as the dough is being incorporated into the bowl. The lighting is bright, illuminating the dough and making its texture more visible. Towards the end of the video, the dough is fully mixed and ready for the next step in the recipe. The mixer continues to spin at high speed, ensuring the dough is well-mixed and evenly spread out. The background remains consistent throughout, showing the same kitchen countertop and utensils. Throughout the video, there are no significant changes or transitions between scenes; the setting and actions remain constant. The overall mood conveyed by the visual elements is one of preparation and cooking, suggesting that the person preparing the dough is likely following a recipe or instructions provided in the video.
Labels with acoustic features: High-pitched and whirring -> Power tool; Rich in frequency modulation -> Speech
The video begins with a close-up of a person's hands playing an electric guitar. The individual is wearing a black shirt and brown pants, seated on what appears to be a bench or chair. The background is out of focus, but it seems to be an indoor setting with artificial lighting. The scene transitions smoothly into the next frame, where the same person is seen holding the same electric guitar. This time, the guitar has a unique design that resembles a long-necked instrument, possibly a custom-built model known as a "Guitar Hero." The guitar features a prominent logo at the bottom that reads "Guitar Hero," indicating its brand identity. The person continues to play the guitar, adjusting their fingers on the fretboard and strumming the strings. In the final part of the video, the person is still engaged in playing the Guitar Hero guitar. They are shown from a slightly different angle, maintaining the same posture and focus on the guitar. The background remains consistent throughout this segment, reinforcing the continuity between scenes. Throughout the video, there are no significant changes in the environment or actions; the setting remains unchanged, and the main action revolves around the guitar player and their interaction with the instrument.
Labels with acoustic features: Rich in frequency modulation -> Speech; Bright and clear -> Harmonica; Smooth and resonant -> Guitar
The video begins with a close-up of a person's hand holding a black speaker. The background shows a wooden block and a green plastic container on a table, suggesting an indoor setting possibly in a kitchen or workshop. The scene transitions to the speaker being turned off, indicating that it is not currently playing music. Next, the focus shifts to a small circuit board placed on the same table as the speaker. A blue LED light is illuminated, which could be part of an electronic project or experiment. The circuit board has several wires connected to it, some of which are visible and appear to be part of the electronics setup. The video then moves forward to show the same circuit board now powered up by a battery, with the blue LED light still on. This suggests that the circuitry is functioning correctly and is likely part of a larger electronic device or system. In the final segment, the video continues with the circuit board now powered up, showing the blue LED light glowing brightly. The surrounding environment remains consistent with the previous scenes, maintaining continuity between the different stages of the project. Throughout the video, there are no significant changes in the actions or events occurring; instead, the sequence of images captures the progression from turning off the speaker to powering up the circuit board and finally illuminating the LED light.
Labels with acoustic features: Rich in frequency modulation -> Speech; Bright and sharp -> Alarm clock
The video begins with a person sitting on a chair, wrapped in a maroon blanket or towel. The background features a wooden floor and various plants, creating an indoor setting that appears cozy and lived-in. As the scene develops, the person starts to move around, bending down and adjusting their position slightly. They eventually sit up straight again, still wrapped in the same maroon item. In the next few frames, the person continues to adjust their posture while still wrapped in the blanket. This suggests they are either getting ready for bed or preparing to get out of bed. The room's decor remains consistent throughout, maintaining a warm and homely atmosphere. Towards the end of this segment, the person is seen standing up from the chair, holding the blanket over their head as if it were a hat. This action indicates they might be about to put it back on or have just taken it off. The final frame shows the person walking away from the chair, leaving the scene empty once more. Throughout the video, there are no significant changes in the environment or actions apart from the person's movements. The lighting is soft and natural, suggesting daytime or well-lit indoor conditions. There are no visible texts or subtitles within the video itself.
Labels with acoustic features: Rich in frequency modulation -> Speech; High-pitched and short -> Sneeze; Generally high pitched, sweet, and giggly -> Child speech, kid speaking; Loud, harsh, and discordant -> Noise; Chaotic and non-structured -> Background noise; Full range and varied -> Human voice; Richly complex -> Human voice
Within the tested configurations, projection raises audio-to-video retrieval and lowers audio grounding and source recall. The main comparison uses frozen LLaVA-1.6-Mistral-7B.
At k=150, B1 decreases by 23–53% and B3 decreases by 31–48% relative to raw tokens across all five encoders. B2 decreases for four encoders and is unchanged at printed precision for Wav2CLIP.
| Encoder | R@1 (%) | B1 | B2 | B3 |
|---|---|---|---|---|
| AudioCLIP | 0.00 → 1.19 | 0.125 → 0.088 | 0.219 → 0.204 | 0.161 → 0.111 |
| Wav2CLIP | 0.00 → 3.88 | 0.113 → 0.086 | 0.211 → 0.211 | 0.163 → 0.112 |
| ImageBind | 0.10 → 15.71 | 0.132 → 0.078 | 0.226 → 0.198 | 0.169 → 0.094 |
| CLAP | 0.10 → 1.99 | 0.129 → 0.071 | 0.225 → 0.199 | 0.185 → 0.107 |
| Whisper | 0.10 → 0.70 | 0.131 → 0.062 | 0.223 → 0.192 | 0.174 → 0.090 |
B1 is caption-to-audio CLAP similarity, B2 is caption-to-middle-frame CLIP similarity, and B3 is semantic recall of active sound sources. CLAP and Whisper R@1 values are lower bounds for their original encoder configurations. Interpret CLAP B1 with B3 because the encoder and scorer share a checkpoint family.
Across 3,568 disjoint paired clips, projection raises centered kernel alignment with visual embeddings and effective rank for every encoder. Orthogonal residual and nearest-neighbor overlap with the raw audio graph decrease. These comparisons do not identify a causal mechanism.
Raw-over-projected generation scores predominate across five frozen backbones from 7B to 34B wherever raw-token generation is stable. Raw-token BLEU exceeds projected-token BLEU in nine of ten audio-aware-reference cells. The study is limited to the tested encoders, backbones, and evaluation pool.
Read the current paper and explore the project repository and AVE-2 documentation.
If you use SoundCLIP or the AVE-2 dataset in your research, please cite our paper:
@unpublished{vosoughi2026soundclip,
title={{Projected Audio Tokens Gain Retrieval and Lose Grounded Generation in Multimodal LLMs}},
author={Vosoughi, Ali and Bi, Jing and Liu, Pinxin and Tang, Yolo Y. and Xu, Chenliang},
note={Preprint; submitted to ICASSP 2027},
year={2026},
url={https://ali-vosoughi.github.io/SoundCLIP/paper/main.pdf}
}