Submitted to ICASSP 2027 ยท submission

Projected Audio Tokens Gain Retrieval and Lose Grounded Generation in Multimodal LLMs

Schematic: projected audio tokens in a multimodal LLM, retrieval versus grounded generation

Schematic illustration of the idea, not a figure from the paper.

Ali Vosoughi, Jing Bi, Pinxin Liu, Yolo Y. Tang, Chenliang Xu

SoundCLIP: audio tokens projected into a LLaVA-style multimodal LLM gain retrieval and lose grounded generation; built on the AVE-2 dataset.

My part. Explored SoundCLIP and WhisperCLIP-style token-substitution methods that exposed a practical retrieval-versus-generation trade-off.

SoundCLIP: audio tokens projected into a LLaVA-style multimodal LLM gain retrieval and lose grounded generation; built on the AVE-2 dataset.

Under review. Listed as a submission; nothing here claims acceptance.

All publications