Microsoft Research · Research Intern · 2024

AVE-2: a 570K-clip audio-visual dataset and its curation stack (Microsoft Research)

Built the weak-supervision pipeline that scores 570K+ clips for temporal alignment, physical causality and source visibility with LLM/VLM scorers, and released the dataset on Hugging Face.

  • Scale 570,138 clips
  • Evaluation retrieval and alignment benchmarks in the paper
  • Who uses it downloaded from Hugging Face every month; access gated with a citation agreement
  • Follow-on SoundCLIP, submitted to ICASSP 2027

Research Intern, Microsoft Research, May to August 2024. The curation stack uses LTU-AS, LLaVA-NeXT and Mistral 7B as scorers, with machine-generated annotation workflows and cross-modal retrieval checks. Published as AVVA (EUSIPCO 2025); the dataset backs the SoundCLIP work (submitted to ICASSP 2027).

All systems