Microsoft Research · Research Intern · 2024
AVE-2: a 570K-clip audio-visual dataset and its curation stack (Microsoft Research)
Built the weak-supervision pipeline that scores 570K+ clips for temporal alignment, physical causality and source visibility with LLM/VLM scorers, and released the dataset on Hugging Face.
- Scale 570,138 clips
- Evaluation retrieval and alignment benchmarks in the paper
- Who uses it downloaded from Hugging Face every month; access gated with a citation agreement
- Follow-on SoundCLIP, submitted to ICASSP 2027
Research Intern, Microsoft Research, May to August 2024. The curation stack uses LTU-AS, LLaVA-NeXT and Mistral 7B as scorers, with machine-generated annotation workflows and cross-modal retrieval checks. Published as AVVA (EUSIPCO 2025); the dataset backs the SoundCLIP work (submitted to ICASSP 2027).