Datasets and benchmarks

Datasets and benchmarks

Datasets, benchmarks and model releases by Ali Vosoughi.

  • AVE-2: 570,138 audio-visual clips scored for temporal alignment, physical causality and source visibility (Hugging Face). Published as AVVA (EUSIPCO 2025).
  • OSCaR: dataset and five checkpoints on Hugging Face (Findings of NAACL 2024).
  • OpenXRD: 217-question crystallography QA benchmark, 74 models (Digital Discovery 2026).
  • VERIFY: 600-item visual reasoning benchmark with human-annotated reasoning paths (COLM 2026); public teaser set on Hugging Face.
  • MMPerspective: 10 tasks, 2,711 images, 5,083 QA pairs (NeurIPS 2025).
  • EAGLE-400K: instruction-tuning data for egocentric video (ACM MM 2024).