Systems

Systems I own

Schematic: a multi-agent system with a feedback loop

Apple · Machine Learning Intern · Jan-Aug 2026

A multi-agent multimodal system (Apple, 2026)

Machine Learning Intern, Apple, January to August 2026. Work on a multi-agent multimodal system.

  • Skills built multi-agent systems; multimodal systems; conversational speech; agentic speech; 3D scenes; validation at scale; explainability; perceptual-quality metrics; social norms
  • Shown here the public outline only
Cover artwork of Digital Discovery, volume 5 issue 5 (May 2026), the issue featuring OPENXRD

University of Rochester · NNSA-funded materials program · 2024-2026

OpenXRD: a benchmark framework for crystallography question answering

Built and maintain the OpenXRD framework and repository: a 217-question X-ray diffraction benchmark that evaluates 74 LLMs and multimodal LLMs under open-book and closed-book settings (Digital Discovery 2026).

  • Scale 217 expert-curated questions, 74 models
  • Ownership all commits in the upstream repository
  • Published Digital Discovery 2026
  • Open repository and benchmark public

Microsoft Research · Research Intern · 2024

AVE-2: a 570K-clip audio-visual dataset and its curation stack (Microsoft Research)

Built the weak-supervision pipeline that scores 570K+ clips for temporal alignment, physical causality and source visibility with LLM/VLM scorers, and released the dataset on Hugging Face.

  • Scale 570,138 clips
  • Evaluation retrieval and alignment benchmarks in the paper
  • Who uses it downloaded from Hugging Face every month; access gated with a citation agreement
  • Follow-on SoundCLIP, submitted to ICASSP 2027
Ali Vosoughi: OSCaR object-state frames from an egocentric video

University of Rochester · NIH-supported accessibility line · 2023-2024

OSCaR: object state captioning for egocentric video

Built the OSCaR codebase and host its five checkpoints and dataset on Hugging Face; object-state captioning and state-change representation for egocentric video (Findings of NAACL 2024), supported in part by an NIH R01 on accessible video description.

  • Ownership 22 of 26 commits in the team's repository
  • Released five checkpoints and the dataset on Hugging Face
  • Who uses it downloaded by other groups every month
  • Published Findings of NAACL 2024

DARPA PTG · University of Rochester · 2022-2024

Real-time multimodal assistant on wearable devices (DARPA PTG)

A vision-language-audio assistant that sees what the wearer sees, listens, answers in speech and guides physical tasks step by step; demonstrated live at the DARPA PTG program review at MIT (Oct 2023) and evaluated by MIT Lincoln Laboratory.

  • Evaluated by MIT Lincoln Laboratory
  • Decision I made a video-to-text prototype that moved the project to language-first models
  • Artifacts EAGLE-400K (ACM MM 2024), OSCaR (NAACL 2024 Findings), MISAR (ICCV 2023 AV4D)
  • Scale live on HoloLens and a helmet-mounted camera-and-microphone rig