Systems
Systems I own

Apple · Machine Learning Intern · Jan-Aug 2026
A multi-agent multimodal system (Apple, 2026)
Machine Learning Intern, Apple, January to August 2026. Work on a multi-agent multimodal system.
- Skills built multi-agent systems; multimodal systems; conversational speech; agentic speech; 3D scenes; validation at scale; explainability; perceptual-quality metrics; social norms
- Shown here the public outline only

University of Rochester · NNSA-funded materials program · 2024-2026
OpenXRD: a benchmark framework for crystallography question answering
Built and maintain the OpenXRD framework and repository: a 217-question X-ray diffraction benchmark that evaluates 74 LLMs and multimodal LLMs under open-book and closed-book settings (Digital Discovery 2026).
- Scale 217 expert-curated questions, 74 models
- Ownership all commits in the upstream repository
- Published Digital Discovery 2026
- Open repository and benchmark public
Microsoft Research · Research Intern · 2024
AVE-2: a 570K-clip audio-visual dataset and its curation stack (Microsoft Research)
Built the weak-supervision pipeline that scores 570K+ clips for temporal alignment, physical causality and source visibility with LLM/VLM scorers, and released the dataset on Hugging Face.
- Scale 570,138 clips
- Evaluation retrieval and alignment benchmarks in the paper
- Who uses it downloaded from Hugging Face every month; access gated with a citation agreement
- Follow-on SoundCLIP, submitted to ICASSP 2027

University of Rochester · NIH-supported accessibility line · 2023-2024
OSCaR: object state captioning for egocentric video
Built the OSCaR codebase and host its five checkpoints and dataset on Hugging Face; object-state captioning and state-change representation for egocentric video (Findings of NAACL 2024), supported in part by an NIH R01 on accessible video description.
- Ownership 22 of 26 commits in the team's repository
- Released five checkpoints and the dataset on Hugging Face
- Who uses it downloaded by other groups every month
- Published Findings of NAACL 2024
DARPA PTG · University of Rochester · 2022-2024
Real-time multimodal assistant on wearable devices (DARPA PTG)
A vision-language-audio assistant that sees what the wearer sees, listens, answers in speech and guides physical tasks step by step; demonstrated live at the DARPA PTG program review at MIT (Oct 2023) and evaluated by MIT Lincoln Laboratory.
- Evaluated by MIT Lincoln Laboratory
- Decision I made a video-to-text prototype that moved the project to language-first models
- Artifacts EAGLE-400K (ACM MM 2024), OSCaR (NAACL 2024 Findings), MISAR (ICCV 2023 AV4D)
- Scale live on HoloLens and a helmet-mounted camera-and-microphone rig