Current work: agents and RL post-training

Agentic speech, full-duplex dialogue, RL with verifiable rewards, agentic radiology, and RL for circuits. Submissions are labelled as submissions; work in preparation is named as such and makes no claims.

Systems I own

Each one answers the same questions: the problem, the component I owned, the scale, how it was evaluated, and who used or evaluated it.

Schematic: a multi-agent system with a feedback loop

Apple · Machine Learning Intern · Jan-Aug 2026

A multi-agent multimodal system (Apple, 2026)

Machine Learning Intern, Apple, January to August 2026. Work on a multi-agent multimodal system.

  • Skills built multi-agent systems; multimodal systems; conversational speech; agentic speech; 3D scenes; validation at scale; explainability; perceptual-quality metrics; social norms
  • Shown here the public outline only
Cover artwork of Digital Discovery, volume 5 issue 5 (May 2026), the issue featuring OPENXRD

University of Rochester · NNSA-funded materials program · 2024-2026

OpenXRD: a benchmark framework for crystallography question answering

Built and maintain the OpenXRD framework and repository: a 217-question X-ray diffraction benchmark that evaluates 74 LLMs and multimodal LLMs under open-book and closed-book settings (Digital Discovery 2026).

  • Scale 217 expert-curated questions, 74 models
  • Ownership all commits in the upstream repository
  • Published Digital Discovery 2026
  • Open repository and benchmark public

Microsoft Research · Research Intern · 2024

AVE-2: a 570K-clip audio-visual dataset and its curation stack (Microsoft Research)

Built the weak-supervision pipeline that scores 570K+ clips for temporal alignment, physical causality and source visibility with LLM/VLM scorers, and released the dataset on Hugging Face.

  • Scale 570,138 clips
  • Evaluation retrieval and alignment benchmarks in the paper
  • Who uses it downloaded from Hugging Face every month; access gated with a citation agreement
  • Follow-on SoundCLIP, submitted to ICASSP 2027
Ali Vosoughi: OSCaR object-state frames from an egocentric video

University of Rochester · NIH-supported accessibility line · 2023-2024

OSCaR: object state captioning for egocentric video

Built the OSCaR codebase and host its five checkpoints and dataset on Hugging Face; object-state captioning and state-change representation for egocentric video (Findings of NAACL 2024), supported in part by an NIH R01 on accessible video description.

  • Ownership 22 of 26 commits in the team's repository
  • Released five checkpoints and the dataset on Hugging Face
  • Who uses it downloaded by other groups every month
  • Published Findings of NAACL 2024

DARPA PTG · University of Rochester · 2022-2024

Real-time multimodal assistant on wearable devices (DARPA PTG)

A vision-language-audio assistant that sees what the wearer sees, listens, answers in speech and guides physical tasks step by step; demonstrated live at the DARPA PTG program review at MIT (Oct 2023) and evaluated by MIT Lincoln Laboratory.

  • Evaluated by MIT Lincoln Laboratory
  • Decision I made a video-to-text prototype that moved the project to language-first models
  • Artifacts EAGLE-400K (ACM MM 2024), OSCaR (NAACL 2024 Findings), MISAR (ICCV 2023 AV4D)
  • Scale live on HoloLens and a helmet-mounted camera-and-microphone rig

Programs and sponsors

Research conducted under programs supported by NSF, NIH, DARPA and DOE/NNSA, as acknowledged in the papers.

Industry research

Industry-supported research

Clinical AI research collaborations through the Rochester radiology group

Co-authors

Names as printed on the papers; each paper page lists the full author list where a public version exists.

Four layers of work

Every paper, published or submitted, has a page with its figure and my part in it.

Audio and speech

PromptReverb (ICASSP 2026 oral), counterfactual audio-language learning (ICASSP 2024, patent application), AVVA (EUSIPCO 2025), AVSA-Sep, three ICASSP 2027 submissions.

Schematic: Bounding Affect-Associated Recall Variation for Turn-Taking Detectors in Full-Duplex Spoken DialogueSchematic: Reflection Markers and Sampling Narrowing in GRPO-Trained Audio Language ModelsSchematic: projected audio tokens in a multimodal LLM, retrieval versus grounded generationFigure: Multimodal Room Impulse Response Generation Through Latent Rectified Flow MatchingFigure: Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation ModelFigure: Learning Audio Concepts from Counterfactual Natural Language

Vision and evaluation

VERIFY (COLM 2026), MMPerspective (NeurIPS 2025), EAGLE (ACM MM 2024), OSCaR, Possible Worlds VQA (IEEE TMM 2024), BVB world-model-style agentic video reconstruction benchmark (arXiv 2026), CAT-V (AAAI 2026, Best Demonstration Runner-up).

Figure: VERIFY: A Benchmark of Visual Reasoning for Multimodal Reasoning FidelityFigure: BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in BlenderFigure: What Language Model Agents Need to Reason Causally over Multivariate Time-Series SignalsFigure: I^2: Generating Instructional Illustrations via Text-Conditioned DiffusionFront cover of Digital Discovery, volume 5 issue 5 (May 2026), featuring OPENXRDFigure: Video Understanding with Large Language Models: A Survey

Medical diagnosis automation

Radiology-report agents (SPIE Emerging Topics in AI 2026), large-scale Granger causality for fMRI (NeuroImage 2025, Scientific Reports 2021), post-cardiac-arrest prognostication (Scientific Reports 2026), video world models (V-JEPA 2) applied to fMRI rendered as video (manuscript prepared), four SPIE Medical Imaging 2027 submissions.

Figure: Large-scale extended Granger causality (lsXGC) for inferring directed dependence in large networks from short multivariate time-series dataFigure: Multi-agent Planning and Generation of Radiology Reports using Multimodal Agentic FrameworksFigure: Large-Scale Augmented Granger Causality for ADHD-Related Resting-State fMRI DysconnectivityFigure: Prior-Guided Sparse Large-Scale Granger Causality for ADHD-Related Resting-State fMRI DysconnectivityFigure: Detecting Developmental Age-Group Signals in Naturalistic fMRI Using Large-Scale Augmented Granger CausalityFigure: Efficient State-Space Modeling with Mamba for CT-Based Brain-Death Classification versus Transformer Baselines

Electronic design automation and hardware accelerators

Cryptographic hardware with side-channel and fault-injection defenses (ISCAS, SLIP, GLSVLSI 2019), analog Ising machines for combinatorial optimization (ISCAS 2020), on-chip power management, SRC TECHCON 2019 selected poster; RL for circuit design in preparation.

Schematic: coupled oscillators lock to solve an optimization problemSchematic: bus-invert encoder hides the power traceSchematic: an on-chip regulator absorbs a glitch before the chipSchematic: correlation and information distinguishers combined to recover a key

All publications, including submissions

News

Exact venue strings; submissions are called submissions.

  • 2026-10Four papers submitted to SPIE Medical Imaging 2027 and four to ICASSP 2027
  • 2026-09VERIFY accepted at COLM 2026
  • 2026-08Radiology-report agents presented at SPIE Emerging Topics in Artificial Intelligence 2026
  • 2026-06I^2 in the CVPR 2026 Workshop Proceedings (AISTORY)
  • 2026-04PromptReverb: oral paper at ICASSP 2026
  • 2026-02Best Demonstration Award Runner-up, AAAI 2026
  • 2026-01Machine Learning Intern at Apple, January to August 2026
  • 2025-12MMPerspective at NeurIPS 2025

All news