
Multimodal ML for audio, vision, medicine and EDA
Systems that see, listen, and can be checked.
Audio-language models, multimodal evaluation, radiology agents, RL post-training. Built in a DARPA program and at Apple, Microsoft Research, Smule and Bosch.
Dissertation complete, University of Rochester · Research in the Department of Computer Science (multimodal lab of Chenliang Xu) and with Axel Wismüller's imaging group · ECE graduate student listing

A multi-agent multimodal system at Apple
Machine Learning Intern, January to August 2026. Skills: multi-agent systems, multimodal systems, conversational and agentic speech.

Live assistant on wearables, DARPA PTG
Led the live demonstration at the MIT program review; evaluated by MIT Lincoln Laboratory.

PromptReverb, ICASSP 2026 oral
48 kHz room impulse responses from text via latent rectified flow matching.

AVE-2 and OSCaR, open releases
570K-clip audio-visual dataset; OSCaR codebase, checkpoints and dataset on Hugging Face.
Current work: agents and RL post-training
Agentic speech, full-duplex dialogue, RL with verifiable rewards, agentic radiology, and RL for circuits. Submissions are labelled as submissions; work in preparation is named as such and makes no claims.
Submitted to ICASSP 2027Full-duplex spoken dialogue: turn-taking under affect
Bounds how turn-taking detector recall moves with speaker affect, across a ten-system battery.
Submitted to ICASSP 2027Reflection markers in GRPO-trained audio language models
Shows that aha-style markers track sampling narrowing rather than new reasoning; a measurement protocol for format and accuracy rewards.
Submitted to ICASSP 2027Agentic speech: audio tokens in a multimodal LLM
Token substitution gains retrieval and loses grounded generation; built on AVE-2.
Submitted to ICASSP 2027What LLM agents need to reason causally over time series
An edge-ledger protocol for causal-discovery agents against PCMCI comparators.
in preparationAgentic radiology: multi-agent report planning and generation
SPIE Emerging Topics in AI 2026, presented August 2026. An open agentic framework follow-on is in preparation.
engineeringRL post-training stack
Online RL with verifiable rewards: GRPO in verl with vLLM rollouts, FSDP2 and LoRA on an 8B vision-language model, multi-GPU H100; reward-structure analysis.
in preparationRL for circuits: schematic understanding with a verifier
Vision-language models read circuit schematics and an RL-trained corrector checks them against simulation.
in preparationClaim-blind rewards for multimodal RL
Verifiable rewards that score reasoning without seeing the claim; judge size, prompt and silence policies measured.
Systems I own
Each one answers the same questions: the problem, the component I owned, the scale, how it was evaluated, and who used or evaluated it.

Apple · Machine Learning Intern · Jan-Aug 2026
A multi-agent multimodal system (Apple, 2026)
Machine Learning Intern, Apple, January to August 2026. Work on a multi-agent multimodal system.
- Skills built multi-agent systems; multimodal systems; conversational speech; agentic speech; 3D scenes; validation at scale; explainability; perceptual-quality metrics; social norms
- Shown here the public outline only

University of Rochester · NNSA-funded materials program · 2024-2026
OpenXRD: a benchmark framework for crystallography question answering
Built and maintain the OpenXRD framework and repository: a 217-question X-ray diffraction benchmark that evaluates 74 LLMs and multimodal LLMs under open-book and closed-book settings (Digital Discovery 2026).
- Scale 217 expert-curated questions, 74 models
- Ownership all commits in the upstream repository
- Published Digital Discovery 2026
- Open repository and benchmark public
Microsoft Research · Research Intern · 2024
AVE-2: a 570K-clip audio-visual dataset and its curation stack (Microsoft Research)
Built the weak-supervision pipeline that scores 570K+ clips for temporal alignment, physical causality and source visibility with LLM/VLM scorers, and released the dataset on Hugging Face.
- Scale 570,138 clips
- Evaluation retrieval and alignment benchmarks in the paper
- Who uses it downloaded from Hugging Face every month; access gated with a citation agreement
- Follow-on SoundCLIP, submitted to ICASSP 2027

University of Rochester · NIH-supported accessibility line · 2023-2024
OSCaR: object state captioning for egocentric video
Built the OSCaR codebase and host its five checkpoints and dataset on Hugging Face; object-state captioning and state-change representation for egocentric video (Findings of NAACL 2024), supported in part by an NIH R01 on accessible video description.
- Ownership 22 of 26 commits in the team's repository
- Released five checkpoints and the dataset on Hugging Face
- Who uses it downloaded by other groups every month
- Published Findings of NAACL 2024
DARPA PTG · University of Rochester · 2022-2024
Real-time multimodal assistant on wearable devices (DARPA PTG)
A vision-language-audio assistant that sees what the wearer sees, listens, answers in speech and guides physical tasks step by step; demonstrated live at the DARPA PTG program review at MIT (Oct 2023) and evaluated by MIT Lincoln Laboratory.
- Evaluated by MIT Lincoln Laboratory
- Decision I made a video-to-text prototype that moved the project to language-first models
- Artifacts EAGLE-400K (ACM MM 2024), OSCaR (NAACL 2024 Findings), MISAR (ICCV 2023 AV4D)
- Scale live on HoloLens and a helmet-mounted camera-and-microphone rig
Programs and sponsors
Research conducted under programs supported by NSF, NIH, DARPA and DOE/NNSA, as acknowledged in the papers.
Industry research
Industry-supported research
Clinical AI research collaborations through the Rochester radiology group
Four layers of work
Every paper, published or submitted, has a page with its figure and my part in it.
Audio and speech
PromptReverb (ICASSP 2026 oral), counterfactual audio-language learning (ICASSP 2024, patent application), AVVA (EUSIPCO 2025), AVSA-Sep, three ICASSP 2027 submissions.
Vision and evaluation
VERIFY (COLM 2026), MMPerspective (NeurIPS 2025), EAGLE (ACM MM 2024), OSCaR, Possible Worlds VQA (IEEE TMM 2024), BVB world-model-style agentic video reconstruction benchmark (arXiv 2026), CAT-V (AAAI 2026, Best Demonstration Runner-up).
Medical diagnosis automation
Radiology-report agents (SPIE Emerging Topics in AI 2026), large-scale Granger causality for fMRI (NeuroImage 2025, Scientific Reports 2021), post-cardiac-arrest prognostication (Scientific Reports 2026), video world models (V-JEPA 2) applied to fMRI rendered as video (manuscript prepared), four SPIE Medical Imaging 2027 submissions.
Electronic design automation and hardware accelerators
Cryptographic hardware with side-channel and fault-injection defenses (ISCAS, SLIP, GLSVLSI 2019), analog Ising machines for combinatorial optimization (ISCAS 2020), on-chip power management, SRC TECHCON 2019 selected poster; RL for circuit design in preparation.
News
Exact venue strings; submissions are called submissions.
- 2026-10Four papers submitted to SPIE Medical Imaging 2027 and four to ICASSP 2027
- 2026-09VERIFY accepted at COLM 2026
- 2026-08Radiology-report agents presented at SPIE Emerging Topics in Artificial Intelligence 2026
- 2026-06I^2 in the CVPR 2026 Workshop Proceedings (AISTORY)
- 2026-04PromptReverb: oral paper at ICASSP 2026
- 2026-02Best Demonstration Award Runner-up, AAAI 2026
- 2026-01Machine Learning Intern at Apple, January to August 2026
- 2025-12MMPerspective at NeurIPS 2025

















