Dissertation
Common-Sense Reasoning and Generation across Data Types
PhD dissertation, University of Rochester. Dissertation complete.
University of Rochester · Dissertation complete · Advisors: Chenliang Xu (Computer Science) and Axel Wismüller (Imaging Sciences) · Listed: ECE graduate students (Wismüller Lab) · NRT PhD trainees 2021-22
Abstract
Human intelligence is multimodal by design: when one sense fails, the others carry the load, and reasoning runs ahead of perception and then reshapes it. Machine intelligence remains largely device-centric, each sensor reporting alone. This dissertation argues that closing that gap is a measurement problem before it is a modeling problem. Its thesis is its title, common-sense reasoning and generation across data types, together with how to measure whether either is there. Each technical chapter pairs a system that reasons across more than one kind of data with a measurement built to decide whether what the system does deserves to be called reasoning.
Chapter 2 recovers directed, causal structure from functional brain imaging with large-scale Granger causality methods across clinical populations, and shows that a language-model agent's stated confidence about a causal claim is not calibrated to its accuracy unless every claim is forced to be earned against recorded probe evidence. Chapter 3 builds a live system that guides a person through a physical task from what it is watching, and introduces benchmarks showing that accuracy on a final answer and the fidelity of the reasoning behind it are separable. Chapter 4 puts language models over expert measurements, radiology reports with a radiologist in the loop and a diffraction benchmark built with a national materials program, and finds that expert-written context helps mid-range models more than the largest ones. Chapter 5 binds audio, video, and language without forcing every signal through language, from counterfactual audio-language supervision to a text-free audio-video representation and generated room acoustics checked against measurement. Chapter 6 measures when a conversational system should speak, bounds how far affective context shifts a turn-taking detector, and reports as a negative result that reflection language from reinforcement-trained audio-language models does not correspond to actual revision of the answer.
Four findings recur: reasoning and its surface markers come apart under measurement; correspondence between modalities is available as a check before any language-mediated judgment; grounding in evidence and physics keeps that check honest; and negative results, properly measured, are findings. The dissertation closes with the limits of these measurements and the questions they leave open.
Chapters
- Chapter 1. Introduction.
- Chapter 2. Causal Structure in the Brain and in Signals. large-scale Granger causality (lsGC, lsAGC, lsKGC) across clinical populations; agentic causal discovery over time series. Published: Scientific Reports 2021, NeuroImage 2025; submitted: lsXGC (Scientific Reports), causal agents (ICASSP 2027), four SPIE Medical Imaging 2027 papers.
- Chapter 3. Vision Shapes the Language Model. the live DARPA PTG assistant (MISAR), EAGLE-400K, OSCaR, CAT-V, BVB; evaluation with VERIFY (COLM 2026), MMPerspective (NeurIPS 2025) and the survey of video understanding with LLMs (IEEE TCSVT 2026).
- Chapter 4. Language Models over Expert Measurements. radiology-report agents (SPIE Emerging Topics in AI 2026) and OpenXRD (Digital Discovery 2026) under an NNSA-funded materials program.
- Chapter 5. Multimodality and Interaction. Possible Worlds VQA (IEEE TMM 2024), counterfactual audio-language learning (ICASSP 2024), invisible-sound separation, AVVA (EUSIPCO 2025), SoundCLIP (submitted to ICASSP 2027), PromptReverb (ICASSP 2026 oral), I^2 (CVPRW 2026).
- Chapter 6. Reinforcement Learning, Conversation, and Emotion. full-duplex turn-taking under affect and reflection markers in GRPO-trained audio language models (both submitted to ICASSP 2027); claim-blind rewards in preparation.
- Chapter 7. Conclusion and Outlook.
The dissertation text is not distributed before the defense. Each chapter's published work is on the publications page.