Submitted to ICASSP 2027 ยท submission

Reflection Markers and Sampling Narrowing in GRPO-Trained Audio Language Models

Schematic: Reflection Markers and Sampling Narrowing in GRPO-Trained Audio Language Models

Schematic illustration of the idea, not a figure from the paper.

Reflection markers in GRPO-trained audio language models track sampling narrowing rather than new reasoning; a measurement protocol for format and accuracy rewards.

My part. GRPO reward-structure analysis: how format reward and accuracy reward are assessed, how they interact, and how they mislead.

Reflection markers in GRPO-trained audio language models track sampling narrowing rather than new reasoning; a measurement protocol for format and accuracy rewards.

Under review. Listed as a submission; nothing here claims acceptance.

All publications