Submitted to ICASSP 2027 ยท submission
Reflection Markers and Sampling Narrowing in GRPO-Trained Audio Language Models

Reflection markers in GRPO-trained audio language models track sampling narrowing rather than new reasoning; a measurement protocol for format and accuracy rewards.
My part. GRPO reward-structure analysis: how format reward and accuracy reward are assessed, how they interact, and how they mislead.
Reflection markers in GRPO-trained audio language models track sampling narrowing rather than new reasoning; a measurement protocol for format and accuracy rewards.
Under review. Listed as a submission; nothing here claims acceptance.