IEEE/CVF CVPR 2026 Workshop Proceedings (AISTORY), CVPRW 2026
I^2: Generating Instructional Illustrations via Text-Conditioned Diffusion

Coherent instructional illustrations from procedural text with text-conditioned diffusion, pairwise cross-image attention, long-step text encoding and preference optimization.
My part. Contributed core multimodal outputs of the DARPA/PTG instructional-vision line, including I^2.
Coherent instructional illustrations from procedural text with text-conditioned diffusion, pairwise cross-image attention, long-step text encoding and preference optimization.