IEEE/CVF CVPR 2026 Workshop Proceedings (AISTORY), CVPRW 2026

I^2: Generating Instructional Illustrations via Text-Conditioned Diffusion

Figure: I^2: Generating Instructional Illustrations via Text-Conditioned Diffusion

Jing Bi, Yunlong (Yolo) Tang, Pinxin Liu, Chao Huang, Ali Vosoughi, Jiarui Wu, Jinxi He, Chenliang Xu

Coherent instructional illustrations from procedural text with text-conditioned diffusion, pairwise cross-image attention, long-step text encoding and preference optimization.

My part. Contributed core multimodal outputs of the DARPA/PTG instructional-vision line, including I^2.

Coherent instructional illustrations from procedural text with text-conditioned diffusion, pairwise cross-image attention, long-step text encoding and preference optimization.

All publications