OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

Zhou, Pengfei; Peng, Xiaopeng; Song, Jiajun; Li, Chuanhao; Xu, Zhaopan; Yang, Yue; Guo, Ziyao; Zhang, Hao; Lin, Yuqi; He, Yefei; Zhao, Lirui; Liu, Shuo; Li, Tianhua; Xie, Yuxuan; Chang, Xiaojun; Qiao, Yu; Shao, Wenqi; Zhang, Kaipeng

Pengfei Zhou, Xiaopeng Peng, Jiajun Song, Chuanhao Li, Zhaopan Xu, Yue Yang, Ziyao Guo, Hao Zhang, Yuqi Lin, Yefei He, Lirui Zhao, Shuo Liu, Tianhua Li, Yuxuan Xie, Xiaojun Chang, Yu Qiao, Wenqi Shao, Kaipeng Zhang; Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 56-66

Abstract

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to data size and diversity limitations. To bridge this gap, we introduce OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks. OpenING covers diverse daily scenarios such as travel guide, design, and brainstorming, offering a robust platform for challenging interleaved generation methods. In addition, we present IntJudge, a judge model for evaluating open-ended multimodal generation methods. Trained with a novel data pipeline, our IntJudge achieves an agreement rate of 82.42% with human judgments, outperforming GPT-based evaluators by 11.34%. Extensive experiments on OpenING reveal that current interleaved generation methods still have substantial room for improvement. Key findings on interleaved image-text generation are further presented to guide the development of next-generation models.

Related Material

[pdf] [supp] [arXiv]

[bibtex]

@InProceedings{Zhou_2025_CVPR, author = {Zhou, Pengfei and Peng, Xiaopeng and Song, Jiajun and Li, Chuanhao and Xu, Zhaopan and Yang, Yue and Guo, Ziyao and Zhang, Hao and Lin, Yuqi and He, Yefei and Zhao, Lirui and Liu, Shuo and Li, Tianhua and Xie, Yuxuan and Chang, Xiaojun and Qiao, Yu and Shao, Wenqi and Zhang, Kaipeng}, title = {OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, month = {June}, year = {2025}, pages = {56-66} }