Kaiwen Zhu (朱楷文)

I am a Ph.D. student at Shanghai Jiao Tong University (SJTU) supervised by Prof. Chao Dong. I have been working as a research intern at Shanghai AI Lab since 2023, mentored by Dr. Yihao Liu. I obtained my B.Eng. degree in computer science and technology from SJTU in 2024.

Currently, my research mainly focuses on multi-modal generation, especially masked generative models. Prior to this, I worked on low-level vision, including image restoration and image quality assessment.

Google Scholar  /  Github

Email: zhukaiwensq [at] gmail [dot] com

profile photo

Selected Research

*: Equal contribution     : Corresponding author

2026

Accelerating Masked Image Generation by Learning Controlled Latent Dynamics
Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin, Jiayang Li, Juncheng Yan, Yu Qiao, Jinjin Gu, Yihao Liu
ACM International Conference on Multimedia (ACM MM), 2026
project page / arXiv / paper / code

We propose MIGM-Shortcut, a lightweight model that leverage the current feature to predict the next feature in Masked Image Generation Models (MIGMs). MIGM-Shortcut bypasses the cumbersome base model by taking a shortcut in the latent space to accelerate generation. It aims to address the computation redundancy in MIGMs: the rich semantics in the continuous features are lost when sampling discrete tokens.

2025

PICABench: How Far Are We from Physically Realistic Image Editing?
Yuandong Pu, Le Zhuo, Songhao Han, Jinbo Xing, Kaiwen Zhu, Shuo Cao, Bin Fu, Si Liu, Hongsheng Li, Yu Qiao, Wenlong Zhang, Xi Chen, Yihao Liu
International Conference on Learning Representations (ICLR), 2026
project page / arXiv / paper / code

PICABench probes how far current editing models are from physically realistic image manipulation. It ties together: PICABench benchmark: physics-aware editing cases spanning eight laws across Optics, Mechanics, and State Transition. PICAEval metric: region-grounded, QA-based verification with human-annotated regions of interest and spatially anchored yes/no questions. PICA-100K dataset: synthetic, video-derived training data that boosts physics consistency when used for fine-tuning.

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Yi Xin*, Qi Qin*, Siqi Luo, Kaiwen Zhu, Juncheng Yan, Yan Tai, Jiayi Lei, Yuewen Cao, Keqi Wang, Yibin Wang, Jinbin Bai, Qian Yu, Dengyang Jiang, Yuandong Pu, Haoxing Chen, Le Zhuo, Junjun He, Gen Luo, Tianbin Li, Ming Hu, Jin Ye, Shenglong Ye, Bo Zhang, Chang Xu, Wenhai Wang, Hongsheng Li, Guangtao Zhai, Tianfan Xue, Bin Fu, Xiaohong Liu, Yu Qiao, Yihao Liu
Technical report
project page / arXiv / paper / code

We introduce Lumina-DiMOO, an open-source foundational model for seamless multimodal generation and understanding. It utilizes a fully discrete diffusion modeling to handle inputs and outputs across various modalities.

ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
Shuo Cao, Nan Ma, Jiayang Li, Xiaohui Li, Lihao Shao, Kaiwen Zhu, Yu Zhou, Yuandong Pu, Jiarui Wu, Jiaquan Wang, Bo Qu, Wenhai Wang, Yu Qiao, Dajuin Yao, Yihao Liu
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
project page / arXiv / paper / code

We present ArtiMuse: an innovative MLLM-based IAA model with Joint Scoring and Expert-Level Understanding capabilities; ArtiMuse-10K: the first expert-curated image aesthetic dataset comprising 10,000 images spanning 5 main categories and 15 subcategories, each annotated by professional experts with 8-dimensional attributes analysis and a holistic score.

Exploring Scalable Unified Modeling for General Low-Level Vision
Xiangyu Chen*, Kaiwen Zhu*, Yuandong Pu*, Shuo Cao, Xiaohui Li, Wenlong Zhang, Yihao Liu, Yu Qiao, Jiantao Zhou, Chao Dong
Preprint
arXiv / paper / code

We build GenLV, a universal model that unifies cross-domain tasks such as image restoration, enhancement, stylization, and feature extraction. We validate the effectiveness and scalability of GenLV across over 100 tasks. GenLV demonstrates strong performance in zero-shot, few-shot, and task-specific fine-tuning settings, highlighting its potential as a foundation model for low-level vision.

2024

An Intelligent Agentic System for Complex Image Restoration Problems
Kaiwen Zhu*, Jinjin Gu*, Zhiyuan You, Yu Qiao, Chao Dong
International Conference on Learning Representations (ICLR), 2025
project page / arXiv / paper / code

Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics humans' approach to image processing by following five key stages: Perception, Scheduling, Execution, Reflection, and Rescheduling. AgenticIR leverages large language models and vision-language models that interact via text generation to dynamically operate a toolbox of image restoration models.

Services

Conference Reviewer

Conference on Computer Vision and Pattern Recognition (CVPR)

European Conference on Computer Vision (ECCV)

Journal Reviewer

IEEE Transactions on Multimedia (TMM)

International Journal of Computer Vision (IJCV)


Template from JonBarron