Kaiwen Zhu (朱楷文)

I am a Ph.D. student at Shanghai Jiao Tong University (SJTU) supervised by Prof. Chao Dong. I have been working as a research intern at Shanghai AI Lab since 2023, mentored by Dr. Yihao Liu. I obtained my B.Eng. degree in computer science and technology from SJTU in 2024.

Currently, my research mainly focuses on multi-modal generation, especially masked generative models. Prior to this, I worked on low-level vision, including image restoration and image quality assessment.

Google Scholar  /  Github

Email: zhukaiwensq [at] gmail [dot] com

profile photo

Selected Research

*: Equal contribution     : Corresponding author

2026

Accelerating Masked Image Generation by Learning Controlled Latent Dynamics
Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin, Jiayang Li, Juncheng Yan, Yu Qiao, Jinjin Gu, Yihao Liu
ACM International Conference on Multimedia (ACM MM), 2026
project page / arXiv / paper / code

We propose MIGM-Shortcut, a lightweight model that leverage the current feature to predict the next feature in Masked Image Generation Models (MIGMs). MIGM-Shortcut bypasses the cumbersome base model by taking a shortcut in the latent space to accelerate generation. It aims to address the computation redundancy in MIGMs: the rich semantics in the continuous features are lost when sampling discrete tokens.

2025

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Yi Xin*, Qi Qin*, Siqi Luo, Kaiwen Zhu, Juncheng Yan, Yan Tai, Jiayi Lei, Yuewen Cao, Keqi Wang, Yibin Wang, Jinbin Bai, Qian Yu, Dengyang Jiang, Yuandong Pu, Haoxing Chen, Le Zhuo, Junjun He, Gen Luo, Tianbin Li, Ming Hu, Jin Ye, Shenglong Ye, Bo Zhang, Chang Xu, Wenhai Wang, Hongsheng Li, Guangtao Zhai, Tianfan Xue, Bin Fu, Xiaohong Liu, Yu Qiao, Yihao Liu
Technical report
project page / arXiv / paper / code

We introduce Lumina-DiMOO, an open-source foundational model for seamless multimodal generation and understanding. It utilizes a fully discrete diffusion modeling to handle inputs and outputs across various modalities.

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
Shuo Cao*, Jiayang Li*, Xiaohui Li, Yuandong Pu, Kaiwen Zhu, Yuanting Gao, Siqi Luo, Yi Xin, Qi Qin, Yu Zhou, Xiangyu Chen, Wenlong Zhang, Bin Fu, Yu Qiao, Yihao Liu
International Conference on Machine Learning (ICML) Spotlight, 2026
project page / arXiv / paper / code

Perceptual-level image understanding focuses on how an image looks and feels—capturing aesthetics, quality degradations, structural regularity, and surface texture. These fine-grained perceptual cues differ fundamentally from semantic recognition, yet remain underexplored in MLLMs. To address this, we introduce UniPercept, a unified framework that defines, evaluates, and improves perceptual-level visual understanding across the IAA (Image Aesthetics Assessment), IQA (Image Quality Assessment), and ISTA (Image Structure and Texture Assessment) domains.

ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
Shuo Cao, Nan Ma, Jiayang Li, Xiaohui Li, Lihao Shao, Kaiwen Zhu, Yu Zhou, Yuandong Pu, Jiarui Wu, Jiaquan Wang, Bo Qu, Wenhai Wang, Yu Qiao, Dajuin Yao, Yihao Liu
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
project page / arXiv / paper / code

We present ArtiMuse: an innovative MLLM-based IAA model with Joint Scoring and Expert-Level Understanding capabilities; ArtiMuse-10K: the first expert-curated image aesthetic dataset comprising 10,000 images spanning 5 main categories and 15 subcategories, each annotated by professional experts with 8-dimensional attributes analysis and a holistic score.

Exploring Scalable Unified Modeling for General Low-Level Vision
Xiangyu Chen*, Kaiwen Zhu*, Yuandong Pu*, Shuo Cao, Xiaohui Li, Wenlong Zhang, Yihao Liu, Yu Qiao, Jiantao Zhou, Chao Dong
Preprint
arXiv / paper / code

We build GenLV, a universal model that unifies cross-domain tasks such as image restoration, enhancement, stylization, and feature extraction. We validate the effectiveness and scalability of GenLV across over 100 tasks. GenLV demonstrates strong performance in zero-shot, few-shot, and task-specific fine-tuning settings, highlighting its potential as a foundation model for low-level vision.

2024

An Intelligent Agentic System for Complex Image Restoration Problems
Kaiwen Zhu*, Jinjin Gu*, Zhiyuan You, Yu Qiao, Chao Dong
International Conference on Learning Representations (ICLR), 2025
project page / arXiv / paper / code

Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics humans' approach to image processing by following five key stages: Perception, Scheduling, Execution, Reflection, and Rescheduling. AgenticIR leverages large language models and vision-language models that interact via text generation to dynamically operate a toolbox of image restoration models.


Template from JonBarron