|
Kaiwen Zhu (朱楷文)
I am a Ph.D. student at Shanghai Jiao Tong
University (SJTU) supervised by Prof. Chao
Dong. I have been working as a research intern at Shanghai AI Lab since 2023, mentored by Dr. Yihao Liu. I obtained my B.Eng. degree in computer
science
and technology from SJTU in 2024.
Currently, my research mainly focuses on multi-modal generation, especially masked generative
models. Prior to this, I worked on low-level vision, including image restoration and image quality
assessment.
Google
Scholar
/
Github
Email: zhukaiwensq [at] gmail [dot] com
|
|
Selected Research
*: Equal contribution †: Corresponding author
|
2026
|
|
Accelerating Masked Image Generation by Learning Controlled Latent Dynamics
Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin,
Jiayang Li, Juncheng Yan, Yu Qiao, Jinjin Gu, Yihao Liu†
ACM International Conference on Multimedia (ACM MM), 2026
project page
/
arXiv
/
paper
/
code
We propose MIGM-Shortcut, a lightweight model that leverage the current feature to predict the next feature in Masked Image Generation Models (MIGMs). MIGM-Shortcut bypasses the cumbersome base model by taking a shortcut in the latent space to accelerate generation. It aims to address the computation redundancy in MIGMs: the rich semantics in the continuous features are lost when sampling discrete tokens.
|
2025
|
|
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Yi Xin*, Qi Qin*, Siqi Luo, Kaiwen Zhu, Juncheng Yan, Yan Tai, Jiayi Lei, Yuewen Cao, Keqi Wang, Yibin Wang, Jinbin Bai, Qian Yu, Dengyang Jiang, Yuandong Pu, Haoxing Chen, Le Zhuo, Junjun He, Gen Luo, Tianbin Li, Ming Hu, Jin Ye, Shenglong Ye, Bo Zhang, Chang Xu, Wenhai Wang, Hongsheng Li, Guangtao Zhai, Tianfan Xue, Bin Fu†, Xiaohong Liu†, Yu Qiao†, Yihao Liu†
Technical report
project page
/
arXiv
/
paper
/
code
We introduce Lumina-DiMOO, an open-source foundational model for seamless multimodal generation and understanding. It utilizes a fully discrete diffusion modeling to handle inputs and outputs across various modalities.
|
|
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
Shuo Cao*, Jiayang Li*, Xiaohui Li, Yuandong Pu, Kaiwen Zhu, Yuanting Gao, Siqi Luo, Yi Xin, Qi Qin, Yu Zhou, Xiangyu Chen, Wenlong Zhang, Bin Fu, Yu Qiao, Yihao Liu†
International Conference on Machine Learning (ICML) Spotlight, 2026
project page
/
arXiv
/
paper
/
code
Perceptual-level image understanding focuses on how an image looks and feels—capturing aesthetics, quality degradations, structural regularity, and surface texture. These fine-grained perceptual cues differ fundamentally from semantic recognition, yet remain underexplored in MLLMs. To address this, we introduce UniPercept, a unified framework that defines, evaluates, and improves perceptual-level visual understanding across the IAA (Image Aesthetics Assessment), IQA (Image Quality Assessment), and ISTA (Image Structure and Texture Assessment) domains.
|
|
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
Shuo Cao, Nan Ma, Jiayang Li, Xiaohui Li, Lihao Shao, Kaiwen Zhu, Yu Zhou, Yuandong Pu, Jiarui Wu, Jiaquan Wang, Bo Qu, Wenhai Wang, Yu Qiao, Dajuin Yao†, Yihao Liu†
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
project page
/
arXiv
/
paper
/
code
We present
ArtiMuse: an innovative MLLM-based IAA model with Joint Scoring and Expert-Level Understanding capabilities;
ArtiMuse-10K: the first expert-curated image aesthetic dataset comprising 10,000 images spanning 5 main categories and 15 subcategories, each annotated by professional experts with 8-dimensional attributes analysis and a holistic score.
|
|
Exploring Scalable Unified Modeling for General Low-Level Vision
Xiangyu Chen*, Kaiwen Zhu*, Yuandong Pu*, Shuo Cao, Xiaohui Li, Wenlong Zhang, Yihao Liu, Yu Qiao, Jiantao Zhou†, Chao Dong†
Preprint
arXiv
/
paper
/
code
We build GenLV, a universal model that unifies cross-domain tasks such as image restoration, enhancement, stylization, and
feature extraction. We validate the effectiveness and scalability of GenLV across over 100 tasks. GenLV demonstrates strong performance in zero-shot, few-shot, and task-specific fine-tuning settings, highlighting its potential as a foundation model for low-level vision.
|
2024
|
|
An Intelligent Agentic System for Complex Image Restoration Problems
Kaiwen Zhu*,
Jinjin Gu*,
Zhiyuan You, Yu Qiao, Chao Dong†
International Conference on Learning Representations (ICLR), 2025
project page
/
arXiv
/
paper
/
code
Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics humans'
approach to image processing by following five key stages: Perception, Scheduling, Execution,
Reflection, and Rescheduling. AgenticIR leverages large language models and vision-language models
that interact via text generation to dynamically operate a toolbox of image restoration models.
|
|