|
Kaiwen Zhu (朱楷文)
I am a Ph.D. student at Shanghai Jiao Tong
University (SJTU) supervised by Prof. Chao
Dong. I have been working as a research intern at Shanghai AI Lab since 2023, mentored by Dr. Yihao Liu. I obtained my B.Eng. degree in computer
science
and technology from SJTU in 2024.
Currently, my research mainly focuses on multi-modal generation, especially masked generative
models. Prior to this, I worked on low-level vision, including image restoration and image quality
assessment.
Google
Scholar
/
Github
Email: zhukaiwensq [at] gmail [dot] com
|
|
Selected Research
*: Equal contribution †: Corresponding author
|
2026
|
|
Accelerating Masked Image Generation by Learning Controlled Latent Dynamics
Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin,
Jiayang Li, Juncheng Yan, Yu Qiao, Jinjin Gu, Yihao Liu†
ACM International Conference on Multimedia (ACM MM), 2026
project page
/
arXiv
/
paper
/
code
We propose MIGM-Shortcut, a lightweight model that leverage the current feature to predict the next feature in Masked Image Generation Models (MIGMs). MIGM-Shortcut bypasses the cumbersome base model by taking a shortcut in the latent space to accelerate generation. It aims to address the computation redundancy in MIGMs: the rich semantics in the continuous features are lost when sampling discrete tokens.
|
2025
|
|
PICABench: How Far Are We from Physically Realistic Image Editing?
Yuandong Pu, Le Zhuo, Songhao Han, Jinbo Xing, Kaiwen Zhu, Shuo Cao, Bin Fu, Si Liu, Hongsheng Li, Yu Qiao, Wenlong Zhang, Xi Chen†, Yihao Liu†
International Conference on Learning Representations (ICLR), 2026
project page
/
arXiv
/
paper
/
code
PICABench probes how far current editing models are from physically realistic image manipulation. It ties together:
PICABench benchmark: physics-aware editing cases spanning eight laws across Optics, Mechanics, and State Transition.
PICAEval metric: region-grounded, QA-based verification with human-annotated regions of interest and spatially anchored yes/no questions.
PICA-100K dataset: synthetic, video-derived training data that boosts physics consistency when used for fine-tuning.
|
|
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Yi Xin*, Qi Qin*, Siqi Luo, Kaiwen Zhu, Juncheng Yan, Yan Tai, Jiayi Lei, Yuewen Cao, Keqi Wang, Yibin Wang, Jinbin Bai, Qian Yu, Dengyang Jiang, Yuandong Pu, Haoxing Chen, Le Zhuo, Junjun He, Gen Luo, Tianbin Li, Ming Hu, Jin Ye, Shenglong Ye, Bo Zhang, Chang Xu, Wenhai Wang, Hongsheng Li, Guangtao Zhai, Tianfan Xue, Bin Fu†, Xiaohong Liu†, Yu Qiao†, Yihao Liu†
Technical report
project page
/
arXiv
/
paper
/
code
We introduce Lumina-DiMOO, an open-source foundational model for seamless multimodal generation and understanding. It utilizes a fully discrete diffusion modeling to handle inputs and outputs across various modalities.
|
|
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
Shuo Cao, Nan Ma, Jiayang Li, Xiaohui Li, Lihao Shao, Kaiwen Zhu, Yu Zhou, Yuandong Pu, Jiarui Wu, Jiaquan Wang, Bo Qu, Wenhai Wang, Yu Qiao, Dajuin Yao†, Yihao Liu†
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
project page
/
arXiv
/
paper
/
code
We present
ArtiMuse: an innovative MLLM-based IAA model with Joint Scoring and Expert-Level Understanding capabilities;
ArtiMuse-10K: the first expert-curated image aesthetic dataset comprising 10,000 images spanning 5 main categories and 15 subcategories, each annotated by professional experts with 8-dimensional attributes analysis and a holistic score.
|
|
Exploring Scalable Unified Modeling for General Low-Level Vision
Xiangyu Chen*, Kaiwen Zhu*, Yuandong Pu*, Shuo Cao, Xiaohui Li, Wenlong Zhang, Yihao Liu, Yu Qiao, Jiantao Zhou†, Chao Dong†
Preprint
arXiv
/
paper
/
code
We build GenLV, a universal model that unifies cross-domain tasks such as image restoration, enhancement, stylization, and
feature extraction. We validate the effectiveness and scalability of GenLV across over 100 tasks. GenLV demonstrates strong performance in zero-shot, few-shot, and task-specific fine-tuning settings, highlighting its potential as a foundation model for low-level vision.
|
2024
|
|
An Intelligent Agentic System for Complex Image Restoration Problems
Kaiwen Zhu*,
Jinjin Gu*,
Zhiyuan You, Yu Qiao, Chao Dong†
International Conference on Learning Representations (ICLR), 2025
project page
/
arXiv
/
paper
/
code
Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics humans'
approach to image processing by following five key stages: Perception, Scheduling, Execution,
Reflection, and Rescheduling. AgenticIR leverages large language models and vision-language models
that interact via text generation to dynamically operate a toolbox of image restoration models.
|
Services
Conference Reviewer
Conference on Computer Vision and Pattern Recognition (CVPR)
European Conference on Computer Vision (ECCV)
Journal Reviewer
IEEE Transactions on Multimedia (TMM)
International Journal of Computer Vision (IJCV)
|
|