About me

Hi, I'm Chenxi Song, currently a postdoctoral researcher at AGI Lab, Westlake University (Jan 2025 – Jan 2027), supervised by Prof. Chi Zhang. I received my Ph.D. degree in Engineering from Jilin University in 2024, where I focused on 3D Computer Vision and Computer Graphics under the supervision of Prof. Shigang Wang, with co-supervision from Prof. Jian Wei and Prof. Yan Zhao.

My current research interests lie in world models, video generation models, and 3D & 4D scene reconstruction and generation, with the hope of advancing world models for world understanding, real-world simulation, and the game, film and television industries. I am actively engaged in the academic community, serving as a reviewer for top-tier AI conferences and journals including NeurIPS, CVPR, ECCV, AAAI, MM, and SIGGRAPH. I am honored that our earlier world model work, WorldForge, was selected as a CVPR 2026 Highlight. More recently, we released World in World, a 4D world model that explores the real world inside a video. I will continue to focus on controllable video generation and world model research. Welcome to collaborate and exchange ideas!

News

  • Sep 2026🚀 World in World is released — paper, project page and code are out!
  • Apr 2026🏆 WorldForge is selected as CVPR Highlight!
  • Feb 2026🎉 5 papers are accepted by CVPR 2026!
  • Feb 2026🔥 WorldForge code is now open-sourced!
  • Sep 2025🔥 Released WorldForge, a training-free world model project.
  • Jan 2025Joined Westlake University School of Engineering as a postdoctoral researcher.
  • Sep 2024Graduated from Jilin University with Ph.D. degree.
  • May 2024Our work FewarNet on sparse-view multi-view synthesis was published in T-CSVT.

Publications

Full list on Google Scholar
World in World teaser
World in World: Explore the World with World Models
arXiv preprint, 2026
page|pdf|code
WorldForge teaser
WorldForge: Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
CVPR 2026 — Highlight
page|pdf|code
FewarNet teaser
FewarNet: An efficient few-shot view synthesis network based on trend regularization
IEEE Transactions on Circuits and Systems for Video Technology, 2024
pdf
SwitchCraft teaser
SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
CVPR 2026
page|pdf|code
Free-Lunch teaser
Free-Lunch Long Video Generation via Layer-Adaptive O.O.D Correction
CVPR 2026
pdf|code
Wide-baseline view synthesis teaser
Wide-baseline view synthesis for light-field display based on plane-depth-fused sweep volume
Displays, 2023
pdf
FlowDirector teaser
FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
CVPR 2026
page|pdf|code
Fast3Dcache teaser
Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration
CVPR 2026
page|pdf|code
AppAgentX teaser
AppAgentX: Evolving GUI agents as proficient smartphone users
arXiv preprint, 2025
page|pdf|code
DyWeight teaser
DyWeight: Dynamic Gradient Weighting for Few-Step Diffusion Sampling
arXiv preprint, 2026
pdf|code
StyleAvatar3D teaser
StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation
IEEE Journal of Selected Topics in Signal Processing (J-STSP), 2026
pdf|code
Elemental image array teaser
Elemental image array generation based on BVH structure combined with spatial partition and display optimization
Displays, 2024
pdf
Light field display teaser
Efficiently enhancing co-occurring details while avoiding artifacts for light field display
Applied Optics, 2020
pdf