7 papers
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
Junhao Chen, Mingjin Chen, Jingjia Mao +12
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been meas…
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
Junhao Chen, Mingjin Chen, Henghaofan Zhang +10
Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D ge…
PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation
Junhao Chen, Xiang Li, Mingjin Chen +9
Code is the medium through which large language models generate structured artifacts: charts, scientific figures, vector graphics, CAD models, 3D scenes, and hardware designs are a…
One Video, One World: Turning Monocular Video into Physical 4D Scenes
Junhao Chen, Boran Zhang, Mingjin Chen +7
We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconst…
LottieGPT: Tokenizing Vector Animation for Autoregressive Generation
Junhao Chen, Kejun Gao, Yuehan Cui +8
Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multimedia on the Internet. Vector…
HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
Mingjin Chen, Junhao Chen, Zhaoxin Fan +6
Recent methods have made notable progress in the visual quality of hand-object interaction video synthesis. However, most approaches rely on 2D control signals that lack spatial ex…