works on

From the 1 of 20 linked papers with an AI index.

collaborators

20 papers

cs.SD2026

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Junhao Chen, Mingjin Chen, Jingjia Mao +12

Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been meas…

cs.CV2026

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

Junhao Chen, Mingjin Chen, Henghaofan Zhang +10

Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D ge…

cs.CV2026

Bunraku: Turning a Single Illustration into an Editable Live2D Character

Junhao Chen, Jingjia Mao, Dayong Li +6

The paper introduces Bunraku, a system that automatically creates a complete Live2D character—including layered RGBA images, deformation meshes, and animation keyposes—from a singl…

cs.CV2026

Engine-Native Editable 3D World Reconstruction with Objects and Lighting

Junhao Chen, Xinghao Chen, Henghaofan Zhang +8

Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-image methods largely stop at…

cs.CL2026

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation

Junhao Chen, Xiang Li, Mingjin Chen +9

Code is the medium through which large language models generate structured artifacts: charts, scientific figures, vector graphics, CAD models, 3D scenes, and hardware designs are a…

cs.CV2026

One Video, One World: Turning Monocular Video into Physical 4D Scenes

Junhao Chen, Boran Zhang, Mingjin Chen +7

We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconst…