From the 1 of 12 linked papers with an AI index.
12 papers
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation
Yiming Wang, Ye Chen, Hanqi Chen +1
Multimodal large models are increasingly used to generate scalable vector graphics (SVG), but reliable evaluation remains underexplored. Existing protocols are often code-centric o…
UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
Liming Tan, Ye Chen, Hao Zhang +3
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occl…
World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration
Ye Chen, Xuanhong Chen, Yupeng Zhu +23
The paper proposes the World Narrative Model, a framework that separates the specification of a 4D physical scene (geometry, motion, camera, lighting) from pixel generation, enabli…
VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
Weijie Wang, Yeqing Chen, Zeyu Zhang +9
Feed-forward 3D Gaussian Splatting (3DGS) has emerged as a highly effective solution for novel view synthesis. Existing methods predominantly rely on a \emph{pixel-aligned} Gaussia…
ProxyImg: Towards Highly-Controllable Image Representation via Hierarchical Disentangled Proxy Embedding
Ye Chen, Yupeng Zhu, Xiongzhen Zhang +4
Prevailing image representation methods, including explicit representations such as raster images and Gaussian primitives, as well as implicit representations such as latent images…
3DProxyImg: Controllable 3D-Aware Animation Synthesis from Single Image via 2D-3D Aligned Proxy Embedding
Yupeng Zhu, Xiongzhen Zhang, Ye Chen +1
3D animation is central to modern visual media, yet traditional production pipelines remain labor-intensive, expertise-demanding, and computationally expensive. Recent AIGC-based a…