16 papers
MoGeFlow: Flowing Through Motion Codebook Geometry for Text-to-Motion Generation
Pengcheng Fang, Tengjiao Sun, Xiaoyu Zhan +2
Vector-quantized motion tokenizers provide a compact discrete interface for text-to-motion generation, but most motion-code priors treat code indices as unordered categorical label…
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
Pengcheng Fang, Yuxia Chen, Xiaohao Cai
Video temporal grounding (VTG) aims to localize the start and end timestamps of the event described by a given query within an untrimmed video. Despite the strong open-world video…
AsyncLane: Decoupling Refinement from Advancement in Diffusion Language Model Decoding
Yingxuan Ren, Yuxuan Lou, Yong Liu +4
Block-wise semi-autoregressive decoding is the standard inference paradigm for diffusion large language models (DLMs), but it imposes a strict dependency between blocks: the next b…
KV-Control: Parameter-Efficient K/V Injection for Trajectory-Controlled Text-to-Motion
Tengjiao Sun, Pengcheng Fang, Xiaoyu Zhan +4
Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop at text: a character may nee…
SO-Mamba: State-Ownership Mamba for Unrolled MRI Reconstruction
Pengcheng Fang, Hongli Chen, Fangfang Tang +3
Accelerated MRI reconstruction requires recovering missing details while preserving anatomically coherent structures across large spatial regions. State-space models such as Mamba…
MotionDuet: Dual-Conditioned 3D Human Motion Generation with Video-Regularized Text Learning
Yi-Yang Zhang, Tengjiao Sun, Pengcheng Fang +4
3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work…