8 papers
One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
Shijun Shi, Jing Xu, Zhihang Li +5
Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matc…
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
Junxin Wang, Dai Guan, Weijie Qiu +7
Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-time scaling. However, they often funct…
Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID
Jiaxuan Li, Xin Wen, Zhihang Li
Any-Time Person Re-identification (AT-ReID) necessitates the robust retrieval of target individuals under arbitrary conditions, encompassing both modality shifts (daytime and night…
ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation
Yinuo Liu, Zi Qian, Heng Zhou +7
Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. C…
Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models
Weijie Qiu, Dai Guan, Junxin Wang +6
Generative reward models (GRMs) for vision-language models (VLMs) often evaluate outputs via a three-stage pipeline: rubric generation, criterion-based scoring, and a final verdict…
OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
Jilei Mao, Jiarui Guan, Yingjuan Tang +7
The visuomotor policy can easily overfit to its training datasets, such as fixed camera positions and backgrounds. This overfitting makes the policy perform well in the in-distribu…