11 papers
PixelIR: Fidelity-Perception Decoupling via Pixel-Space Image-Residual Flow Matching for Efficient One-Step Real-World Super-Resolution
Bingtian Qiao, Yue Shi, Yong Guo +2
Real-world image super-resolution (Real-ISR) aims to preserve structures supported by the degraded observation while reconstructing perceptually realistic details. However, existin…
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
Dingyi Rong, Yue Shi, Chaofan Ma +6
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human ma…
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation
Dayu Xia, Yue Shi, Yao Mu +7
Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models t…
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
Bingtian Qiao, Yue Shi, Yingjie Zhou +3
Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inheri…
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
Jiameng Li, Minye Wu, Jiezhang Cao +2
Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tokens while sparse sampling ri…
EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans
Yingjie Zhou, Xilei Zhu, Siyu Ren +11
Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker…