collaborators

13 papers

cs.CV2026

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Ziyi Wang, Xumin Yu, Yongming Rao +19

Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situatio…

cs.CV2026

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

Xumin Yu, Zuyan Liu, Zhenyu Yang +5

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete s…

cs.CV2026

GEM: Generative Supervision Helps Embodied Intelligence

Ruowen Zhao, Bangguo Li, Zuyan Liu +9

Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a si…

cs.CV2026

HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Tencent Robotics X, HY Vision Team, : +20

We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) an…

cs.CV2026

PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning

Shaoxuan Li, Zhixuan Zhao, Hanze Deng +9

We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is su…

cs.AI2026

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

Yang Shi, Yuhao Dong, Yue Ding +22

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question rem…