activity
20242026
collaborators

5 papers

cs.RO2026

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Kehan Li, Bohan Hou, Minghao Zhu +28

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, Ry…

cs.CV2026

MemPose: Category-level Object Pose Estimation with Memory

Xiao Lin, Minghao Zhu, Yun Peng +4

In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that learn effective representations from data…

cs.CV2026

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition

Minghao Zhu, Xiao Lin, Mengxian Hu +5

Adapting CLIP for open-vocabulary video recognition necessitates a delicate balance between newly acquired video knowledge and the pretrained generalization. While existing studies…

cs.CV2025

CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation

Xiao Lin, Yun Peng, Liuyi Wang +6

Category-level object pose estimation aims to recover the rotation, translation and size of unseen instances within predefined categories. In this task, deep neural network-based m…

cs.CV2024

MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer

Minghao Zhu, Zhengpu Wang, Mengxian Hu +5

Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules…