collaborators

19 papers

cs.CV2026

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

Junpeng Jing, Ronglai Zuo, Zhelun Shen +5

Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computation, or additional foundation-model priors, making them difficul…

cs.CV2026

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation

Chen Si, Yulin Liu, Bo Ai +4

We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation. While recent works with foundation approaches have shown th…

cs.CV2026

Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation

Chrysa Pratikaki, Pablo Ruiz-Ponce, Jiankang Deng +2

Recent advances in 4D Human-Object Interaction (HOI) generation have enabled increasingly realistic motion synthesis, particularly for single-object manipulation. Yet current resea…

cs.CV2026

StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video

Huajian Zeng, Chaohua Yao, Yuantai Zhang +3

Recovering world space 4D motion of two interacting hands from egocentric video is a fundamental capability for supervising robot policy learning, where wrist trajectories track th…

cs.CV2026

HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching

Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +3

Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plau…

cs.CV2026

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

Yura Choi, Roy Miles, Rolandos Alexandros Potamias +3

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Model…