activity
20242026
collaborators

9 papers

cs.CV2026

EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video

Yuan Zeng, Yujia Shi, Tiao Tan +6

Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Exist…

cs.CV2026

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

Yuan Zeng, Yujia Shi, Yuhao Yang +4

Human image animation aims to generate a video from a static reference image, guided by pose information extracted from a driving video. Existing approaches often rely on pose esti…

cs.CV2025

Efficient Feature Aggregation and Scale-Aware Regression for Monocular 3D Object Detection

Yifan Wang, Xiaochen Yang, Fanqi Pu +2

Monocular 3D object detection has attracted great attention due to simplicity and low cost. Existing methods typically follow conventional 2D detection paradigms, first locating ob…

cs.CV2025

TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs

Zhehan Kan, Yanlin Liu, Kun Yin +8

DeepSeek R1 has significantly advanced complex reasoning for large language models (LLMs). While recent methods have attempted to replicate R1's reasoning capabilities in multimoda…

cs.CV2025

UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval

Yating Liu, Yaowei Li, Xiangyuan Lan +3

Text-based Person Retrieval (TPR) as a multi-modal task, which aims to retrieve the target person from a pool of candidate images given a text description, has recently garnered co…

cs.CV2025

DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval

Yating Liu, Zimo Liu, Xiangyuan Lan +3

Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person…