activity
20242026
collaborators

6 papers

cs.RO2026

Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations

Wei-Jin Huang, Yue-Yi Zhang, Yi-Lin Wei +5

Enabling humanoid robots to physically interact with humans is a critical frontier, but progress is hindered by the scarcity of high-quality Human-Humanoid Interaction (HHoI) data.…

cs.CV2025

LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning

Shenghao Fu, Qize Yang, Yuan-Ming Li +3

Long video understanding is still challenging for recent Large Video-Language Models (LVLMs) due to the conflict between long-form temporal understanding and detailed spatial perce…

cs.CV2025

Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks

Wei-Jin Huang, Yuan-Ming Li, Zhi-Wei Xia +4

Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering erro…

cs.CV2025

ViSpeak: Visual Instruction Feedback in Streaming Videos

Shenghao Fu, Qize Yang, Yuan-Ming Li +6

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent mod…

cs.RO2025

Task-Oriented 6-DoF Grasp Pose Detection in Clutters

An-Lan Wang, Nuo Chen, Kun-Yu Lin +2

In general, humans would grasp an object differently for different tasks, e.g., "grasping the handle of a knife to cut" vs. "grasping the blade to hand over". In the field of robot…

cs.CV2024

TechCoach: Towards Technical-Point-Aware Descriptive Action Coaching

Yuan-Ming Li, An-Lan Wang, Kun-Yu Lin +4

To guide a learner in mastering action skills, it is crucial for a coach to 1) reason through the learner's action execution and technical points (TechPoints), and 2) provide detai…