5 papers
Uncertainty-Aware Intention Prediction for Human-to-Robot Assembly Teleoperation
Fnu Heman, Yixuan Wang, Kolin Xu +6
In assisted teleoperation for human-robot collaboration, accurate intention prediction is critical for enabling timely and reliable robotic assistance during long-horizon manipulat…
VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track
Wenming Tu, Jian Gao, Yanru Huo +9
Audio reasoning requires multi-step, evidence-grounded inference over temporally dynamic and acoustically mixed signals, exceeding conventional perception tasks such as ASR or capt…
Effective Model Pruning: Measure The Redundancy of Model Components
Yixuan Wang, Dan P. Guralnik, Saiedeh Akbari +1
This article initiates the study of a basic question about model pruning. Given a vector of importance scores assigned to model components, how many of the scored components co…
Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
Chaoyang Wang, Yangfan He, Yiyang Zhou +6
We reveal a critical yet underexplored flaw in Large Vision-Language Models (LVLMs): even when these models know the correct answer, they frequently arrive there through incorrect…
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Shuhao Gu, Jialing Zhang, Siyuan Zhou +23
Recently, Vision-Language Models (VLMs) have achieved remarkable progress in multimodal tasks, and multimodal instruction data serves as the foundation for enhancing VLM capabiliti…