1 citations · 1 across the 7 of their papers we have counts for
10 papers
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
Simon Holk, Ryosuke Takanami, Tatsuya Matsushima +4
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with…
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
Ziang Yan, Sheng Xia, Jiashuo Yu +10
Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant se…
See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs
Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often inf…
Continuous Reasoning for Vision-Language-Action
Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals o…
MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models
Yue Wu, Changyuan Wang, Zixuan Wang +2
Conventional Post-Training Quantization (PTQ) methods struggle with 4-bit Omni-modal Large Language Models (OLLMs) due to the extreme distribution heterogeneity and disparate outli…
Concurrent Prehensile and Nonprehensile Manipulation: A Practical Approach to Multi-Stage Dexterous Tasks
Hao Jiang, Yue Wu, Yue Wang +2
Dexterous hands enable concurrent prehensile and nonprehensile manipulation, such as holding one object while interacting with another, a capability essential for everyday tasks ye…