4 papers
KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
Yaping Li, Zhaxizhuoma, Qiaojun Yu +3
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-A…
Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
Ruichao Mao, Zhou Fang, Teng Guo +9
User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large lan…
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
Weizhe Chen, Miao Zhang, Junpeng Jiang +3
Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality, making hybrid architecture des…
InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
Yang Tian, Yuyin Yang, Yiman Xie +13
Recent works explore how real and synthetic data contribute to Vision-Language-Action (VLA) models' generalization. While current VLA models have shown the strong effectiveness of…