8 papers
Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches
Yitao Jiang, Yaoqing Yang, Luyang Zhao +2
Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results suggest that the channel basis of each k…
ASCII Art Turns LLMs into VLA Controllers
Yitao Jiang, Roy Xing, Luyang Zhao +3
Vision--Language--Action (VLA) controllers are often built by extending vision--language models (VLMs) with action supervision, relying on multimodal backbones with large data and…
Manipulider: A Multi-Engine Buoyancy-Controlled Robot for Thrusterless Underwater Gliding and Manipulation
Yitao Jiang, Yewei Huang, Weizhi Cao +5
The Manipulider is a buoyancy-actuated underwater robot that enables thrusterless, glide-like locomotion and attitude-based manipulation, while providing a magnetic modular interfa…
FlowMo-WM: A World Model with Object Momentum and Hidden Ambient Drift
Yitao Jiang, Luyang Zhao, Muhao Chen +1
World models in robot learning predict future states from visual observations and actions, enabling agents to reason about the consequences of their controls. However, many action-…
SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles
Mingi Jeong, Arihant Chadda, Ziang Ren +8
This paper introduces the first publicly accessible labeled multi-modal perception dataset for autonomous maritime navigation, focusing on in-water obstacles within the aquatic env…
An Untethered Bioinspired Robotic Tensegrity Dolphin with Multi-Flexibility Design for Aquatic Locomotion
Luyang Zhao, Yitao Jiang, Chun-Yi She +5
This paper presents the first steps toward a soft dolphin robot using a bio-inspired approach to mimic dolphin flexibility. The current dolphin robot uses a minimalist approach, wi…