4 papers
Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches
Yitao Jiang, Yaoqing Yang, Luyang Zhao +2
Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results suggest that the channel basis of each k…
ASCII Art Turns LLMs into VLA Controllers
Yitao Jiang, Roy Xing, Luyang Zhao +3
Vision--Language--Action (VLA) controllers are often built by extending vision--language models (VLMs) with action supervision, relying on multimodal backbones with large data and…
Manipulider: A Multi-Engine Buoyancy-Controlled Robot for Thrusterless Underwater Gliding and Manipulation
Yitao Jiang, Yewei Huang, Weizhi Cao +5
The Manipulider is a buoyancy-actuated underwater robot that enables thrusterless, glide-like locomotion and attitude-based manipulation, while providing a magnetic modular interfa…
FlowMo-WM: A World Model with Object Momentum and Hidden Ambient Drift
Yitao Jiang, Luyang Zhao, Muhao Chen +1
World models in robot learning predict future states from visual observations and actions, enabling agents to reason about the consequences of their controls. However, many action-…