4 papers
CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding
Yuri Ishitoya, Jeremy Siburian, Masashi Hamaya +3
Vision-language-action models (VLAs) inherit semantic capabilities from pretrained VLMs, yet large-scale post-training on robot data and architectural modifications can reshape the…
Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
Rikuto Kotoge, Mai Nishimura, Jiaxin Ma
Reinforcement Learning has emerged as a dominant post-training approach to elicit agentic RAG behaviors such as search and planning from language models. Despite its success with l…
Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft Wrist
Tatsuya Kamijo, Mai Nishimura, Cristian C. Beltran-Hernandez +2
Tactile memory, the ability to store and retrieve touch-based experience, is critical for contact-rich tasks such as key insertion under uncertainty. To replicate this capability,…
PLASMA: A Layout-Aware Benchmark Reveals Memory Layout Matters for Graph-based ANNS on GPU
Yutaro Oguri, Mai Nishimura, Yusuke Matsui
We propose a latform for ayout-ware earch and emory rrangement (), a unified evaluation fra…