23 papers
Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation
Shuitsu Koyama, Kazuki Matsuda, Yuiga Wada +3
Automatic evaluation of image and video captioning is essential for benchmarking multimodal systems, although standard evaluation metrics show limited alignment with human judgment…
MLLM-as-a-Judge Exhibits Model Preference Bias
Shuitsu Koyama, Yuiga Wada, Daichi Yashima +1
Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model performance. If such MLLM-as-a-J…
Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation
Hina Kogure, Kei Katsumata, Taiki Miyanishi +1
Dynamic urban environments are often captured by cameras placed at spatially separated locations with little or no view overlap. However, most existing 4D reconstruction methods as…
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
Daichi Yashima, Koki Seno, Shuhei Kurita +2
Coarse-to-fine autoregressive modeling has recently shown strong promise for visuomotor policy learning, combining the inference efficiency of autoregressive methods with the globa…
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation
Motonari Kambara, Koki Seno, Tomoya Kaichi +2
We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only…
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
Yusuke Takagi, Motonari Kambara, Daichi Yashima +3
In this study, we address the problem of language-guided robotic manipulation, where a robot is required to manipulate a wide range of objects based on visual observations and natu…