1 citations · 1 across the 12 of their papers we have counts for
12 papers
Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control
Runji Cai, Toshihiko Yamasaki, Ling Xiao
Social robot navigation (SRN) requires more than geometric path planning; it demands understanding human intentions, social norms, and contextual cues to generate socially complian…
Video-Mirai: Autoregressive Video Diffusion Models Need Foresight
Yonghao Yu, Lang Huang, Runyi Li +2
Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segment becomes a commitment tha…
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
Ling Xiao, Toshihiko Yamasaki
Most fine-grained fashion image retrieval (FIR) methods assume a static setting, requiring full retraining when new attributes appear, which is costly and impractical for dynamic s…
Enhancing Lightweight Vision Language Models through Group Competitive Learning for Socially Compliant Navigation
Xinyu Zhang, Atsushi Konno, Toshihiko Yamasaki +1
Social robot navigation requires a sophisticated integration of scene semantics and human social norms. Scaling up Vision Language Models (VLMs) generally improves reasoning and de…
Reward Incremental Learning in Text-to-Image Generation
Maorong Wang, Jiafeng Mao, Xueting Wang +1
The recent success of denoising diffusion models has significantly advanced text-to-image generation. While these large-scale pretrained models show excellent performance in genera…
Dealing with Synthetic Data Contamination in Online Continual Learning
Maorong Wang, Nicolas Michel, Jiafeng Mao +1
Image generation has shown remarkable results in generating high-fidelity realistic images, in particular with the advancement of diffusion-based models. However, the prevalence of…