collaborators

5 papers

cs.CV2025

Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos

Shingo Yokoi, Kento Sasaki, Yu Yamaguchi

Recent advances in end-to-end (E2E) autonomous driving have been enabled by training on diverse large-scale driving datasets, yet autonomous driving models still struggle in out-of…

cs.LG2025

Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness

Tsubasa Takahashi, Shojiro Yamabe, Futa Waseda +1

Differential Attention (DA) has been proposed as a refinement to standard attention, suppressing redundant or noisy context through a subtractive structure and thereby reducing con…

cs.CV2025

STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes

Keishi Ishihara, Kento Sasaki, Tsubasa Takahashi +2

Vision-Language Models (VLMs) have been applied to autonomous driving to support decision-making in complex real-world scenarios. However, their training on static, web-sourced ima…

cs.CV2025

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

Fanheng Kong, Jingyuan Zhang, Hongzhi Zhang +7

Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing ben…

cs.CV2025

One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression

Keita Miwa, Kento Sasaki, Hidehisa Arai +2

Current image tokenization methods require a large number of tokens to capture the information contained within images. Although the amount of information varies across images, mos…