2 papers
cs.CV2026
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
Yibo Yuan, Jiacheng Fu, Jiangtong Zhu +8
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalabilit…
cs.AI2026
SportD: How do VLMs physically strategize?
Jasin Cekinmez, Addison J. Wu, Haotian Xia +11
The paper introduces SportD, a benchmark that tests whether vision‑language models can choose optimal shoot or pass actions in soccer situations, comparing model choices to a value…