3 papers
cs.CV2026
CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
Quang Minh Dinh, Tuan Kiet Doan
Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation history and textual descriptions.…
cs.CL2025
BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
Paige TuttösÃ, Mantaj Dhillon, Luna Sang +6
Some speech recognition tasks, such as automatic speech recognition (ASR), are approaching or have reached human performance in many reported metrics. Yet, they continue to struggl…
cs.CV2024
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
Quang Minh Dinh, Minh Khoi Ho, Anh Quan Dang +1
Traffic video description and analysis have received much attention recently due to the growing demand for efficient and reliable urban surveillance systems. Most existing methods…