7 papers
BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering
Yi-Ruei Liu, Jie-Ying Lee, Zheng-Hui Huang +2
Inverse rendering of urban scenes from captured videos enables numerous applications, including content creation and autonomous driving simulation. Physically-based rendering metho…
OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning
Maonan Wang, Zhengyan Huang, Kemou Jiang +13
Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. Howev…
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
Chengwen Liu, Xiaomin Yu, Zhuoyue Chang +15
In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore ne…
Moaw: Unleashing Motion Awareness for Video Diffusion Models
Tianqi Zhang, Ziyi Wang, Wenzhao Zheng +5
Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks suc…
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
Zhe Huang, Hao Wen, Aiming Hao +6
Multimodal Large Language Models (MLLMs) have made remarkable progress in video understanding. However, they suffer from a critical vulnerability: an over-reliance on language prio…
Towards Fine-Grained Human Motion Video Captioning
Guorui Song, Guocun Wang, Zhe Huang +4
Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motio…