2 papers
cs.CV2025
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
Zixuan Dong, Baoyun Peng, Yufei Wang +4
Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video…
cs.CV2025
LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
Xiaodong Wang, Zhirong Wu, Peixi Peng
Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error ac…