2 citations · 3 across the 59 of their papers we have counts for
32 papers · 1 filter
Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models
Xiaoyang Guo, Guoping Luo, Jusheng Zhang +2
Adapting video vision-language models (VLMs) is computationally expensive because video inputs produce a large number of visual tokens, making both fine-tuning and inference costly…
ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models
Yihao Wang, Zijian He, Jie Ren +1
Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts. How…
LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
Jinzhou Tang, Sidi Liu, Waikit Xiu +2
A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critica…
The Fourth Challenge on Image Super-Resolution (4) at NTIRE 2026: Benchmark Results and Method Overview
Zheng Chen, Kai Liu, Jingkai Wang +150
This paper presents the NTIRE 2026 image super-resolution (4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to r…
DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration
Jinzhou Tang, Fan Feng, Minghao Fu +3
Learned world models excel at interpolative generalization but fail at extrapolative generalization to novel physical properties. This limitation arises because they learn statisti…
Process-of-Thought Reasoning for Videos
Jusheng Zhang, Kaitong Cai, Jian Wang +3
Video understanding requires not only recognizing visual content but also performing temporally grounded, multi-step reasoning over long and noisy observations. We propose Process-…