2 citations · 2 across the 2 of their papers we have counts for
4 papers
AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps
Liaoyuan Fan, Zetian Xu, Chen Cao +3
Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial r…
DINO-Tok: Adapting DINO for Visual Tokenizers
Mingkai Jia, Mingxiao Li, Zhijian Shu +12
Recent advances in visual generation have emphasized the importance of Latent Generative Models (LGMs), which critically depend on effective visual tokenizers to bridge pixels and…
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
Fuhao Li, Huan Jin, Bin Gao +3
Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing dat…
SCBench: A Sports Commentary Benchmark for Video LLMs
Kuangzhi Ge, Lingjun Chen, Kevin Zhang +6
Recently, significant advances have been made in Video Large Language Models (Video LLMs) in both academia and industry. However, methods to evaluate and benchmark the performance…