From the 1 of 52 linked papers with an AI index.
52 papers
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
Fan Zhang, Guangming Yao, Jinyang Wu +6
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approx…
ATOM: Geometry-Aware Microgesture towards Object-Agnostic Tangible Interaction
Yinqiao Wang, Hao Xu, Qixuan Liu +3
This paper presents ATOM, an integrated framework towards agnostic and tangible object interactions with microgestures. Our goal is to support microgesture interactions across diff…
Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
Zhongyao Wang, Wanli Ouyang, Taoyong Cui +1
Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can rewar…
LiverPlan: A Stage-Adaptive Immersive Visual Analytics Framework for Anatomical Liver Surgical Planning
Qixuan Liu, Shi Qiu, Xiwen Wu +7
Anatomical liver resection (ALR) surgery is the most important treatment for liver cancer, yet preoperative planning demands complex, multi-stage clinical reasoning under competing…
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Haodong Li, Tianfei Ren, Xiaoxiao Ma +25
The paper presents VideoCoCo, a system that generates physically consistent videos by having a coding agent produce executable Blender code that defines the scene and its dynamics,…
ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
Qiao Yan, Yihan Wang, Zhenghao Xing +2
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or…