From the 1 of 14 linked papers with an AI index.
14 papers
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Yunfeng Liu, Yuandong Yang, Jiarui Han +5
The paper introduces VIABench, a video benchmark built from first‑person recordings by visually impaired users to evaluate multimodal large language models on tasks like proactive…
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
Jiange Yang, Yansong Shi, Haoyi Zhu +6
Unsupervised learning of latent motion from Internet videos is crucial for robot learning. Existing discrete methods generally mitigate the shortcut learning caused by extracting e…
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
Jiaming Zhang, Shengming Cao, Rui Li +8
Preserving first-frame identity while ensuring precise motion control is a fundamental challenge in human image animation. The Image-to-Motion Binding process of the dominant Refer…
SAM 2++: Tracking Anything at Any Granularity
Jiaming Zhang, Cheng Liang, Yichun Yang +7
Due to the varying granularity of target states across different tasks, most existing trackers are tailored to a single task, which specificity limits their generalization, prevent…
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
Boyue Xu, Ruichao Hou, Tongwei Ren +1
UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream methods suffer from isolated feat…
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
Cheng Liang, Haoxian Chen, Liang Hou +4
The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal a…