From the 1 of 9 linked papers with an AI index.
9 papers
MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
Ke Cheng, Hanqiao Ye, Lei Shi +8
The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…
VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning
Li-Heng Chen, Ke Cheng, Yahui Liu +3
Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
Junhao Xiao, Zhiyu Wu, Hao Lin +5
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing me…
Kelix Technical Report
Boyang Ding, Chenglong Chu, Dunju Zang +28
Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which u…
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
Weixin Ye, Wei Wang, Yahui Liu +5
In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Visi…
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
Titong Jiang, Xuefeng Jiang, Yuan Ma +7
We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in exe…