From the 1 of 9 linked papers with an AI index.
7 papers · 1 filter
MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
Ke Cheng, Hanqiao Ye, Lei Shi +8
The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…
VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning
Li-Heng Chen, Ke Cheng, Yahui Liu +3
Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
Junhao Xiao, Zhiyu Wu, Hao Lin +5
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing me…
Kelix Technical Report
Boyang Ding, Chenglong Chu, Dunju Zang +28
Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which u…
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
Weixin Ye, Wei Wang, Yahui Liu +5
In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Visi…
SEGT: A General Spatial Expansion Group Transformer for nuScenes Lidar-based Object Detection Task
Cheng Mei, Hao He, Yahui Liu +1
In the technical report, we present a novel transformer-based framework for nuScenes lidar-based object detection task, termed Spatial Expansion Group Transformer (SEGT). To effici…