works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators

9 papers

cs.CV2026

MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

Ke Cheng, Hanqiao Ye, Lei Shi +8

The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…

cs.CV2026

VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning

Li-Heng Chen, Ke Cheng, Yahui Liu +3

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…

cs.CV2026

Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning

Junhao Xiao, Zhiyu Wu, Hao Lin +5

Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing me…

cs.CV2026

Kelix Technical Report

Boyang Ding, Chenglong Chu, Dunju Zang +28

Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which u…

cs.CV2026

A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models

Weixin Ye, Wei Wang, Yahui Liu +5

In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Visi…

cs.RO2025

The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning

Titong Jiang, Xuefeng Jiang, Yuan Ma +7

We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in exe…