works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

Ke Cheng, Hanqiao Ye, Lei Shi +8

The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…

cs.CV2026

VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning

Li-Heng Chen, Ke Cheng, Yahui Liu +3

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…

cs.CV2026

Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning

Junhao Xiao, Zhiyu Wu, Hao Lin +5

Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing me…

cs.CV2026

Kelix Technical Report

Boyang Ding, Chenglong Chu, Dunju Zang +28

Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which u…

cs.CV2026

A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models

Weixin Ye, Wei Wang, Yahui Liu +5

In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Visi…

cs.CV2025

SEGT: A General Spatial Expansion Group Transformer for nuScenes Lidar-based Object Detection Task

Cheng Mei, Hao He, Yahui Liu +1

In the technical report, we present a novel transformer-based framework for nuScenes lidar-based object detection task, termed Spatial Expansion Group Transformer (SEGT). To effici…