works on

From the 1 of 14 linked papers with an AI index.

collaborators

14 papers

cs.CV2026

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

Muxin Fu, Yifan Zhang, Wentao Zhang +5

Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded…

cs.CV2026

WorkDrive: Roadwork Chain of Causation for Autonomous Driving

Tianyi Jiang, Wen Zhang, Sihan Yang +2

The paper introduces WorkDrive, a framework that adds perception‑grounded causal reasoning to vision‑language models for autonomous driving in roadwork zones, improving trajectory…

cs.CV2026

Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs

Ruichuan An, Kai Zeng, Ming Lu +5

Vision-Language Models (VLMs) have demonstrated exceptional performance in various multi-modal tasks. Recently, there has been an increasing interest in improving the personalizati…

cs.CV2026

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

Kai Zeng, Moran Li, Zhengwei Wang +6

Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain,…

cs.CL2026

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek-AI, Anyi Xu, Bangcai Lin +315

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…

cs.CV2026

PEARL: Personalized Streaming Video Understanding Model

Yuanhong Zheng, Ruichuan An, Xiaopeng Lin +10

Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimoda…