From the 1 of 14 linked papers with an AI index.
14 papers
StreamFlow: Dynamic Memory Flows for Streaming Video Understanding
Muxin Fu, Yifan Zhang, Wentao Zhang +5
Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded…
WorkDrive: Roadwork Chain of Causation for Autonomous Driving
Tianyi Jiang, Wen Zhang, Sihan Yang +2
The paper introduces WorkDrive, a framework that adds perception‑grounded causal reasoning to vision‑language models for autonomous driving in roadwork zones, improving trajectory…
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
Ruichuan An, Kai Zeng, Ming Lu +5
Vision-Language Models (VLMs) have demonstrated exceptional performance in various multi-modal tasks. Recently, there has been an increasing interest in improving the personalizati…
SteerVTE: Seamless Video Text Editing with Style and Glyph Control
Kai Zeng, Moran Li, Zhengwei Wang +6
Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain,…
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-AI, Anyi Xu, Bangcai Lin +315
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…
PEARL: Personalized Streaming Video Understanding Model
Yuanhong Zheng, Ruichuan An, Xiaopeng Lin +10
Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimoda…