18 papers
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
Xiaogang Xu, Jiaqi Tang, Jianmin Chen +12
Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction.…
From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation
Zhefan Rao, Bin Zou, Xuanhua He +5
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-…
USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes
Li-Heng Chen, Haokai Pang, Chengye Su +7
Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-or…
MSEditor: Toward Consistent Multi-Shot Video Editing
Kunyu Feng, Yue Ma, Bingyuan Wang +6
In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos…
Data Pyramid for Embodied Manipulation: A Survey
Yifan Ye, Yankai Fu, Yaoxu Lv +26
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…
LiveLight: Real-time Streaming Video Relighting with Interactive Control
Yue Ma, Jiangming Wang, Yucheng Wang +8
We present LiveLight, the first diffusion-based framework for real-time streaming video relighting with interactive 3D lighting control. Achieving this is non-trivial, as it requir…