activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

Yiran Guan, Sifan Tu, Dingkang Liang +6

Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel…

cs.CV2026

UniFuture: A 4D Driving World Model for Future Generation and Perception

Dingkang Liang, Dingyuan Zhang, Xin Zhou +7

We present UniFuture, a unified 4D Driving World Model designed to simulate the dynamic evolution of the 3D physical world. Unlike existing driving world models that focus solely o…

cs.CV2026

The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey

Sifan Tu, Xin Zhou, Dingkang Liang +4

The Driving World Model (DWM), which focuses on predicting scene evolution during the driving process, has emerged as a promising paradigm in the pursuit of autonomous driving (AD)…

cs.CV2025

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Xin Zhou, Dingkang Liang, Sifan Tu +6

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to inc…

cs.CV2024

A Unified Framework for 3D Scene Understanding

Wei Xu, Chunsheng Shi, Sifan Tu +3

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a…