collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

PhiZero: A World Model Built Around Physical Language

Shuyao Shang, Yuqi Wang, Ruopeng Gao +4

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically…

cs.CV2026

UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer

Shuai Wang, Liang Li, Yang Chen +3

Unified Multimodal Models (UMMs) have emerged as a critical direction for general-purpose multimodal intelligence, integrating understanding and generation into a single framework.…

cs.CV2025

Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval

Chunxu Liu, Jiyuan Yang, Ruopeng Gao +4

Multimodal embeddings are widely used in downstream tasks such as multimodal retrieval, enabling alignment of interleaved modalities in a shared representation space. While recent…

cs.CV2025

History-Aware Transformation of ReID Features for Multiple Object Tracking

Ruopeng Gao, Yuyao Wang, Chunxu Liu +1

In Multiple Object Tracking (MOT), Re-identification (ReID) features are widely employed as a powerful cue for object association. However, they are often wielded as a one-size-fit…

cs.CV2024

Multiple Object Tracking as ID Prediction

Ruopeng Gao, Ji Qi, Limin Wang

Multi-Object Tracking (MOT) has been a long-standing challenge in video understanding. A natural and intuitive approach is to split this task into two parts: object detection and a…

cs.CV2023

MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking

Ruopeng Gao, Limin Wang

As a video task, Multiple Object Tracking (MOT) is expected to capture temporal information of targets effectively. Unfortunately, most existing methods only explicitly exploit the…