collaborators

13 papers

cs.CV2026

EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass

Haomin Peng, Yongkang Li, Zhaoxiang Liu +4

Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and achieves competitive open-vocab…

cs.CV2026

ReWorld: Representation Learning for World Action Models

Tianze Xia, Lijun Zhou, Kaixin Xiong +9

World Action Models (WAMs) unify future environment prediction with action generation for autonomous driving, yet existing approaches optimize only the final outputs, leaving inter…

cs.CV2026

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World

Tianze Xia, Yongkang Li, Lijun Zhou +9

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approa…

cs.CV2026

DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

Lunbin Zeng, Jingfeng Yao, Bencheng Liao +3

Diffusion-based decoding has recently emerged as an appealing alternative to autoregressive (AR) generation, offering the potential to update multiple tokens in parallel and reduce…

cs.CL2026

Mixture-of-Depths Attention

Lianghui Zhu, Yuxin Fang, Bencheng Liao +10

Scaling depth is a key driver for large language models (LLMs). Yet, as LLMs become deeper, they often suffer from signal degradation: informative features formed in shallow layers…

cs.CV2025

DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis

Yueting Zhu, Yuehao Song, Shuai Zhang +2

Whole Slide Images (WSIs) are typically analyzed using multiple instance learning (MIL) methods. However, the scale and heterogeneity of WSIs generate highly redundant and disperse…