collaborators

6 papers

cs.AI2026

IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

Lingyi Meng, Zecong Tang, Haoran Li +12

In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework. Mastering dynamic, multi-turn interleaved ima…

cs.AI2026

Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving

Weitong Lian, Zecong Tang, Haoran Li +12

Autonomous driving is an important and safety-critical task, and recent advances in LLMs/VLMs have opened new possibilities for reasoning and planning in this domain. However, larg…

cs.AI2026

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving

Zecong Tang, Zixu Wang, Yifei Wang +10

Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reasoning and generalization abiliti…

cs.RO2026

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action

Yi Zhang, Yinda Chen, Che Liu +26

We present Pelican-Unify 1.0, the first embodied foundation model trained according to the principle of unification. Pelican-Unify 1.0 uses a single VLM as a unified understanding…

cs.CV2026

Imagine a City: CityGenAgent for Procedural 3D City Generation

Zishan Liu, Zecong Tang, RuoCheng Wu +6

The automated generation of interactive 3D cities is a critical challenge with broad applications in autonomous driving, virtual reality, and embodied intelligence. While recent ad…

cs.CL2025

CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation

Yuyang Jiang, Chacha Chen, Shengyuan Wang +8

Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptim…