collaborators

6 papers

cs.CV2026

Cambrian-P: Pose-Grounded Video Understanding

Jihan Yang, Zifan Zhao, Xichen Pan +6

Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that relates observations across video frames. Yet this signal is large…

cs.CV2026

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

Junjie Wang, Xinghua Lou, Jason Li +8

Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm…

cs.CV2026

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

Bin Kang, Shaoguo Wen, Yang Fan +6

While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the structural mismatch between…

cs.CV2025

Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior

Yulin Li, Haokun Gui, Ziyang Fan +4

Recent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computat…

cs.CV2025

CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval

Bin Kang, Bin Chen, Junjie Wang +3

Existing Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggreg…

cs.CV2025

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception

Junjie Wang, Bin Chen, Yulin Li +3

Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbou…