activity
20212026
most citedDetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

64 citations · 136 across the 43 of their papers we have counts for

collaborators

38 papers

cs.CV2026

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

Yinming Huang, Shuyuan Tu, Xi Yan +6

Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual…

cs.RO2026

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

Zhihao Zhu, Hanlin Shang, Mingwang Xu +6

Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high comp…

cs.CV2026

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

Jiacheng Fu, Yibo Yuan, Meng Tian +8

Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. Howeve…

cs.CV2026

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

Yibo Yuan, Jiacheng Fu, Jiangtong Zhu +8

End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalabilit…

cs.CV2026

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools

Xiuwei Chen, Quanlin Chen, Wentao Hu +8

Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) paradigm by iteratively perfor…

cs.CV2026

Latent Visual States for Efficient Multimodal Reasoning

Xiuwei Chen, Wentao Hu, Yongxin Wang +8

The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete out…