activity
20232026
most citedDriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model

18 citations · 46 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CV2026

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

Xianzhe Fan, Shengliang Deng, Xiaoyang Wu +7

Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D informa…

cs.RO2025

Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping

Zhuoling Li, Jinrong Yang, Yong Zhao +4

Realizing generalizable dynamic object manipulation on conveyor systems is important for enhancing manufacturing efficiency, as it eliminates specialized engineering for different…

cs.CV2025

DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation

Yunhan Yang, Shuo Chen, Yukun Huang +6

Recent advancements in leveraging pre-trained 2D diffusion models achieve the generation of high-quality novel views from a single in-the-wild image. However, existing works face c…

cs.CV2024★ 1 cited

UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation

Lihe Yang, Zhen Zhao, Hengshuang Zhao

Semi-supervised semantic segmentation (SSS) aims at learning rich visual knowledge from cheap unlabeled images to enhance semantic segmentation capability. Among recent works, UniM…

cs.RO2024

VIP: Vision Instructed Pre-training for Robotic Manipulation

Zhuoling Li, Liangliang Ren, Jinrong Yang +5

The effectiveness of scaling up training data in robotic manipulation is still limited. A primary challenge in manipulation is the tasks are diverse, and the trained policy would b…

cs.CV2024

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Kai Chen, Yunhao Gou, Runhui Huang +28

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language…