1.2k citations · 1.7k across the 63 of their papers we have counts for
19 papers · 1 filter
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs
Junzhou Chen, Jindong Wang, Gang Zhou
Vision-language models and vision-language action models endow the robot with unprecedented capabilities. However, the input of video and high-resolution images yields a massive nu…
LatentUMM: Dual Latent Alignment for Unified Multimodal Models
Yinyi Luo, Wenwen Wang, Hayes Bai +2
Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency…
Self-Corrected Image Generation with Explainable Latent Rewards
Yinyi Luo, Hrishikesh Gokhale, Marios Savvides +2
Despite significant progress in text-to-image generation, aligning outputs with complex prompts remains challenging, particularly for fine-grained semantics and spatial relations.…
Self-Ensemble Post Learning for Noisy Domain Generalization
Wang Lu, Jindong Wang
While computer vision and machine learning have made great progress, their robustness is still challenged by two key issues: data distribution shift and label noise. When domain ge…
Exploring Scale Shift in Crowd Localization under the Context of Domain Generalization
Juncheng Wang, Lei Shang, Ziqi Liu +5
Crowd localization plays a crucial role in visual scene understanding towards predicting each pedestrian location in a crowd, thus being applicable to various downstream tasks. How…
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
Chika Maduabuchi, Hao Chen, Yujin Han +1
Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where…