25 papers
StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
Xiao Liu, Yuguang Yang, Xi Wang +6
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…
Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features
Sining Ang, Yuguang Yang, Yan Wang
Large video diffusion models provide rich spatiotemporal priors for autonomous driving, but existing world-action models often inherit the cost of iterative future-video generation…
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation
Jinting Wang, Yuguang Yang, Shengyu Li +4
Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated…
InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection
Zichao Feng, Haodong Zhu, Jingying Yang +8
Robust object detection under adverse visual conditions remains a long-standing challenge for multi-modal perception systems. Existing fusion-based methods typically require both R…
Teaching Vision-Language-Action Models What to See and Where to Look
Yuguang Yang, Canyu Chen, Zhewen Tan +10
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual q…
CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection
Zihan Liu, Yuguang Yang, Shengjie Su +5
Continual Object Detection (COD) requires a detector to acquire new categories over time while preserving previously learned ones. This goal is closely related to open-vocabulary d…