8 papers
READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations
Bo Fang, Xinyao Zhang, Yuxin Song +3
Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on promptin…
CogniRoute: Learning to Route Social Evidence in Omni-Modal Models
Yifan Shen, Pei Tian, Xinzhuo Li +8
Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee that a model uses the right evidence. This gap is especially prono…
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
Jiabing Yang, Yixiang Chen, Yuan Xu +14
Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demonstrating remarkable potential for…
LaPA: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation
Jiabing Yang, Yixiang Chen, Zichen Wen +8
Prefix-based methods have emerged as a promising paradigm for Controllable Text Generation (CTG) due to their parameter efficiency. However, while effective in short sequences, the…
EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation
Yuan Xu, Jiabing Yang, Xiaofeng Wang +16
Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-…
Learning to Explore: Policy-Guided Outlier Synthesis for Graph Out-of-Distribution Detection
Li Sun, Lanxu Yang, Jiayu Tian +6
Detecting out-of-distribution (OOD) graphs is crucial for ensuring the safety and reliability of Graph Neural Networks. In unsupervised graph-level OOD detection, models are typica…