collaborators

7 papers

cs.AI2025

User-Feedback-Driven Adaptation for Vision-and-Language Navigation

Yongqiang Yu, Xuhui Li, Hazza Mahmood +6

Real-world deployment of Vision-and-Language Navigation (VLN) agents is constrained by the scarcity of reliable supervision after offline training. While recent adaptation methods…

cs.CV2025

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

Changlin Li, Jiawei Zhang, Shuhao Liu +4

Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training th…

cs.CV2025

Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models

Changlin Li, Jiawei Zhang, Zeyi Shi +3

Large-scale vision generative models, including diffusion and flow models, have demonstrated remarkable performance in visual generation tasks. However, transferring these pre-trai…

cs.CV2025

Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection

Mingfei Han, Haihong Hao, Jinxing Zhou +5

Vision-language models often hallucinate details, generating non-existent objects or inaccurate attributes that compromise output reliability. Existing methods typically address th…

cs.CV2025

Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models

Longtao Jiang, Jie Huang, Mingfei Han +5

Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant,…

cs.CV2025

Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation

Jinxing Zhou, Zhihui Li, Yongqiang Yu +7

We present \textbf{Met}a-\textbf{T}oken \textbf{Le}arning (Mettle), a simple and memory-efficient method for adapting large-scale pretrained transformer models to downstream audio-…