4 papers
KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing
Mingshu Cai, Miao Zhang, Chenghe Yang +3
In recent years, training-free video generation has progressed remarkably. However, when handling complex textual instructions, existing methods still suffer from semantic ambiguit…
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
Yixuan Li, Yuhui Chen, Mingcai Zhou +3
Spatial perception and reasoning are crucial for Vision-Language-Action (VLA) models to accomplish fine-grained manipulation tasks. However, existing approaches often lack the abil…
Q-Tacit: Image Quality Assessment via Latent Visual Reasoning
Yuxuan Jiang, Yixuan Li, Hanwei Zhu +3
Vision-Language Model (VLM)-based image quality assessment (IQA) has been significantly advanced by incorporating Chain-of-Thought (CoT) reasoning. Recent work has refined image qu…
Dual-Granularity Contrastive Reward via Generated Episodic Guidance for Efficient Embodied RL
Xin Liu, Yixuan Li, Yuhui Chen +3
Designing suitable rewards poses a significant challenge in reinforcement learning (RL), especially for embodied manipulation. Trajectory success rewards are suitable for human jud…