4 papers
AgentVLN: Towards Agentic Vision-and-Language Navigation
Zihao Xin, Wentong Li, Yixuan Jiang +6
Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-La…
DECIDER: A Dual-System Rule-Controllable Decoding Framework for Language Generation
Chen Xu, Tian Lan, Yu Ji +8
Constrained decoding approaches aim to control the meaning or style of text generated by the pre-trained large language models (LLMs or also PLMs) for various tasks at inference ti…
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
Fan Yuan, Xiaoyuan Fang, Rong Quan +4
Visual Commonsense Reasoning, which is regarded as one challenging task to pursue advanced visual scene comprehension, has been used to diagnose the reasoning ability of AI systems…
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
Fan Yuan, Chi Qin, Xiaogang Xu +1
Large Vision-Language Models (LVLMs) have shown remarkable performance on many visual-language tasks. However, these models still suffer from multimodal hallucination, which means…