3 papers
cs.AI2026
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
Ye Mo, Kai Ye, Xianwei Mao +9
Document understanding aims to perform question answering and information extraction over document images, where the visual content is highly information-dense and most queries rel…
cs.CV2026
Semantic Audio-Visual Navigation in Continuous Environments
Yichen Zeng, Hebaixu Wang, Meng Liu +4
Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on pre…
cs.CL2024
Meta-Reflection: A Feedback-Free Reflection Learning Framework
Yaoke Wang, Yun Zhu, Xintong Bao +7
Despite the remarkable capabilities of large language models (LLMs) in natural language understanding and reasoning, they often display undesirable behaviors, such as generating ha…