10 papers
VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
Jiabin Lou, Haopeng Wang, Yuanshuai Wang +6
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encode spatial priors, such as orie…
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
Siwei Wen, Zhangcheng Wang, Xingjian Zhang +2
Online video understanding requires models to perform continuous perception and long-range reasoning within potentially infinite visual streams. Its fundamental challenge lies in t…
Parallel Layer Normalization for Universal Approximation
Yunhao Ni, Yuxin Guo, Yuhe Liu +4
This paper studies the approximation capabilities of neural networks that combine layer normalization (LN) with linear layers. We prove that networks consisting of two linear layer…
RLLaVA: An RL-central Framework for Language and Vision Assistants
Lei Zhao, Zihao Ma, Boyu Lin +3
We present an RL-central framework for Language and Vision Assistants (RLLaVA) with its formulation of Markov decision process (MDP). RLLaVA decouples RL algorithmic logic from mod…
VoQA: Visual-only Question Answering
Jianing An, Luyang Jiang, Jie Luo +2
Visual understanding requires interpreting both natural scenes and the textual information that appears within them, motivating tasks such as Visual Question Answering (VQA). Howev…
What Color Is It? A Text-Interference Multimodal Hallucination Benchmark
Jinkun Zhao, Lei Huang, Haixin Ge +1
With the rapid advancement of Large Models, numerous text-and-vision-fused Multimodal Large Models (MLMs) have emerged. However, these MLMs remain susceptible to informational inte…