12 papers
TransMem: Transforming Hidden States into Memory for Large Language Models
Haodong Lei, Junming Liu, Yirong Chen +4
Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distrib…
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
Yang Chen, Yunwen Li, Yufan Shen +6
Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks:…
SemFlowRAG: Directed Semantic Flow from Abstraction to Evidence for Complex Reasoning
Houyuan Qin, Rong Wu, Qinyuan Qin +4
Retrieval-Augmented Generation (RAG) enhanced by Knowledge Graphs has shown promise in complex multi-hop reasoning tasks. However, existing graph-based retrieval methods typically…
MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
Weihua Cheng, Junming Liu, Yifei Sun +3
Multimodal Large Language Models (MLLMs) have significantly advanced GUI agents, yet long-horizon automation remains constrained by two critical bottlenecks: context overload from…
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
Yizhou Wang, Song Mao, Yang Chen +8
Recent multimodal large language models (MLLMs) increasingly integrate multiple vision encoders to improve performance on various benchmarks, assuming that diverse pretraining obje…
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
Junming Liu, Siyuan Meng, Yanting Gao +7
Multimodal reasoning in Large Language Models (LLMs) struggles with incomplete knowledge and hallucination artifacts, challenges that textual Knowledge Graphs (KGs) only partially…