4 papers
MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models
Yuncheng Yang, Feiyang Ye, Shixian Luo +7
Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving effici…
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
Shixian Luo, Zezhou Zhu, Yu Yuan +3
Geometric spatial reasoning forms the foundation of many applications in artificial intelligence, yet the ability of large language models (LLMs) to operate over geometric spatial…
PEVLM: Parallel Encoding for Vision-Language Models
Letian Kang, Shixian Luo, Yiqiang Li +5
Vision-Language Models (VLMs) have demonstrated strong capabilities in multimodal understanding and generation tasks. However, their application to long video understanding remains…
Cognitive Memory in Large Language Models
Lianlei Shan, Shixian Luo, Zezhou Zhu +2
This paper examines memory mechanisms in Large Language Models (LLMs), emphasizing their importance for context-rich responses, reduced hallucinations, and improved efficiency. It…