6 papers
GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning
Xin Xiao, Jiang Zhong, Junnan Zhu +4
Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows ar…
From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation
Qin Lei, Jiang Zhong, Xin Xiao +2
Curvilinear object segmentation, including vessels and cracks, is challenging due to extreme spatial sparsity and topological fragility, where small local errors can cause severe s…
Do Models See in Line with Human Vision? Probing the Correspondence Between LVLM Representations and EEG Signals
Xin Xiao, Yang Lei, Haoyang Zeng +6
Large Vision Language Models (LVLMs) exhibit strong visual understanding and reasoning abilities. However, whether their internal representations reflect human visual cognition is…
Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
Haoyun Yang, Xin Xiao, Jiang Zhong +5
Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal represe…
Exploring Similarity between Neural and LLM Trajectories in Language Processing
Xin Xiao, Kaiwen Wei, Jiang Zhong +2
Understanding the similarity between large language models (LLMs) and human brain activity is crucial for advancing both AI and cognitive neuroscience. In this study, we provide a…
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
Kaiwen Wei, Xiao Liu, Jie Zhang +11
Multimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence, and numerous video-based…