7 papers
CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models
Haiping Liu, Qian Zhao, Lijing Lin +2
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation…
TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents
Jingyu Sun, Yuyang Xue, Mingyang Li +9
Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the la…
PMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM Agents
Jingyu Sun, Yan Lin, Yuyang Xue +10
Long-term memory is essential for LVLM agents to maintain consistency and integrate information across extended multimodal interactions. Existing agent memory systems, however, oft…
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs
Jingyu Sun, Jiachen Tu, Yuyang Xue +8
Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself. Counterfactual i…
Computational Lesions in Multilingual Language Models Separate Shared and Language-specific Brain Alignment
Yang Cui, Jingyuan Sun, Yizheng Sun +8
How the brain supports language across different languages is a basic question in neuroscience and a useful test for multilingual artificial intelligence. Neuroimaging has identifi…
Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study
Yizheng Sun, Hao Li, Chang Xu +4
Vision-Language Models (VLMs) are powerful yet computationally intensive for widespread practical deployments. To address such challenge without costly re-training, post-training a…