8 papers
SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models
Dongxu Zhang, Yiding Sun, Zihao Guo +5
Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may…
Dual Latent Memory for Visual Multi-agent System
Xinlei Yu, Chengming Xu, Zhangquan Chen +8
While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall":…
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
Dongxu Zhang, Yiding Sun, Cheng Tan +4
While Chain-of-Thought (CoT) reasoning significantly enhances the performance of Multimodal Large Language Models (MLLMs), its autoregressive nature incurs prohibitive latency cons…
SCP: Accelerating Discovery with a Global Web of Autonomous Scientific Agents
Yankai Jiang, Wenjie Lou, Lilong Wang +17
We introduce SCP: the Science Context Protocol, an open-source standard designed to accelerate discovery by enabling a global network of autonomous scientific agents. SCP is built…
Socrates-Mol: Self-Oriented Cognitive Reasoning through Autonomous Trial-and-Error with Empirical-Bayesian Screening for Molecules
Xiangru Wang, Zekun Jiang, Heng Yang +4
Molecular property prediction is fundamental to chemical engineering applications such as solvent screening. We present Socrates-Mol, a framework that transforms language models in…
NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction
Zhongmin Li, Runze Ma, Jiahao Tan +2
Nucleotide sequence variation can induce significant shifts in functional fitness. Recent nucleotide foundation models promise to predict such fitness effects directly from sequenc…