activity
20242026
collaborators

16 papers

cs.AI2026

WFM: Wiki Foundation Model for Complex Agentic Reasoning

Junnan Dong, Linhao Luo, Senlei Zhang +9

Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs have show…

cs.AI2026

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

Jiayi Kuang, Yinghui Li, Yunze Song +11

Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate onl…

cs.LG2026

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents

Qiang Liu, Taian Guo, Ruizhi Qiao +1

Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. However, in long-horizon, multi-turn tasks char…

cs.IR2026

Skills Know Their Neighbors: Cluster-Contrastive Capability Pages for Skill Retrieval

Zifei Wang, Wei Wen, Qiang Ji +1

As skill libraries grow, large language model agents must retrieve reusable skills from candidates that often share the same topic and vocabulary but implement different capabiliti…

cs.LG2026

Training-Free Hashing-Based Attention via Binary Principal Components

Daohai Yu, Zhanpeng Zeng, Keyu Chen +6

Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decodi…

cs.IR2026

Breaking the Evaluation Paradox: Evaluating High-Entropy Search with Computationally Irreducible Constraints

Juntao Wu, Wei Wen, Xianting Huang +4

Evaluating the exhaustive search capabilities of large language models (LLMs) is plagued by a fundamental paradox: verifying completeness requires complete ground truth, yet high-e…