7 papers
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
Guoqiang Zhang, Kexin Tan, Ming Zhang +12
Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a sing…
Agents in the Large: Perception-Centered Architecture for Persistent Agents
Shihan Dou, Haoxiang Jia, Shichun Liu +14
Cognitive language agents have achieved substantial progress by equipping language models with memory, tools, and decision-making procedures, enabling agents to reason and act in i…
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Junjie Ye, Zhuohui Sheng, Shaofan Liu +12
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps)…
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Shaowen Wang, Yuke Zheng, Tansheng Zhu +4
Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling a…
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
Jiazheng Zhang, Ziche Fu, Junrui Shen +17
Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (RLVR) for LLMs. However, existi…
CL-bench Life: Can Language Models Learn from Real-Life Context?
Shihan Dou, Yujiong Shen, Chenhao Huang +35
Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for models. As these systems move be…