From the 1 of 20 linked papers with an AI index.
7 papers · 1 filter
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
Sen Xu, Wei Wang, Shixi Liu +5
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallo…
A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions
Jianghan Shen, Siqi Luo, Yue Li +9
Policy gradient algorithms for language models optimize the same objective , which has exactly two factors: the trajectory probability…
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
Sen Xu, Shixi Liu, Wei Wang +6
This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly smal…
MemVerse: Multimodal Memory for Lifelong Learning Agents
Junming Liu, Yifei Sun, Weihua Cheng +11
Despite rapid progress in large-scale language and vision models, AI agents still suffer from a fundamental limitation: they cannot remember. Without reliable memory, agents catast…
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
Jianghan Shen, Siqi Luo, Xinyu Cheng +6
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only superv…
MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
Weihua Cheng, Junming Liu, Yifei Sun +3
Multimodal Large Language Models (MLLMs) have significantly advanced GUI agents, yet long-horizon automation remains constrained by two critical bottlenecks: context overload from…