Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Measuring AI Reasoning: A Guide for Researchers
Munachiso Samuel Nwadike, Zangir Iklassov, Kareem Ali +2
In this paper, we offer a guide for researchers on evaluating reasoning in language models, building the case that reasoning should be assessed through evidence of adaptive, multi-…
cs.AI2025
AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents
Renxi Wang, Rifo Ahmad Genadi, Bilal El Bouardi +5
Language model (LM) agents have gained significant attention for their ability to autonomously complete tasks through interactions with environments, tools, and APIs. LM agents are…