Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Reinforcement Learning with Decomposed Subtasks
Mattie Terzolo, Mikolaj Sacha, Ayan Sinha +1
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajecto…
cs.AI2025
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
Darvin Yi, Teng Liu, Mattie Terzolo +4
As large language model (LLM) agents increasingly undertake digital work, reliable frameworks are needed to evaluate their real-world competence, adaptability, and capacity for hum…