co-evolution training 1computer-use agents 1reinforcement learning 1stateful applications 1synthetic environments 1
From the 1 of 5 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
Sarthak Harne, Chinmay Karkar, Yash Pandya +2
Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI)…
cs.AI2026
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Yash Pandya, Sahil Gupta, Sarthak Harne +10
Echoverse introduces a pipeline that compiles specifications into deep, stateful synthetic applications for training computer-use agents, using a co‑evolution loop that repairs env…