Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
Sarthak Harne, Chinmay Karkar, Yash Pandya +2
Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI)…
cs.AI2026
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Yash Pandya, Sahil Gupta, Sarthak Harne +10
Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and s…