4 papers
Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
Sarthak Harne, Chinmay Karkar, Yash Pandya +2
Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI)…
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Yash Pandya, Sahil Gupta, Sarthak Harne +10
Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and s…
Dialogue to Discovery: Attribute-Aware Preference Elicitation for Conversational Product Search Assistants
Sarthak Harne, Natwar Modani, Debabrata Mahapatra +1
Conversational product search assistants offer a more expressive, natural, and interactive alternative to traditional keyword-based product search. With limited screen space, showi…
LLM driven Text-to-Table Generation through Sub-Tasks Guidance and Iterative Refinement
Rajmohan C, Sarthak Harne, Arvind Agarwal
Transforming unstructured text into structured data is a complex task, requiring semantic understanding, reasoning, and structural comprehension. While Large Language Models (LLMs)…