1 citations · 1 across the 2 of their papers we have counts for
5 papers
On the Off-Policy Teacher in On-Policy Distillation
Langlin Huang, Hao Liu, Mononito Goswami +6
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teache…
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu +9
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its…
Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling
Xinyu Li, Mononito Goswami, Hao Liu +5
Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how…
LLMs for Customized Marketing Content Generation and Evaluation at Scale
Haoran Liu, Amir Tahmasbi, Ehtesham Sam Haque +1
Offsite marketing is essential in e-commerce, enabling businesses to reach customers through external platforms and drive traffic to retail websites. However, most current offsite…
SERP Interference Network and Its Applications in Search Advertising
Purak Jain, Sandeep Appala
Search Engine marketing teams in the e-commerce industry manage global search engine traffic to their websites with the aim to optimize long-term profitability by delivering the be…