1 citations · 1 across the 4 of their papers we have counts for
4 papers
OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
Divij Handa, David Blincoe, Orson Adams +1
Deploying capable and user-aligned LLM-based systems necessitates reliable evaluation. While LLMs excel in verifiable tasks like coding and mathematics, where gold-standard solutio…
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
Zehua Zhang, Ati Priya Bajaj, Divij Handa +13
Automatically compiling open-source software (OSS) projects is a vital, labor-intensive, and complex task, which makes it a good challenge for LLM Agents. Existing methods rely on…
ThinkTuning: Instilling Cognitive Reflections without Distillation
Aswin RRV, Jacob Dineen, Divij Handa +4
Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning. While RL drives this self-improveme…
Can NLP Models Correctly Reason Over Contexts that Break the Common Assumptions?
Neeraj Varshney, Mihir Parmar, Nisarg Patel +4
Pre-training on large corpora of text enables the language models to acquire a vast amount of factual and commonsense knowledge which allows them to achieve remarkable performance…