activity
20242026
collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Sangwoo Cho, Kushal Chawla, Pengshan Cai +4

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, an…

cs.AI2026

SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation

Chenyang Zhu, Jiayu Yao, Kushal Chawla +10

As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows.…

cs.AI2026

Know Thy Reasoner: Not All Language Models Explore Alike

Moulik Choraria, Argyrios Gerogiannis, Anirban Das +4

Compute scaling for LLM reasoning trades off exploring solution approaches (\emph{breadth}) against refining promising ones (\emph{depth}), yet why a given trade-off works, and why…

cs.AI2026

A History-Aware Visually Grounded Critic for Computer Use Agents

Jaewoo Lee, Zaid Khan, Archiki Prasad +7

Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in compl…

cs.AI2025

RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization

Hanyang Zhao, Genta Indra Winata, Anirban Das +4

Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully a…