8 papers
FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows
Darshan Deshpande, Yoshinari Fujinuma, Martyna Markiewicz +5
Vision Language Models have recently shown improvements in several objective and verifiable domains such as object detection but continue to underperform on subjective and creative…
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Darshan Deshpande
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties b…
DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
Li Siyan, Darshan Deshpande, Anand Kannappan +1
When recalling information in conversation, people often arrive at the recollection after multiple turns. However, existing benchmarks for evaluating agent capabilities in such tip…
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
Darshan Deshpande, Anand Kannappan, Rebecca Qian
Recent advances in reinforcement learning for code generation have made robust environments essential to prevent reward hacking. As LLMs increasingly serve as evaluators in code-ba…
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
Darshan Deshpande, Varun Gangal, Hersh Mehta +3
Recent works on context and memory benchmarking have primarily focused on conversational instances but the need for evaluating memory in dynamic enterprise environments is crucial…
TRAIL: Trace Reasoning and Agentic Issue Localization
Darshan Deshpande, Varun Gangal, Hersh Mehta +3
The increasing adoption of agentic workflows across diverse domains brings a critical need to scalably and systematically evaluate the complex traces these systems generate. Curren…