activity
20242026
collaborators

10 papers

cs.CL2026

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion

Haoyi Qiu, Yilun Zhou, Pranav Narayanan Venkit +4

As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored the dynamics of Agent-to-Agent (A2…

cs.AI2026

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey +4

Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limited by the lack of scalable,…

cs.CL2025

Direct Judgement Preference Optimization

Peifeng Wang, Austin Xu, Yilun Zhou +2

Auto-evaluation is crucial for assessing response quality and offering feedback for model development. Recent studies have explored training large language models (LLMs) as generat…

cs.CL2025

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou +3

Generative search engines and deep research LLM agents promise trustworthy, source-grounded synthesis, yet users regularly encounter overconfidence, weak sourcing, and confusing ci…

cs.CL2025

J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization

Austin Xu, Yilun Zhou, Xuan-Phi Nguyen +2

To keep pace with the increasing pace of large language models (LLM) development, model output evaluation has transitioned away from time-consuming human evaluation to automatic ev…

cs.CL2025

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Yilun Zhou, Austin Xu, Peifeng Wang +2

Scaling test-time computation, or affording a generator large language model (LLM) extra compute during inference, typically employs the help of external non-generative evaluators…