activity
20242026
most citedCRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions

1 citations · 1 across the 6 of their papers we have counts for

collaborators

10 papers

cs.AI2026

Agentic Confidence Calibration

Jiaxin Zhang, Caiming Xiong, Chien-Sheng Wu

AI agents are rapidly advancing from passive language models to autonomous systems executing complex, multi-step tasks. Yet their overconfidence in failure remains a fundamental ba…

cs.AI2026

Agentic Uncertainty Quantification

Jiaxin Zhang, Prafulla Kumar Choubey, Kung-Hsiang Huang +2

Although AI agents have demonstrated impressive capabilities in long-horizon reasoning, their reliability is severely hampered by the ``Spiral of Hallucination,'' where early epist…

cs.CL2025

GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness

Kung-Hsiang Huang, Haoyi Qiu, Yutong Dai +2

Graphical user interface (GUI) agents built on vision-language models have emerged as a promising approach to automate human-computer workflows. However, they also face the ineffic…

cs.CL2025

Benchmarking Deep Search over Heterogeneous Enterprise Data

Prafulla Kumar Choubey, Xiangyu Peng, Shilpa Bhagavath +3

We present a new benchmark for evaluating Deep Search--a realistic and complex form of retrieval-augmented generation (RAG) that requires source-aware, multi-hop reasoning over div…

cs.CL20251 cited

CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions

Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat +6

While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platform…

cs.CL2025

BingoGuard: LLM Content Moderation Tools with Risk Levels

Fan Yin, Philippe Laban, Xiangyu Peng +7

Malicious content generated by large language models (LLMs) can pose varying degrees of harm. Although existing LLM-based moderators can detect harmful content, they struggle to as…