4 citations · 8 across the 11 of their papers we have counts for
4 papers · 1 filter
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
Nikita Gupta, Riju Chatterjee, Lukas Haas +9
We introduce DeepSearchQA, a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional benchmarks…
Me-Agent: A Personalized Mobile Agent with Two-Level User Habit Learning for Enhanced Interaction
Shuoxin Wang, Chang Liu, Gowen Loo +5
Large Language Model (LLM)-based mobile agents have made significant performance advancements. However, these agents often follow explicit user instructions while overlooking perso…
SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
Yushen Fang, Jianjun Li, Mingqian Ding +3
Although Large language Model (LLM)-powered information extraction (IE) systems have shown impressive capabilities, current fine-tuning paradigms face two major limitations: high t…
POLIS-Bench: Towards Multi-Dimensional Evaluation of LLMs for Bilingual Policy Tasks in Governmental Scenarios
Tingyue Yang, Junchi Yao, Yuhui Guo +1
We introduce POLIS-Bench, the first rigorous, systematic evaluation suite designed for LLMs operating in governmental bilingual policy scenarios. Compared to existing benchmarks, P…