3 citations · 3 across the 12 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations
Yunfan Zhou, Qiming Shi, Yizhou Yang +2
While recent Large Language Model (LLM)-based text-to-SQL systems achieve impressive performance on standard benchmarks, they struggle when user queries implicitly rely on domain-s…
cs.AI2026
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Qiming Shi, Yulong Tao, Linbo Jin +10
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments…
cs.AI2026
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Qiming Shi, Yibo Dou, Jiawen Zhu +5
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summ…