collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

Yongchang Peng, Qingshui Gu, Liya Zhu +31

Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world…

cs.AI2026

Repo2Skill-Evo: Repository Skills Go Stale in Silence

Chenyuan Duan, Ge Shi, Zineng Mao +10

Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural knowledge: which APIs to call, w…

cs.AI2026

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Liya Zhu, Xin Ma, Tao Liu +35

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on re…

cs.AI2026

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

Zining Huang, Haoran Que, Hong Zeng +8

When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in t…

cs.AI2026

Knowledge Index of Noah's Ark

Sheng Jin, Minghao Liu, Yunze Xiao +24

Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consen…