1 citations · 1 across the 6 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
ContextWeave: A Real-World Workflow Benchmark
Bo Wang, Yuqian Yao, Enxi Wang +25
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We…
cs.AI2026
Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling
Rongman Xu, Yifei Li, Tianzhe Zhao +3
Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-time scaling faces challenges in trade-off b…
cs.AI2026
Steering LLMs via Scalable Interactive Oversight
Enyu Zhou, Zhiheng Xi, Long Ma +9
As Large Language Models increasingly automate complex, long-horizon tasks such as \emph{vibe coding}, a supervision gap has emerged. While models excel at execution, users often s…