1 citations · 1 across the 4 of their papers we have counts for
4 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
HyperD: Hybrid Periodicity Decoupling Framework for Traffic Forecasting
Minlan Shao, Zijian Zhang, Yili Wang +3
Accurate traffic forecasting plays a vital role in intelligent transportation systems, enabling applications such as congestion control, route planning, and urban mobility optimiza…
Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems
Xu Shen, Yixin Liu, Yiwei Dai +5
The communication topology in large language model-based multi-agent systems fundamentally governs inter-agent collaboration patterns, critically shaping both the efficiency and ef…
Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning
Ruxue Shi, Hengrui Gu, Hangting Ye +3
Few-shot tabular learning, in which machine learning models are trained with a limited amount of labeled data, provides a cost-effective approach to addressing real-world challenge…