Publications (5)
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
Junjie Ye, Zhengyin Du, Xuesong Yao +11
Effective evaluation of multi-hop tool use is critical for analyzing the understanding, reasoning, and function-calling capabilities of large language models (LLMs). However, progr…
Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
ByteDance Seed, :, Jiaze Chen +267
We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 8…
RubyStar: A Non-Task-Oriented Mixture Model Dialog System
Huiting Liu, Tao Lin, Hanfei Sun +4
RubyStar is a dialog system designed to create "human-like" conversation by combining different response generation strategies. RubyStar conducts a non-task-oriented conversation o…
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
Zhan Ling, Kang Liu, Kai Yan +6
Large language models (LLMs) have demonstrated remarkable progress in understanding long-context inputs. However, benchmarks for evaluating the long-context reasoning abilities of…
Bayesian Optimization Algorithms for Accelerator Physics
Ryan Roussel, Auralee L. Edelen, Tobias Boltz +23
Accelerator physics relies on numerical algorithms to solve optimization problems in online accelerator control and tasks such as experimental design and model calibration in simul…