2 papers
cs.LG2026
Zenith: Scaling up Ranking Models for Billion-scale Livestreaming Recommendation
Ruifeng Zhang, Zexi Huang, Zikai Wang +11
Accurately capturing feature interactions is essential in recommender systems, and recent trends show that scaling up model capacity could be a key driver for next-level predictive…
cs.SE2026
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…