From the 2 of 17 linked papers with an AI index.
17 papers
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Yansen Zhang, Yilu Liu, Tianyu Liu +6
The paper proposes CostAda, a cost‑aware controller that guides large language model‑based discovery by evaluating frontier progress relative to the token cost incurred, enabling m…
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
Bowei He, Yankai Chen, Xiaokun Zhang +1
The paper proposes Branching Policy Optimization (BPO), a reinforcement learning method for large language model agents operating in deterministic, snapshottable sandboxes, which l…
LLM-as-a-Judge for Reliable and Explainable Offline Evaluation in Top-K Recommendation
Yue Que, Junyi Zhou, Xiaokun Zhang +3
Recommendation evaluation plays a crucial role in guiding the refinement and deployment of recommender systems. Most existing trials rely on offline evaluation using Top-K metrics…
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
Jiamin Chen, Yidi Wu, Qiexiang Wang +6
Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than construct…
DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation
Jiamin Chen, Qianben Chen, Jiawen Zhang +5
Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinematic control, audio, and cros…
Looking Farther with Confidence: Uncertainty-Guided Future Learning for Sequential Recommendation
Ziqiang Cui, Xing Tang, Peiyang Liu +4
Sequential recommendation effectively models dynamic user interests but continues to face challenges related to data sparsity. While self-supervised learning has alleviated this is…