6 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
Control-R: Towards controllable test-time scaling
Di Zhang, Weida Wang, Junxian Li +10
This paper target in addressing the challenges of underthinking and overthinking in long chain-of-thought (CoT) reasoning for Large Reasoning Models (LRMs) by introducing Reasoning…
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
Di Zhang, Junxian Li, Jingdi Lei +10
Vision-language models (VLMs) have shown remarkable advancements in multimodal reasoning tasks. However, they still often generate inaccurate or irrelevant responses due to issues…
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
Xi Wang, Jie Liu, Jianbo Wu +4
Compute eXpress Link (CXL) is emerging as a promising memory interface technology. However, its performance characteristics remain largely unclear due to the limited availability o…
LoRA Diffusion: Zero-Shot LoRA Synthesis for Diffusion Model Personalization
Ethan Smith, Rami Seid, Alberto Hojel +2
Low-Rank Adaptation (LoRA) and other parameter-efficient fine-tuning (PEFT) methods provide low-memory, storage-efficient solutions for personalizing text-to-image models. However,…
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Di Zhang, Jianbo Wu, Jingdi Lei +9
This paper presents an advanced mathematical problem-solving framework, LLaMA-Berry, for enhancing the mathematical reasoning ability of Large Language Models (LLMs). The framework…