activity
20242026
collaborators

14 papers

cs.LG2026

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short

Han Zhou, Adam X. Yang, Laurence Aitchison +2

Reinforcement learning with verifiable rewards (RLVR) has become a leading paradigm for improving the reasoning ability of large language models through outcome-based supervision.…

cs.SE2026

SEW: Self-Evolving Agentic Workflows for Automated Code Generation

Siwei Liu, Jinyuan Fang, Han Zhou +2

Large Language Models (LLMs) have demonstrated effectiveness in code generation tasks. To enable LLMs to address more complex coding challenges, existing research has focused on cr…

cs.AI2026

Voxtral TTS

Mistral-AI, :, Alexander H. Liu +186

We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid…

cs.AI2026

Voxtral Realtime

Mistral-AI, :, Alexander H. Liu +166

We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adap…

cs.LG2026

Visual Planning: Let's Think Only with Images

Yi Xu, Chengzu Li, Han Zhou +4

Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse tasks. However, these model…

cs.LG2026

Agentic Policy Optimization via Instruction-Policy Co-Evolution

Han Zhou, Xingchen Wan, Ivan Vulić +1

Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capability of large language models (LLMs), enabling autonomous agents that can conduct effective m…