activity
20242026
collaborators

15 papers

cs.AI2026

Evaluating Multiple LLM Generations with Validated Task Coverage

Florian Le Bronnec, Rio Yokota

Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination. Predominant evaluation settings, however, still focus…

cs.CL2026

On the Optimal Reasoning Length for RL-Trained Language Models

Daisuke Nohara, Taishi Nakamura, Rio Yokota

Reinforcement learning substantially improves reasoning in large language models, but it also tends to lengthen chain-of-thought outputs and increase computational cost. Although l…

cs.CL2026

Beyond pass@k: Redundancy-Aware RLVR for Multi-Sample Code Generation

Le Bronnec Florian, Alexandre Verine, Rio Yokota +1

LLMs for code generation are commonly evaluated in repeated-sampling settings using Pass@k, where multiple candidate programs are executed against unit tests under a finite samplin…

cs.LG2026

Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks

Taishi Nakamura, Satoki Ishikawa, Masaki Kawamura +4

Empirical scaling laws have driven the evolution of large language models (LLMs), yet their coefficients shift whenever the model architecture or data pipeline changes. Mixture-of-…

cs.LG2026

Rewriting Pre-Training Data Boosts LLM Performance in Math and Code

Kazuki Fujii, Yukito Tajima, Sakae Mizuki +14

The performance of large language models (LLMs) in program synthesis and mathematical reasoning is fundamentally limited by the quality of their pre-training corpora. We introduce…

cs.NE2026

Evolutionary Context Search for Automated Skill Acquisition

Qi Sun, Stefan Nielsen, Rio Yokota +1

Large Language Models cannot reliably acquire new knowledge post-deployment -- even when relevant text resources exist, models fail to transform them into actionable knowledge with…