most citedMiMo-V2-Flash Technical Report

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.DL2026

The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding

Yifan Qian, Zhe Wen, Alexander C. Furnas +3

Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapidly diffusing into scientific practice, h…

cs.LG2026

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1

Post-training LLMs with RLHF and preference optimization methods (e.g., DPO, IPO) has greatly improved alignment, yet these approaches assume a single objective. In reality, humans…

cs.LG2026

Best Policy Learning from Trajectory Preference Feedback

Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1

Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned reward models makes it vulnerable t…

cs.CL20261 cited

MiMo-V2-Flash Technical Report

Core Team, Bangjun Xiao, Bingquan Xia +123

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-…

cs.CL2025

Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation

Wenzhen Luo, Wei Guan, Yifan Yao +6

We introduce Falcon, a cross-domain Chinese text-to-SQL benchmark grounded in an enterprise-compatible dialect (MaxCompute/Hive). It contains 600 Chinese questions over 28 database…

cs.LG2025

Online Bandit Learning with Offline Preference Data for Improved RLHF

Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1

Reinforcement Learning with Human Feedback (RLHF) is at the core of fine-tuning methods for generative AI models for language and images. Such feedback is often sought as rank or p…