1 citations · 1 across the 4 of their papers we have counts for
6 papers
The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding
Yifan Qian, Zhe Wen, Alexander C. Furnas +3
Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapidly diffusing into scientific practice, h…
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
Post-training LLMs with RLHF and preference optimization methods (e.g., DPO, IPO) has greatly improved alignment, yet these approaches assume a single objective. In reality, humans…
Best Policy Learning from Trajectory Preference Feedback
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned reward models makes it vulnerable t…
MiMo-V2-Flash Technical Report
Core Team, Bangjun Xiao, Bingquan Xia +123
We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-…
Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation
Wenzhen Luo, Wei Guan, Yifan Yao +6
We introduce Falcon, a cross-domain Chinese text-to-SQL benchmark grounded in an enterprise-compatible dialect (MaxCompute/Hive). It contains 600 Chinese questions over 28 database…
Online Bandit Learning with Offline Preference Data for Improved RLHF
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
Reinforcement Learning with Human Feedback (RLHF) is at the core of fine-tuning methods for generative AI models for language and images. Such feedback is often sought as rank or p…