9 papers
Are Your Generated Instances Truly Useful? GenBench-MILP: A Benchmark Suite for MILP Instance Generation
Yidong Luo, Chenguang Wang, Dong Li +1
The proliferation of machine learning-based methods for Mixed-Integer Linear Programming (MILP) instance generation has surged, driven by the need for diverse training datasets. Ho…
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
Xin Luo, Jiahao Wang, Chenyuan Wu +6
Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce…
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…
ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
Xinhang Chen, Chao Zhang, Jiahuan He +9
DeepSeek-V3.2-Exp introduces a sparse attention mechanism that significantly reduces inference latency in long-context scenarios. Although the overall throughput has improved great…
Modeling Hypergraph Using Large Language Models
Bingqiao Gu, Jiale Zeng, Xingqin Qi +1
Due to the advantages of hypergraphs in modeling high-order relationships in complex systems, they have been applied to higher-order clustering, hypergraph neural networks and comp…
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
MiniMax, :, Aili Chen +125
We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combin…