benchmark compilation 1capability probing 1evaluation methods 1large language models 1representation learning 1
From the 1 of 9 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
Anxiang Zeng, Haibo Zhang, Hailing Zhang +13
We present CompassMax-V3-Thinking, a hundred-billion-scale MoE reasoning model trained with a new RL framework built on one principle: each prompt must matter. Scaling RL to this s…
cs.AI2025
Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications
Shuyi Xie, Ziqin Liew, Hailing Zhang +5
Large Language Models (LLMs) excel on general-purpose NLP benchmarks, yet their capabilities in specialized domains remain underexplored. In e-commerce, existing evaluations-such a…
cs.AI2025
Compass-Thinker-7B Technical Report
Anxiang Zeng, Haibo Zhang, Kaixiang Mo +6
Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning i…