1 citations · 1 across the 6 of their papers we have counts for
7 papers
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
Zijun Min, Bingshuai Liu, Ante Wang +4
Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus…
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
Pakorn Ueareeworakul, Shuman Liu, Jinghao Feng +7
As global e-commerce rapidly expands into emerging markets, the lack of high-quality semantic representations for low-resource languages has become a decisive bottleneck for retrie…
Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications
Shuyi Xie, Ziqin Liew, Hailing Zhang +5
Large Language Models (LLMs) excel on general-purpose NLP benchmarks, yet their capabilities in specialized domains remain underexplored. In e-commerce, existing evaluations-such a…
Mid-Training of Large Language Models: A Survey
Kaixiang Mo, Yuxin Shi, Weiwei Weng +4
Large language models (LLMs) are typically developed through large-scale pre-training followed by task-specific fine-tuning. Recent advances highlight the importance of an intermed…
Compass-Thinker-7B Technical Report
Anxiang Zeng, Haibo Zhang, Kaixiang Mo +6
Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning i…
Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization
Meng Li, Guangda Huzhang, Haibo Zhang +2
Direct Preference Optimization (DPO) has emerged as a promising framework for aligning Large Language Models (LLMs) with human preferences by directly optimizing the log-likelihood…