works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CL2026

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Yanshi Li, Xueru Bai, Shuman Liu +1

The paper introduces RepBench, a framework that aggregates thousands of benchmark datasets into a large set of probe texts to evaluate capability-aligned representations of large l…

cs.CV2026

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

Yu Chen, Caorui Li, Ziyu Xiong +8

Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity in…

cs.LG2026

AMO: Adaptive Muon Orthogonalization

Xinlin Zhuang, Panyi Ouyang, Yichen Li +7

Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iterations as its core operation. Existi…

cs.CL2026

Decomposing and Steering Functional Metacognition in Large Language Models

Yanshi Li, Xueru Bai, Shuman Liu +2

Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies in benchmark settings. Prior…

cs.LG2026

ESPO: Entropy Importance Sampling Policy Optimization

Yuepeng Sheng, Yuwei Huang, Shuman Liu +2

Reinforcement learning (RL) has become a central component of post-training for large language models (LLMs), particularly for complex reasoning tasks that require stable optimizat…

cs.CL2025

Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings

Pakorn Ueareeworakul, Shuman Liu, Jinghao Feng +7

As global e-commerce rapidly expands into emerging markets, the lack of high-quality semantic representations for low-resource languages has become a decisive bottleneck for retrie…