1 citations · 1 across the 6 of their papers we have counts for
6 papers · 1 filter
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
Zheyue Tan, Zhiyuan Li, Tao Yuan +13
Mixture-of-Experts (MoE) architectures have emerged as a promising approach to scale Large Language Models (LLMs). MoE boosts the efficiency by activating a subset of experts per t…
FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
Haotian Wu, Shufan Jiang, Chios Chen +5
As large language models (LLMs) advance in role-playing (RP) tasks, existing benchmarks quickly become obsolete due to their narrow scope, outdated interaction paradigms, and limit…
Implicit Reasoning in Large Language Models: A Comprehensive Survey
Jindong Li, Yali Fu, Li Fan +6
Large Language Models (LLMs) have demonstrated strong generalization across a wide range of tasks. Reasoning with LLMs is central to solving multi-step problems and complex decisio…
Thinking with Nothinking Calibration: A New In-Context Learning Paradigm in Reasoning Large Language Models
Haotian Wu, Bo Xu, Yao Shu +2
Reasoning large language models (RLLMs) have recently demonstrated remarkable capabilities through structured and multi-step reasoning. While prior research has primarily focused o…
dots.llm1 Technical Report
Bi Huo, Bin Tu, Cheng Qin +24
Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this…
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
Jialin Chen, Aosong Feng, Ziyu Zhao +7
Understanding the relationship between textual news and time-series evolution is a critical yet under-explored challenge in applied data science. While multimodal learning has gain…