4 papers · 1 filter
Position-Aware Depth Decay Decoding (): Boosting Large Language Model Inference Efficiency
Siqi Fan, Xuezhi Fang, Xingrun Xing +3
Due to the large number of parameters, the inference phase of Large Language Models (LLMs) is resource-intensive. Unlike traditional model compression, which needs retraining, rece…
The Price of a Second Thought: On the Evaluation of Reasoning Efficiency in Large Language Models
Siqi Fan, Bowen Qin, Peng Han +3
Recent thinking models trained with reinforcement learning and backward-checking CoT often suffer from overthinking: they produce excessively long outputs even on simple problems,…
52B to 1T: Lessons Learned via Tele-FLM Series
Xiang Li, Yiqun Yao, Xin Jiang +17
Large Language Models (LLMs) represent a significant stride toward Artificial General Intelligence. As scaling laws underscore the potential of increasing model sizes, the academic…
Tele-FLM Technical Report
Xiang Li, Yiqun Yao, Xin Jiang +17
Large language models (LLMs) have showcased profound capabilities in language understanding and generation, facilitating a wide array of applications. However, there is a notable p…