5 papers · 1 filter
RouteSparse: Input-Conditional Pattern Routing for Budgeted Long-Context Prefilling
Chao Zhang, Yifan Ji, Ziyan Zhang +2
Dynamic sparse attention can reduce the quadratic cost of long-context prefilling without changing model weights. MInference assigns each attention head one pattern offline and est…
Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation
Xin Zhang, Yang Cao, Baoxing Wu +2
Large language models have shown strong performance in natural language generation and downstream reasoning tasks, but they still struggle with logical consistency, factual groundi…
SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation
Xin Zhang, Yang Cao, Baoxing Wu +2
Large Language Models (LLMs) have demonstrated strong capabilities across diverse NLP applications, such as translation, text generation, and question answering. Nevertheless, they…
MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
Chenglong Wang, Yang Gan, Hang Zhou +10
Recent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behin…
Revealing the Parallel Multilingual Learning within Large Language Models
Yongyu Mu, Peinan Feng, Zhiquan Cao +8
In this study, we reveal an in-context learning (ICL) capability of multilingual large language models (LLMs): by translating the input to several languages, we provide Parallel In…