4 papers
RouteSparse: Input-Conditional Pattern Routing for Budgeted Long-Context Prefilling
Chao Zhang, Yifan Ji, Ziyan Zhang +2
Dynamic sparse attention can reduce the quadratic cost of long-context prefilling without changing model weights. MInference assigns each attention head one pattern offline and est…
Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation
Xin Zhang, Yang Cao, Baoxing Wu +2
Large language models have shown strong performance in natural language generation and downstream reasoning tasks, but they still struggle with logical consistency, factual groundi…
SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation
Xin Zhang, Yang Cao, Baoxing Wu +2
Large Language Models (LLMs) have demonstrated strong capabilities across diverse NLP applications, such as translation, text generation, and question answering. Nevertheless, they…
MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
Chenglong Wang, Yang Gan, Hang Zhou +10
Recent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behin…