1 citations · 1 across the 7 of their papers we have counts for
9 papers
Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
Jinwu Hu, Dongjin Yang, Langyu Bian +6
Large language models (LLMs) have demonstrated impressive performance across various language tasks. However, existing LLM reasoning strategies mainly rely on the LLM itself with f…
Beyond Model Scaling: Test-Time Intervention for Efficient Deep Reasoning
Qianyue Wang, Jinwu Hu, Yufeng Wang +5
Large Reasoning Models (LRMs) excel at multi-step reasoning but often suffer from inefficient reasoning processes like overthinking and overshoot, where excessive or misdirected re…
Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance
Yufeng Wang, Jinwu Hu, Ziteng Huang +10
Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal ass…
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
Jinwu Hu, Wei Zhang, Yufeng Wang +4
Large Language Models (LLMs) have shown outstanding performance across a variety of tasks, partly due to advanced prompting techniques. However, these techniques often require leng…
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
Chen Zhang, Kuntai Du, Shu Liu +10
Large language models (LLMs) are widely used but expensive to run, especially as inference workloads grow. To lower costs, maximizing the request batch size by managing GPU memory…
Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement
Qianyue Wang, Jinwu Hu, Zhengping Li +4
Long-form story generation task aims to produce coherent and sufficiently lengthy text, essential for applications such as novel writingand interactive storytelling. However, exist…