5 papers
TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models
Jinkun Zhao, Kui Zhang, Wenjun Wu
The paper introduces Tone‑Pressure Contrastive Decoding (TPCD), which subtracts logits from high‑pressure prompts from those of neutral prompts to reduce commitment bias in vision‑…
Learning to Adapt SFT Data for Better Reasoning Generalization
Lisong Sun, Li Wang, Chen Zhang +4
Large language models (LLMs) have achieved remarkable progress, with post-training playing a crucial role in enhancing their reasoning capabilities. Among post-training paradigms,…
TAMTRL: Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning in Long-Context Compression
Li Wang, Yandong Wang, Xin Yu +3
The rapid progress of large language models (LLMs) has led to remarkable performance gains across a wide range of tasks. However, when handling long documents that exceed the model…
Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
Li Wang, Changhao Zhang, Zengqi Xiu +4
Despite recent advances in the reasoning capabilities of Large Language Models (LLMs), improving the reasoning ability of Small Language Models (SLMs, e.g., up to 1.5B parameters)…
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
Jinkun Zhao, Yuanshuai Wang, Xingjian Zhang +6
With the rapid evolution of artificial intelligence, AIOps has emerged as a prominent paradigm in DevOps. Lots of work has been proposed to improve the performance of different AIO…