#chain-of-thought reasoning
11 papers · 1 filter
Cybersecurity Detection Classification with Reasoning-enabled Language Models
Amol Khanna, Manu Nandan, Cristian Viorel Popa +10
The paper introduces a chain-of-thought reasoning classifier built on large language models to triage Windows endpoint security alerts, using a calibrated confidence estimator to i…
Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models
Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta +1
The paper introduces CoT-Mediate, a framework that edits a medical vision-language model's own generated reasoning to test whether the model's predictions follow that reasoning, re…
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Xinyu Tang, Qianggang Cao, Yurou Liu +13
The paper introduces a training pipeline that scales zero‑reinforcement‑learning to a trillion‑parameter language model, revealing emergent chain‑of‑thought reasoning abilities and…
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
Hefeng Zhou, Jinxuan Zhang, Jiong Lou +4
The paper introduces Deep Interaction, a method that lets users directly edit the chain‑of‑thought output of large language models to fix reasoning errors, resulting in higher corr…
Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models
Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5
The paper presents a large-scale structured reasoning dataset created via slice‑wise synthesis that encodes chain‑of‑thought explanations for 3D medical images, and uses it to inst…
Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration
Shuhao Li, Guodong Du, Anhao Zhao +3
The paper examines how supervised fine-tuning, reinforcement learning, and on‑policy distillation affect confidence estimates of large language models during chain‑of‑thought reaso…