#chain-of-thought reasoning

topicchain-of-thought reasoning

11 papers · 1 filter

cs.LG2026

Cybersecurity Detection Classification with Reasoning-enabled Language Models

Amol Khanna, Manu Nandan, Cristian Viorel Popa +10

The paper introduces a chain-of-thought reasoning classifier built on large language models to triage Windows endpoint security alerts, using a calibrated confidence estimator to i…

cs.LG2026

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta +1

The paper introduces CoT-Mediate, a framework that edits a medical vision-language model's own generated reasoning to test whether the model's predictions follow that reasoning, re…

cs.CL2026

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Xinyu Tang, Qianggang Cao, Yurou Liu +13

The paper introduces a training pipeline that scales zero‑reinforcement‑learning to a trillion‑parameter language model, revealing emergent chain‑of‑thought reasoning abilities and…

cs.AI2026

Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models

Hefeng Zhou, Jinxuan Zhang, Jiong Lou +4

The paper introduces Deep Interaction, a method that lets users directly edit the chain‑of‑thought output of large language models to fix reasoning errors, resulting in higher corr…

cs.CV2026

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5

The paper presents a large-scale structured reasoning dataset created via slice‑wise synthesis that encodes chain‑of‑thought explanations for 3D medical images, and uses it to inst…

cs.CL2026

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration

Shuhao Li, Guodong Du, Anhao Zhao +3

The paper examines how supervised fine-tuning, reinforcement learning, and on‑policy distillation affect confidence estimates of large language models during chain‑of‑thought reaso…