14 papers
DRIFT: Direct-Recursive Intervention-Conditioned Forecasting of ICU Physiological Trajectories
Weixin Liu, Juming Xiong, Congning Ni +4
Many time-series forecasts depend not only on prior observations but also on actions specified during the forecast period. In intensive care units (ICUs), future vital signs and la…
Coverage-Controlled Preference Mining from Noisy Claim Verification for Evidence-Grounded Generation
Weixin Liu, Congning Ni, Qingyuan Song +4
Evidence-grounded generation produces summaries whose claims should be supported by supplied evidence, but claim-level verifiers provide noisy feedback and can reward models that s…
Learning When to Sample: Confidence-Aware Selective Sampling for Efficient Chain-of-Thought Reasoning
Juming Xiong, Kevin Guo, Congning Ni +7
Large language models (LLMs) can achieve strong reasoning performance through chain-of-thought (CoT) reasoning, yet they often generate unnecessarily long reasoning paths that incu…
CoRA: Confidence-Rationale Alignment for Reliable Chain-of-Thought Reasoning
Juming Xiong, Weixin Liu, Kevin Guo +9
Chain-of-thought (CoT) reasoning can improve LLM performance, but high answer confidence may be misleading when the accompanying CoT rationale is plausible yet incomplete or poorly…
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
Kevin H. Guo, Chao Yan, Avinash Baidya +5
Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behavior to sycophancy learned duri…
Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization
Weixin Liu, Bowen Qu, Juming Xiong +3
Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, or analytic workflows. Even w…