9 papers
Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility
Siyang Wu, Yibo Jiang, Bryon Aragam
Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. W…
Contemporary AI lacks the imagination to diverge or negate in science
Honglin Bao, Siyang Wu, Xiao Liu +3
Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mo…
Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction
Zehan Li, Yutong Zhu, Siyang Wu +2
Large language models produce fluent fiction, yet their creative output is widely seen as flat. We ask where this quality originates in the training and whether it affects differen…
Zero-Forgetting CISS via Dual-Phase Cognitive Cascades
Yuquan Lu, Yifu Guo, Zishan Xu +6
Continual semantic segmentation (CSS) is a cornerstone task in computer vision that enables a large number of downstream applications, but faces the catastrophic forgetting challen…
Mapping Overlaps in Benchmarks through Perplexity in the Wild
Siyang Wu, Honglin Bao, Sida Li +2
We introduce benchmark signatures to characterize the capacity demands of LLM benchmarks and their overlaps. Signatures are sets of salient tokens from in-the-wild corpora whose mo…
EDIS: Diagnosing LLM Reasoning via Entropy Dynamics
Chenghua Zhu, Siyan Wu, Xiangkang Zeng +6
Entropy-based confidence signals are increasingly leveraged to improve reasoning in large language models (LLMs), yet existing approaches treat confidence as a static quantity -- t…