Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
nDNA -- the Semantic Helix of Artificial Cognition
Amitava Das
As AI foundation models grow in capability, a deeper question emerges: What shapes their internal cognitive identity -- beyond fluency and output? Benchmarks measure behavior, but…
cs.AI2025
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
Amitava Das, Vinija Jain, Aman Chadha
Large Language Models (LLMs) fine-tuned to align with human values often exhibit alignment drift, producing unsafe or policy-violating completions when exposed to adversarial promp…