Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Iterative Finetuning is Mostly Idempotent
Zephaniah Roe, Jack Sanderson, Dang Nguyen +5
If a model has some behavioral tendency, such as sycophancy or misalignment, and it is trained on its own outputs, will the tendency be amplified in the next generation of models?…
cs.AI2025
Know Thyself? On the Incapability and Implications of AI Self-Recognition
Xiaoyan Bai, Aryan Shrivastava, Ari Holtzman +1
Self-recognition is a crucial metacognitive capability for AI systems, relevant not only for psychological analysis but also for safety, particularly in evaluative scenarios. Motiv…