Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
Amr Moustafa, Max Feser, Florian Mai
Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain…
cs.AI2025
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
Tianyi Zhang, Florian Mai, Lucie Flek
Continual pretraining promises to adapt large language models (LLMs) to new domains using only unlabeled test-time data, but naively applying standard self-supervised objectives to…
cs.AI2025
Superalignment with Dynamic Human Values
Florian Mai, David Kaczér, Nicholas Kluge Corrêa +1
Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they…