24 papers
Detecting AI-Generated Video: A Vision-Language Dual-View Survey
Dylan Xinming Hou, Juntian Zhang, Xu Gu +5
The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspect…
SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation
Mohammed Talha Alam, Nada Saadi, Fahad Shamshad +4
Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety per…
Entropy-Gated Latent Recursion
Soham Bhattacharjee, Dushyant Singh Chauhan, Salem Lahlou +2
Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-le…
A Gravitational Interpretation of Fine-Tuning Reversion
Samuele Poppi, Nils Lukas
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge,…
PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity
Muhammad Waseem, Nurbek Tastan, Andrej Jovanovic +4
Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foundation models. Heterogeneous hardware res…
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
Sai Kartheek Reddy Kasu, Nils Lukas, Samuele Poppi
Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialogue, yet its final-turn refu…