3 papers
cs.LG2026
Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding
Yingxuan Zhuang, Jingxiao Yang, Miao Pan +7
MLLMs frequently hallucinate objects inconsistent with visual inputs. This issue is typically attributed to the over-reliance on language priors, which can override the visual cont…
cs.CR2026
Toward a Principled Framework for Agent Safety Measurement
Shuyi Lin, Anshuman Suri, Alina Oprea +1
LLM agents emit actions, not just text, and once taken, those actions often cannot be undone. Yet today's agent-safety evaluations run greedy or a few sampled rollouts and report a…
cs.CR2026
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
Shuyi Lin, Anshuman Suri, Alina Oprea +1
As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their vulnerability to jailbreak attacks pres…