6 papers
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Lena Libon, Ben Rank, Jehyeok Yeon +5
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such…
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
Jehyeok Yeon, Ben Rank, Maksym Andriushchenko
AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces. Even no…
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
Jehyeok Yeon, Federico Cinus, Yifan Wu +1
Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity objective treats latent features as indep…
Quantitative Certification of Agentic Tool Selection
Jehyeok Yeon, Isha Chaudhary, Gagandeep Singh
Large language models (LLMs) are increasingly deployed in agentic systems, where a fundamental task is mapping user intents to relevant external tools. Errors in tool selection can…
Securing Multimodal AI through Internal Information Decomposition
Jehyeok Yeon, Hyeonjeong Ha, Qiusi Zhan +1
Multimodal large language models introduce attack surfaces absent in unimodal systems: adversaries can distribute malicious intent across modalities to evade unimodal safeguards. T…
TRAP: Targeted Redirecting of Agentic Preferences
Hangoo Kang, Jehyeok Yeon, Gagandeep Singh
Autonomous agentic AI systems powered by vision-language models (VLMs) are rapidly advancing toward real-world deployment, yet their cross-modal reasoning capabilities introduce ne…