4 papers
ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm
Kefan Song, Yanjun Qi
Autonomous CLI agents can now execute hundreds of actions across multi-hour sessions: writing code, executing shell commands, browsing the web, and managing cloud infrastructure, a…
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
Kefan Song, Amir Moeini, Peng Wang +4
Reinforcement learning (RL) is a framework for solving sequential decision-making problems. In this work, we demonstrate that, surprisingly, RL emerges during the inference time of…
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
Aman Goel, Daniel Schwartz, Yanjun Qi
Large language models (LLMs) have demonstrated impressive capabilities across diverse tasks, but they remain susceptible to hallucinations--generating content that appears plausibl…
A Comprehensive Survey on Concept Erasure in Text-to-Image Diffusion Models
Changhoon Kim, Yanjun Qi
Text-to-Image (T2I) models have made remarkable progress in generating high-quality, diverse visual content from natural language prompts. However, their ability to reproduce copyr…