activity
20242026
collaborators

5 papers

cs.HC2026

UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization

Aeree Cho, Alexander D. Greenhalgh, Jonathan Bodea +2

Reinforcement learning has emerged as a dominant technique for fine-tuning the behavior of large language models, with policy optimization (PO) algorithms such as GRPO, DAPO, and D…

cs.CV2025

ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages

Matthew Hull, Haoyang Yang, Pratham Mehta +8

As 3D Gaussian Splatting (3DGS) gains rapid adoption in safety-critical tasks for efficient novel-view synthesis from static images, how might an adversary tamper images to cause h…

cs.SE2025

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

Seongmin Lee, Aeree Cho, Grace C. Kim +3

As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe out…

cs.CR2025

3D Gaussian Splat Vulnerabilities

Matthew Hull, Haoyang Yang, Pratham Mehta +8

With 3D Gaussian Splatting (3DGS) being increasingly used in safety-critical applications, how can an adversary manipulate the scene to cause harm? We introduce CLOAK, the first at…

cs.LG2024

Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation

Aeree Cho, Grace C. Kim, Alexander Karpekov +6

The Transformer architecture underpins modern large language models powering state-of-the-art text generation and AI applications. However, its complexity makes it difficult for no…