activity
20162025
most citedRECAST: Enabling User Recourse and Interpretability of Toxicity Detection Models with Interactive Visualization

31 citations · 142 across the 40 of their papers we have counts for

collaborators

63 papers

cs.HC2025

LitForager: Exploring Multimodal Literature Foraging Strategies in Immersive Sensemaking

Haoyang Yang, Elliott H. Faa, Weijian Liu +3

Exploring and comprehending relevant academic literature is a vital yet challenging task for researchers, especially given the rapid expansion in research publications. This task f…

cs.LG2025

Diffusion Explorer: Interactive Exploration of Diffusion Models

Alec Helbling, Duen Horng Chau

Diffusion models have been central to the development of recent image, video, and even text generation systems. They posses striking geometric properties that can be faithfully por…

cs.SE2025

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

Seongmin Lee, Aeree Cho, Grace C. Kim +3

As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe out…

cs.LG2025

Shape it Up! Restoring LLM Safety during Finetuning

ShengYun Peng, Pin-Yu Chen, Jianfeng Chi +2

Finetuning large language models (LLMs) enables user-specific customization but introduces critical safety risks: even a few harmful examples can compromise safety alignment. A com…

cs.HC2025

HybridCollab: Unifying In-Person and Remote Collaboration for Cardiovascular Surgical Planning in Mobile Augmented Reality

Pratham Darrpan Mehta, Rahul Ozhur Narayanan, Vidhi Kulkarni +3

Surgical planning for congenital heart disease traditionally relies on collaborative group examinations of a patient's 3D-printed heart model, a process that lacks flexibility and…

cs.CV2025

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Alec Helbling, Tuna Han Salih Meral, Ben Hoover +2

Do the rich representations of multi-modal diffusion transformers (DiTs) exhibit unique properties that enhance their interpretability? We introduce ConceptAttention, a novel metho…