activity
20242026
most citedEnhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders

4 citations · 9 across the 11 of their papers we have counts for

collaborators

21 papers

cs.AI2026

Same Answer, Different Representations: Hidden instability in VLMs

Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena +6

The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions reflect stable multimodal processi…

cs.LG2025

Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value

Joe Edelman, Tan Zhi-Xuan, Ryan Lowe +30

Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to…

cs.CL2025

HACK: Hallucinations Along Certainty and Knowledge Axes

Adi Simhi, Jonathan Herzig, Itay Itzhak +7

Hallucinations in LLMs present a critical barrier to their reliable usage. Existing research usually categorizes hallucination by their external properties rather than by the LLMs'…

cs.AI2025

VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models

Aman Gupta, Denny O'Shea, Fazl Barez

Large language models (LLMs) are increasingly being used for tasks where outputs shape human decisions, so it is critical to verify that their responses consistently reflect desire…

cs.CY2025

Embodied AI: Emerging Risks and Opportunities for Policy Action

Jared Perlo, Alexander Robey, Fazl Barez +2

The field of embodied AI (EAI) is rapidly advancing. Unlike virtual AI, EAI systems can exist in, learn from, reason about, and act in the physical world. With recent advances in A…

cs.LG2025

Rethinking Safety in LLM Fine-tuning: An Optimization Perspective

Minseon Kim, Jin Myung Kwak, Lama Alssum +5

Fine-tuning language models is commonly believed to inevitably harm their safety, i.e., refusing to respond to harmful user requests, even when using harmless datasets, thus requir…