collaborators

12 papers

cs.LG2026

Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems

Shichang Zhang, Hongzhe Du, Jiaqi W. Ma +1

Modern AI systems are typically developed through multiple stages-pretraining, fine-tuning rounds, and subsequent adaptation or alignment, where each stage builds on the previous o…

cs.CL2026

Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability

Usha Bhalla, Alex Oesterling, Claudio Mayrink Verdun +2

Translating the internal representations and computations of models into concepts that humans can understand is a key goal of interpretability. While recent dictionary learning met…

cs.LG2026

Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders

Aaron J. Li, Suraj Srinivas, Usha Bhalla +1

Sparse autoencoders (SAEs) are commonly used to interpret the internal activations of large language models (LLMs) by mapping them to human-interpretable concept representations. W…

cs.AI2025

Computational Copyright: Towards A Royalty Model for Music Generative AI

Junwei Deng, Xirui Jiang, Shiyuan Zhang +5

The rapid rise of generative AI has intensified copyright and economic tensions in creative industries, particularly in music. Current approaches addressing this challenge often fo…

cs.CL2025

How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence

Hongzhe Du, Weikai Li, Min Cai +5

Post-training is essential for the success of large language models (LLMs), transforming pre-trained base models into more useful and aligned post-trained models. While plenty of w…

cs.LG2025

Inference-Time Reward Hacking in Large Language Models

Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling +2

A common paradigm to improve the performance of large language models is optimizing for a reward model. Reward models assign a numerical score to an LLM's output that indicates, fo…