activity
20242026
collaborators

8 papers

cs.CL2026

Debiasing Without Protected Attributes: Latent Concept Erasure from Textual Profiles

Shun Shao, Zheng Zhao, Anna Korhonen +2

Most fairness research in NLP assumes direct access to protected attributes such as gender, race, or nationality. In practice, however, such information is often unavailable due to…

cs.CV2026

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

Yifu Qiu, Yftah Ziser, Anna Korhonen +2

Can unified vision-language models (VLMs) perform forward dynamics prediction (FDP), i.e., predicting the future state (in image form) given the previous observation and an action…

cs.CL2026

Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer

Shun Shao, Binxu Wang, Shay B. Cohen +2

Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expensive, model-specific, and dif…

cs.LG2026

Self-Improving World Modelling with Latent Actions

Yifu Qiu, Zheng Zhao, Waylon Li +4

Internal modelling of the world -- predicting transitions between previous states and next states under actions -- is essential to reasoning and planning for LLMs and V…

cs.CL2026

Bolmo: Byteifying the Next Generation of Language Models

Benjamin Minixhofer, Tyler Murray, Tomasz Limisiewicz +6

Recent advances in generative AI have been largely driven by large language models (LLMs), deep neural networks that operate over discrete units called tokens. To represent text, t…

cs.CL2025

Iterative Multilingual Spectral Attribute Erasure

Shun Shao, Yftah Ziser, Zheng Zhao +3

Multilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between langu…