◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Andreas Madsen

Mila

5 papers hereh-index 7579 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • first author3
  • middle author1

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.CL4
  • cs.LG1
affiliations
  • Mila
HomepageORCID 0000-0002-1487-2796

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CLShow all

4 papers · 1 filter

cs.CL2026

Scaling Inherently Interpretable Language Models

Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7

Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult…

cs.CL2024

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing

Andreas Madsen

As machine learning becomes more widespread and is used in more critical applications, it's important to provide explanations for these models, to prevent unintended behavior. Unfo…

cs.CL2024

Faithfulness Measurable Masked Language Models

Andreas Madsen, Siva Reddy, Sarath Chandar

A common approach to explaining NLP models is to use importance measures that express which tokens are important for a prediction. Unfortunately, such explanations are often wrong…

cs.CL2024

Are self-explanations from Large Language Models faithful?

Andreas Madsen, Sarath Chandar, Siva Reddy

Instruction-tuned Large Language Models (LLMs) excel at many tasks and will even explain their reasoning, so-called self-explanations. However, convincing and wrong self-explanatio…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.