works on

From the 4 of 36 linked papers with an AI index.

activity
20242026
most citedDistributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models

1 citations · 1 across the 29 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios

Xinyi Li, Zhen Fang, Yongxin Deng +12

Hallucination detection is essential for the reliable deployment of large language models (LLMs). However, existing evaluations face two core challenges: inconsistent inference con…

cs.CL2026

DECOR: Auditing LLM Deception via Information Manipulation Theory

Linyue Cai, Samuel Yeh, Jwala Dhamala +2

Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such behavior difficult to detec…

cs.CL2026

Auditing Agent Harness Safety

Chengzhi Liu, Yichen Guo, Yepeng Liu +8

LLM agents increasingly run inside execution harnesses that dispatch tools, allocate resources, and route messages between specialized components. However, a harness can return a c…

cs.CL2026

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability

Shawn Im, Changdae Oh, Zhen Fang +1

Semantic associations such as the link between "bird" and "flew" are foundational for language modeling as they enable models to go beyond memorization and instead generalize and g…

cs.CL2026

On Safety Risks in Experience-Driven Self-Evolving Agents

Weixiang Zhao, Yichen Zhang, Yingshuo Wang +8

Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduc…

cs.CL2026

How Retrieved Context Shapes Internal Representations in RAG

Samuel Yeh, Sharon Li

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by conditioning generation on retrieved external documents, but the effect of retrieved context is often…