works on

From the 1 of 18 linked papers with an AI index.

collaborators

18 papers

cs.CL2026

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

Tianci Liu, Zihan Dong, Tianchun Li +8

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a…

cs.CR2026

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4

A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent…

cs.AI2026

How Benchmarks Mis-Score Computer-Use Agents

Zihan Dong, Zhiyuan Ma, Zekun Wang +5

The paper examines how current benchmarks for computer-use agents often give inaccurate scores due to issues in task design, trajectory observation, scoring, and reporting, and pro…

cs.CL2026

RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation

Pei Tian, Zihan Dong, Tianci Liu +2

Small-scale language models (SLMs) are attractive for retrieval-augmented generation (RAG) in resource-constrained settings, but their limited capacity makes them highly sensitive…

cs.AI2026

Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising

Tianci Liu, Zihan Dong, Linjun Zhang +6

Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespeci…

cs.LG2026

Decompose Sparsely Where You Should, Absorb Densely Where You Should No

Ruixuan Deng, Zehao Jin, Zekun Wang +1

Sparse autoencoders (SAEs) are typically trained to reconstruct the \textbf{entire} residual stream through a sparse dictionary, implicitly assuming that all activation content is…