works on

From the 1 of 18 linked papers with an AI index.

collaborators

18 papers

cs.CL2026

Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

Haokai Zhao, Yunze Xiao, Weihao Xuan +3

Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-pro…

cs.HC2026

ExpressionCueLens: A Cross-Cultural Analysis of Human-AI Companion Conversations on Social Media

Lynnette Hui Xian Ng, Yunze Xiao, Lionel Z. Wang +2

The paper presents ExpressionCueLens, a framework for categorizing anthropomorphic expressions in human‑AI companion conversations, and uses it to compare how Reddit and XiaoHongSh…

cs.AI2026

PACE: A Proxy for Agentic Capability Evaluation

Yueqi Song, Lintang Sutawika, Jiarui Liu +8

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars…

cs.AI2026

Knowledge Index of Noah's Ark

Sheng Jin, Minghao Liu, Yunze Xiao +24

Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consen…

cs.CL2026

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

Minglai Yang, Xinyan Velocity Yu, Pengyuan Li +22

Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OC…

cs.CL2026

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs

Weijia Zhang, Ruiqi Chen, Yunze Xiao +1

Existing LLM moral benchmarks usually ask which isolated moral act, value, or foundation a model prefers. This is useful but incomplete. Realistic judgments often require a model t…