From the 1 of 18 linked papers with an AI index.
18 papers
Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment
Haokai Zhao, Yunze Xiao, Weihao Xuan +3
Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-pro…
ExpressionCueLens: A Cross-Cultural Analysis of Human-AI Companion Conversations on Social Media
Lynnette Hui Xian Ng, Yunze Xiao, Lionel Z. Wang +2
The paper presents ExpressionCueLens, a framework for categorizing anthropomorphic expressions in human‑AI companion conversations, and uses it to compare how Reddit and XiaoHongSh…
PACE: A Proxy for Agentic Capability Evaluation
Yueqi Song, Lintang Sutawika, Jiarui Liu +8
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars…
Knowledge Index of Noah's Ark
Sheng Jin, Minghao Liu, Yunze Xiao +24
Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consen…
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
Minglai Yang, Xinyan Velocity Yu, Pengyuan Li +22
Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OC…
Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs
Weijia Zhang, Ruiqi Chen, Yunze Xiao +1
Existing LLM moral benchmarks usually ask which isolated moral act, value, or foundation a model prefers. This is useful but incomplete. Realistic judgments often require a model t…