From the 2 of 12 linked papers with an AI index.
6 papers · 1 filter
UXBench: Benchmarking User Experience in AI Assistants
Mengze Hong, Xia Zeng, Zeyang Lei +26
UXBench is a user‑centric benchmark that uses real interaction logs to evaluate how well AI assistants align with user preferences and generate engaging dialogue, featuring three t…
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
Jiamin Chen, Yidi Wu, Qiexiang Wang +6
Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than construct…
DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation
Jiamin Chen, Qianben Chen, Jiawen Zhang +5
Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinematic control, audio, and cros…
PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
Bowei He, Lihao Yin, Hui-Ling Zhen +3
Model pruning is an effective approach for compressing large language models (LLMs). However, this process often leads to significant degradation of model capabilities. While post-…
Grounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation
Jiamin Chen, Yuchen Li, Xinyu Ma +5
Retrieval-Augmented Generation (RAG) has become an essential approach for extending the reasoning and knowledge capacity of large language models (LLMs). While prior research has p…
Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
Bowei He, Lihao Yin, Huiling Zhen +5
Post-training compression has been a widely employed approach to scale down large language model (LLM) and facilitate efficient inference. In various proposed compression methods,…