6 papers
From Role Prompt to Infinite Thinking: Exploiting Persona Conditioning for Inference Cost Attacks in LLMs
Zhiyi Mou, Wangze Ni, Tianfang Xiao +6
LLMs are increasingly deployed in real-world applications, making inference efficiency and service reliability critical concerns due to their substantial computational costs. Howev…
SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation
Beichen Guo, Zhiyuan Wen, Jia Gu +6
Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialize…
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
Dingjie Song, Tianlong Xu, Yi-Fan Zhang +6
Assessing student handwritten scratchwork is crucial for personalized educational feedback but presents unique challenges due to diverse handwriting, complex layouts, and varied pr…
SRBench: A Comprehensive Benchmark for Sequential Recommendation with Large Language Models
Jianhong Li, Zeheng Qian, Wangze Ni +4
LLM development has aroused great interest in Sequential Recommendation (SR) applications. However, comprehensive evaluation of SR models remains lacking due to the limitations of…
FEANEL: A Benchmark for Fine-Grained Error Analysis in K-12 English Writing
Jingheng Ye, Shen Wang, Jiaqi Chen +9
Large Language Models (LLMs) have transformed artificial intelligence, offering profound opportunities for educational applications. However, their ability to provide fine-grained…
Protein as a Second Language for LLMs
Xinhui Chen, Zuchao Li, Mengqi Gao +4
Deciphering the function of unseen protein sequences is a fundamental challenge with broad scientific impact, yet most existing methods depend on task-specific adapters or large-sc…