30 citations · 84 across the 69 of their papers we have counts for
17 papers · 1 filter
GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation
Yifan Chen, Haitao Li, Qingyao Ai +4
Large language models are increasingly used as scalable evaluators for open-ended tasks. However, many LLM judges derive query-specific criteria during scoring, leaving the evaluat…
Metis: Memory Foundation Model
Zeyu Zhang, Ziliang Guo, Yihang Sun +14
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reaso…
Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
Chengbing Wang, Yang Zhang, Wenjie Wang +4
Preference alignment has enabled large language models (LLMs) to better reflect human expectations, but current methods mostly optimize for population-level preferences, overlookin…
SteerX: Disentangled Steering for LLM Personalization
Xiaoyan Zhao, Ming Yan, Yilun Qiu +5
Large language models (LLMs) have shown remarkable success in recent years, enabling a wide range of applications, including intelligent assistants that support users' daily life a…
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
Fengbin Zhu, Xiang Yao Ng, Ziyang Liu +19
Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks.…
VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
Ziang Ye, Yang Zhang, Wentao Shi +3
Graphical User Interface (GUI) agents powered by Large Vision-Language Models (LVLMs) have emerged as a revolutionary approach to automating human-machine interactions, capable of…