5 citations · 10 across the 14 of their papers we have counts for
14 papers
Beyond Top- Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
Wang Wei, Tiankai Yang, Samyadeep Basu +7
Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill registries is difficult because many skills are functionally redu…
Personalized Auto-Research: Towards a True AI Co-Scientist
Bo Ni, Franck Dernoncourt, Hongjie Chen +5
AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out. Despi…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in…
A Survey on LLM-based Conversational User Simulation
Bo Ni, Leyao Wang, Yu Wang +27
User simulation has long played a vital role in computer science due to its potential to support a wide range of applications. Language, as the primary medium of human communicatio…
Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling
Ruijie Ye, Jiayi Zhang, Zhuoxin Liu +10
We study instruction-based image editing under professional workflows and identify three persistent challenges: (i) editors often over-edit, modifying content beyond the user's int…
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
Runzhou Liu, Hailey Weingord, Sejal Mittal +18
Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important…