8 papers
SIQA: Toward Reliable Scientific Image Quality Assessment
Wenzhe Li, Liang Chen, Junying Wang +6
Scientific images fundamentally differ from natural and AI-generated images in that they encode structured domain knowledge rather than merely depict visual scenes. Assessing their…
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
Minheng Ni, Yutao Fan, Zhengyuan Yang +6
Recent advances in large multimodal models (LMMs) have enabled instruction-based image editing, allowing users to modify visual content via natural language descriptions. However,…
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
Qi Jia, Ye Shen, Xiujie Song +5
Evaluating LLMs' instruction-following ability in multi-topic dialogues is essential yet challenging. Existing benchmarks are limited to a fixed number of turns, susceptible to sat…
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Ye Shen, Dun Pei, Yiqiu Guo +6
Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…
QoNext: Towards Next-generation QoE for Foundation Models
Yijin Guo, Zicheng Zhang, Ye Shen +4
Existing evaluations of foundation models, including recent human-centric approaches, fail to capture what truly matters: user's experience during interaction. Current methods trea…
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
Junying Wang, Zicheng Zhang, Ye Shen +8
High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…