2 citations · 4 across the 14 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Rethinking the Evaluation of Harness Evolution for Agents
Yike Wang, Huaisheng Zhu, Zhengyu Hu +7
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report…
cs.AI2025
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
Jihan Yao, Yushi Hu, Yujie Yi +9
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex…