12 papers · 1 filter
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors
Chenghao Yang, Yuning Zhang, Zhoufutu Wen +4
Model distillation is a primary driver behind the rapid progress of LLM agents, yet it often leads to behavioral homogenization. Many emerging agents share nearly identical reasoni…
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
Zexuan Wang, Chenghao Yang, Yingqi Que +18
Real-world autonomous planning requires coordinating tightly coupled constraints where a single decision dictates the feasibility of all subsequent actions. However, existing bench…
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
Shuangshuang Ying, Yunwen Li, Xingwei Qu +21
Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We intr…
MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity
Kaiyuan Zhang, Chenghao Yang, Zhoufutu Wen +19
As reasoning models scale rapidly, the essential role of multimodality in human cognition has come into sharp relief, driving a growing need to probe vision-centric cognitive behav…
DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains
Xiying Zhao, Zhoufutu Wen, Zhixuan Chen +20
The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. Whi…
COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes
Yunwen Li, Shuangshuang Ying, Xingwei Qu +16
Large language models exhibit systematic deficiencies in creative writing, particularly in non-English contexts where training data is scarce and lacks process-level supervision. W…