8 citations · 9 across the 10 of their papers we have counts for
11 papers · 1 filter
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
Yiming Liang, Yizhi Li, Yantao Du +14
Benchmarks play a crucial role in tracking the rapid advancement of large language models (LLMs) and identifying their capability boundaries. However, existing benchmarks predomina…
Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
Tyler Loakman, Joseph James, Chenghua Lin
With the rise of Large Language Models (LLMs) and their vision-enabled counterparts (VLMs), numerous works have investigated their capabilities in tasks that fuse the modalities of…
COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes
Yunwen Li, Shuangshuang Ying, Xingwei Qu +16
Large language models exhibit systematic deficiencies in creative writing, particularly in non-English contexts where training data is scarce and lacks process-level supervision. W…
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
Shuangshuang Ying, Yunwen Li, Xingwei Qu +21
Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We intr…
A Survey on Latent Reasoning
Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng +30
Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, especially when guided by explicit chain-of-thought (CoT) reasoning that verbalizes intermediate s…
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
Yizhi Li, Ge Zhang, Hanhua Hong +2
As natural language processing for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques, such as pre-trained language models, suffer from…