most citedSee What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses

2 citations · 3 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Learning to Reason under Off-Policy Guidance

Jianhao Yan, Yafu Li, Zican Hu +5

Recent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcement learning wit…

cs.CL2025

RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction

Jianhao Yan, Yun Luo, Yue Zhang

In the multi-turn interaction schema, large language models (LLMs) can leverage user feedback to enhance the quality and relevance of their responses. However, evaluating an LLM's…

cs.CL2025

Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing

Zhilin Wang, Yafu Li, Jianhao Yan +2

Dynamical systems theory provides a framework for analyzing iterative processes and evolution over time. Within such systems, repetitive transformations can lead to stable configur…

cs.CL20241 cited

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

Jianhao Yan, Pingchuan Yan, Yulong Chen +3

This study presents a comprehensive evaluation of GPT-4's translation capabilities compared to human translators of varying expertise levels. Through systematic human evaluation us…

cs.CL20242 cited

See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses

Yulong Chen, Yang Liu, Jianhao Yan +6

The impressive performance of Large Language Models (LLMs) has consistently surpassed numerous human-designed benchmarks, presenting new challenges in assessing the shortcomings of…