9 papers
TACOMORE: Exploring a replicable prompting protocol for LLM-assisted corpus analysis
Bingru Li, Han Wang, Nicholas Groom
As corpus linguistics continues to scale, researchers are facing a growing methodological bottleneck: while computational tools can easily count billions of words, the qualitative…
DecepChain: Inducing Deceptive Reasoning in Large Language Models
Wei Shen, Han Wang, Haoyu Li +1
Large Language Models (LLMs) have been demonstrating strong reasoning capability with their chain-of-thoughts (CoT), which are routinely used by humans to judge answer quality. Thi…
On The Fragility of Benchmark Contamination Detection in Reasoning Models
Han Wang, Haoyu Li, Brian Ko +1
Leaderboards for LRMs have turned evaluation into a competition, incentivizing developers to optimize directly on benchmark suites. A shortcut to achieving higher rankings is to in…
daVinci-Dev: Agent-native Mid-training for Software Engineering
Ji Zeng, Dayuan Fu, Tiantian Mi +14
Recently, the frontier of Large Language Model (LLM) capabilities has shifted from single-turn code generation to agentic software engineering-a paradigm where models autonomously…
JoyAgent-JDGenie: Technical Report on the GAIA
Jiarun Liu, Shiyue Xu, Shangkun Liu +13
Large Language Models are increasingly deployed as autonomous agents for complex real-world tasks, yet existing systems often focus on isolated improvements without a unifying desi…
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Junyu Zhang, Runpei Dong, Han Wang +8
This paper presents AlphaOne (1), a universal framework for modulating reasoning progress in large reasoning models (LRMs) at test time. 1 first introduces moment, whi…