collaborators

6 papers

cs.CL2025

Semi-structured LLM Reasoners Can Be Rigorously Audited

Jixuan Leng, Cassandra A. Cohen, Zhixian Zhang +2

Although Large Language Models (LLMs) have become capable reasoners, the problem of faithfulness persists: their reasoning can contain errors and omissions that are difficult to de…

cs.CL2025

CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation

Jixuan Leng, Chengsong Huang, Langlin Huang +4

Existing reasoning evaluation frameworks for Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) predominantly assess either text-based reasoning or vision-langua…

cs.CL2025

POSS: Position Specialist Generates Better Draft for Speculative Decoding

Langlin Huang, Chengsong Huang, Jixuan Leng +2

Speculative decoding accelerates Large Language Model (LLM) inference by using a small draft model to predict multiple tokens, and a large target model to verify these tokens in pa…

cs.CL2025

Taming Overconfidence in LLMs: Reward Calibration in RLHF

Jixuan Leng, Chengsong Huang, Banghua Zhu +1

Language model calibration refers to the alignment between the confidence of the model and the actual performance of its responses. While previous studies point out the overconfide…

cs.LG2025

Efficient Test-Time Scaling via Self-Calibration

Chengsong Huang, Langlin Huang, Jixuan Leng +2

Increasing test-time computation is a straightforward approach to enhancing the quality of responses in Large Language Models (LLMs). While Best-of-N sampling and Self-Consistency…

cs.LG2024

SFT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity

Xinyu Yang, Jixuan Leng, Geyang Guo +5

Current PEFT methods for LLMs can achieve either high quality, efficient training, or scalable serving, but not all three simultaneously. To address this limitation, we investigate…