6 papers · 1 filter
A Survey on Diffusion Language Models
Tianyi Li, Mingda Chen, Bowei Guo +1
Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm. By generating tokens in parallel through…
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
Jiacheng Guo, Suozhi Huang, Zixin Yao +16
This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large Language Model (LLM) agents in t…
Sink-Aware Pruning for Diffusion Language Models
Aidar Myrzakhan, Tianyi Li, Bowei Guo +2
Diffusion Language Models (DLMs) incur high inference cost due to iterative denoising, motivating efficient pruning. Existing pruning heuristics largely inherited from autoregressi…
Estimating the Error of Large Language Models at Pairwise Text Comparison
Tianyi Li
We measure LLMs' output error at pairwise text comparison, noting the probability of error in their preferences. Our method does not rely on the ground truth and supports two scena…
A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
Tianyi Li, Yu Qin, Olivia R. Liu Sheng
How much large language models (LLMs) can aid scientific discovery, notably in assisting academic peer review, is in heated debate. Between a literature digest and a human-comparab…
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
Tianyi Li, Erenay Dayanik, Shubhi Tyagi +1
In this paper, we present HalluCana, a canary lookahead to detect and correct factuality hallucinations of Large Language Models (LLMs) in long-form generation. HalluCana detects a…