activity
20232026
most citedProbabilistic Transformer: A Probabilistic Dependency Model for Contextual Word Representation

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

Renwei Meng, Bowen Zhang, Jian Wang +4

Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, o…

cs.CL2026

GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs

Junjie Peng, You Wu, Haoyi Wu +4

Large language models (LLMs) with extended context lengths rely on the key-value (KV) cache to support attention over prior tokens. However, maintaining the KV cache incurs substan…

cs.CL2026

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer

Penghao Kuang, Haoyi Wu, Kewei Tu

Probabilistic Transformer (PT), a white-box probabilistic model for contextual word representation, has demonstrated substantial similarity to standard Transformers in both computa…

cs.CL2025

Parallel Continuous Chain-of-Thought with Jacobi Iteration

Haoyi Wu, Zhihao Teng, Kewei Tu

Continuous chain-of-thought has been shown to be effective in saving reasoning tokens for large language models. By reasoning with continuous latent thought tokens, continuous CoT…

cs.CL2024

A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference

You Wu, Haoyi Wu, Kewei Tu

Recently, sharing key-value (KV) cache across layers has been found effective in efficient inference of large language models (LLMs). To systematically investigate different techni…

cs.CL2024

Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Haoyi Wu, Kewei Tu

Huge memory consumption has been a major bottleneck for deploying high-throughput large language models in real-world applications. In addition to the large number of parameters, t…