activity
20242026
most cited100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

1 citations · 1 across the 2 of their papers we have counts for

collaborators

11 papers

cs.AI2026

Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System

Shaochen Zhong

With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the…

cs.CL20261 cited

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

Wang Yang, Hongye Jin, Shaochen Zhong +4

Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effortlessly process many originally exhaust…

cs.CL2026

AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models

Feng Luo, Yu-Neng Chuang, Guanchu Wang +8

Reasoning-capable large language models (LLMs) achieve strong performance on complex tasks but often exhibit overthinking after distillation, generating unnecessarily long chain-of…

cs.LG2026

70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)

Tianyi Zhang, Mohsen Hariri, Shaochen Zhong +4

Large-scale AI models, such as Large Language Models (LLMs) and Diffusion Models (DMs), have grown rapidly in size, creating significant challenges for efficient deployment on reso…

cs.CL2025

Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly

Wenya Xie, Shaochen, Zhong +4

Large Reasoning Models (LRMs) are often bottlenecked by the high cost of output tokens. We show that a significant portion of these tokens are useless self-repetitions - what we ca…

cs.LG2025

Quantize What Counts: More for Keys, Less for Values

Mohsen Hariri, Alan Luo, Weicong Chen +6

Large Language Models (LLMs) suffer inference-time memory bottlenecks dominated by the attention Key-Value (KV) cache, which scales with model size and context length. While KV-cac…