4 papers
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
Yufeng Du, Phillip Harris, Minyang Tian +5
We identify intrinsic limitations of Rotary Positional Embeddings (RoPE) in Transformer-based long-context language models. Our theoretical analysis abstracts away from the specifi…
Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
Minhui Zhu, Minyang Tian, Xiaocheng Yang +61
While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason effectively through complex, ope…
AttenGW: A Lightweight Attention-Based Multi-Detector Gravitational-Wave Detection Pipeline
Victoria Tiki, Eliu Huerta
We present AttenGW, an attention-based multi-detector gravitational-wave detection model and accompanying software stack designed for analysis of real LIGO data. AttenGW combines a…
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
Yufeng Du, Minyang Tian, Srikanth Ronanki +7
Large language models (LLMs) often fail to scale their performance on long-context tasks performance in line with the context lengths they support. This gap is commonly attributed…