4 papers
Trust the Mass: Forced Weights in KV-Cache Eviction
Jack Shi, Jerry Gu
Every deployed sparse-attention or KV-cache-eviction rule keeps a subset of the keys, discards the rest, and renormalizes the attention weights over the kept set. Enumerating the e…
Reinforcement learning to improve large language model-based automated code compliance systems
Jack Wei Lun Shi, Minghao Dang, Wawan Solihin +2
Large language model (LLM)-based approaches for automated code compliance (ACC) of building regulations are prone to generating incorrect and hallucinated computer-processable rule…
LLM attribution analysis across different fine-tuning strategies and model scales for automated code compliance
Jack Wei Lun Shi, Minghao Dang, Wawan Solihin +1
Existing research on large language models (LLMs) for automated code compliance has primarily focused on performance, treating the models as black boxes and overlooking how trainin…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…