Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Memory Retrieval in Transformers: Insights from The Encoding Specificity Principle
Viet Hung Dinh, Ming Ding, Youyang Qu +1
While explainable artificial intelligence (XAI) for large language models (LLMs) remains an evolving field with many unresolved questions, increasing regulatory pressures have spur…
cs.LG2025
Virtual Width Networks
Seed, Baisheng Li, Banggu Wu +115
We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN d…