1 paper
Ziang Duan, Jiajun Wu, Zetian Chen +8
Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separa…