2 papers
cs.DC2026
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models
Li Wang, Yi Su, Xiabao Wu +9
Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memo…
cs.AI2026
Text2GraphQuery-Bench: A Text to Graph Query Benchmark
Songlin Lyu, Lujie Ban, Zihang Wu +14
Graph models are fundamental to data analysis in domains rich with complex relationships. Unlike SQL, which benefits from a rel- atively unified standard and widespread familiarity…