Publications (12)
Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens
Ziqian Zeng, Jiahong Yu, Qianshi Pang +4
Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their widespread application is hindered by the resource-intensive decoding pr…
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
Peng Chen, Jiaji Zhang, Hailiang Zhao +11
In modern GPU inference, cache efficiency remains a major bottleneck, and heuristic policies such as \textsc{LRU} can perform far worse than the offline optimum. Existing learning-…
Almost Vector Bundles over Perfectoid Spaces
Yuntong Cui, Guo Li, Shuhan Jiang +1
In this paper, we define vector bundles within the framework of almost mathematics (referred to as almost vector bundles) and establish the -descent theorem together with a stru…
Lift-independence problem in the -adic Simpson correspondence for curves
Xiangyu Pan, Jiahong Yu
Let be a proper smooth rigid analytic variety over a complete algebraically closed field -adic field . Fix an continuation of . Faltings (in…
A Conjecture of Bhatt--Lurie and weakly -nilpotent Hodge--Tate stacks
Jiahong Yu
Let be a perfect field of characteristic , and let be a smooth variety. It is known that given a Frobenius lifting of , we can identify prismatic crystals and nilpo…
Prismatic Crystals for schemes in characteristic
Jiahong Yu
Let be a crystalline prism and let be a finite type -scheme admitting a Koszul-regular closed immersion into a smooth formal -scheme . We constr…