266 citations · 459 across the 47 of their papers we have counts for
41 papers · 1 filter
Kalman Delta Networks: Uncertainty-aware Associative Memory
Ngoc Bui, Tinglin Huang, Rex Ying
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requi…
FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space
Jiahong Liu, Ram Samarth B B, Xinyu Fu +4
Federated learning enables privacy-preserving collaborative training, but highly heterogeneous client data remain challenging, especially in graph federated learning where clients…
Hyperbolic Multimodal Continual Learning
Jiahong Liu, Ming Shen, Xiaohao Liu +4
Hyperbolic geometry has recently emerged as a powerful representation space for multimodal learning, as it naturally captures hierarchical semantic structure across modalities. Des…
Variational Learning for Insertion-based Generation
Yangtian Zhang, Zhe Wang, Arthur Gretton +4
Non-monotonic sequence generation methods, such as masked diffusion models, provide a flexible alternative to left-to-right autoregressive modeling by allowing tokens to be generat…
Reasoning through Verifiable Forecast Actions: Consistency-Grounded RL for Financial LLMs
Jialin Chen, Aosong Feng, Harshit Verma +7
Financial markets are characterized by extreme non-stationarity, low signal-to-noise ratios, and strong dependence on external information such as news, company fundamentals, and m…
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1
The key-value (KV) cache is a major bottleneck in long-context inference, where memory and computation grow with sequence length. Existing KV eviction methods reduce this cost but…