2 papers
cs.CL2025
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
Danush Khanna, Aditya Kumar Guru, Srivarshinee Sridhar +7
Inference accounts for the majority of latency and energy consumption in large language model (LLM) deployments, often exceeding 90% of total cost. While training-time efficiency h…
cs.CL2025
Bridging Emotions and Architecture: Sentiment Analysis in Modern Distributed Systems
Mahak Shah, Akaash Vishal Hazarika, Meetu Malhotra +2
Sentiment analysis is a field within NLP that has gained importance because it is applied in various areas such as; social media surveillance, customer feedback evaluation and mark…