4 papers
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
Anubhab Banerjee
Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameterized by step count, tool invocations, and…
When Words Predict Workload
Anubhab Banerjee
Standard distributed \ac{llm} schedulers rely on static token counts or rolling latency averages, making them susceptible to failures on statutorily constrained text. On \ac{epo} c…
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
Anubhab Banerjee
The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted content under Article 83 and Rul…
Inductive Latent Context Persistence: Closing the Post-Handover Cold Start in 6G Radio Access Networks
Anubhab Banerjee, Daniyal Amir Awan
In modern radio access networks (RANs), rule-based handover (HO) decisions (e.g., A3/A5) depend on user equipment (UE) measurements only, so UEs at the same location can receive in…