3 papers
cs.LG2026
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention
Vishesh Tripathi, Abhay Kumar
Self-attention is central to Transformer performance and is often the most expensive part of the Transformer at long context lengths because its pairwise token interactions scale q…
cs.CL2025
The Instruction Gap: LLMs get lost in Following Instruction
Vishesh Tripathi, Uday Allu, Biddwan Ahmed
Large Language Models (LLMs) have shown remarkable capabilities in natural language understanding and generation, yet their deployment in enterprise environments reveals a critical…
cs.LG2025
Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding
Vishesh Tripathi, Tanmay Odapally, Indraneel Das +2
Retrieval-Augmented Generation (RAG) systems have revolutionized information retrieval and question answering, but traditional text-based chunking methods struggle with complex doc…