4 papers
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention
Vishesh Tripathi, Abhay Kumar
Self-attention is central to Transformer performance and is often the most expensive part of the Transformer at long context lengths because its pairwise token interactions scale q…
The Instruction Gap: LLMs get lost in Following Instruction
Vishesh Tripathi, Uday Allu, Biddwan Ahmed
Large Language Models (LLMs) have shown remarkable capabilities in natural language understanding and generation, yet their deployment in enterprise environments reveals a critical…
Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding
Vishesh Tripathi, Tanmay Odapally, Indraneel Das +2
Retrieval-Augmented Generation (RAG) systems have revolutionized information retrieval and question answering, but traditional text-based chunking methods struggle with complex doc…
Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS
Onkar Kishor Susladkar, Vishesh Tripathi, Biddwan Ahmed
This research introduces a comprehensive Bahasa text-to-speech (TTS) dataset and a novel TTS model, EnGen-TTS, designed to enhance the quality and versatility of synthetic speech i…