3 papers
cs.LG2025
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
Raghav Sharma, Manan Mehta, Sai Tiger Raina
Reinforcement Learning from Human Feedback (RLHF) is the standard for aligning Large Language Models (LLMs), yet recent progress has moved beyond canonical text-based methods. This…
cs.AI2025
Adaptive and Explainable AI Agents for Anomaly Detection in Critical IoT Infrastructure using LLM-Enhanced Contextual Reasoning
Raghav Sharma, Manan Mehta
Ensuring that critical IoT systems function safely and smoothly depends a lot on finding anomalies quickly. As more complex systems, like smart healthcare, energy grids and industr…
cs.AI2025
Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
Raghav Sharma, Manan Mehta
Small language models (SLMs; 1-12B params, sometimes up to 20B) are sufficient and often superior for agentic workloads where the objective is schema- and API-constrained accuracy…