7 papers
Workload-Aware Caching for Multi-Agent Systems
Anas Mohamed, Kaizan Haque, Azal Ahmad Khan +3
Multi-agent systems decompose complex tasks into directed acyclic graphs (DAGs) of specialized agent executions, creating natural opportunities for caching intermediate results acr…
Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing
Azal Ahmad Khan, Ammar Ahmed, Zeshan Fayyaz +3
Synchronous reinforcement learning methods such as Group Relative Policy Optimization (GRPO) provide stable and reproducible on-policy training, but they are highly vulnerable to s…
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
Ammar Ahmed, Azal Ahmad Khan, Ayaan Ahmad +3
Large reasoning models improve accuracy by producing long reasoning traces, but this inflates latency and cost, motivating inference-time efficiency. We propose Retrieval-of-Though…
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
Seyyed Saeid Cheshmi, Azal Ahmad Khan, Xinran Wang +2
Large Language Models (LLMs) are increasingly relied upon for solving complex reasoning tasks in domains such as mathematics, logic, and multi-step question answering. A growing li…
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
Anas Mohamed, Azal Ahmad Khan, Xinran Wang +5
Generative AI can now synthesize strikingly realistic images from text, yet output quality remains highly sensitive to how prompts are phrased. Direct Preference Optimization (DPO)…
Safety Aware Task Planning via Large Language Models in Robotics
Azal Ahmad Khan, Michael Andrev, Muhammad Ali Murtaza +5
The integration of large language models (LLMs) into robotic task planning has unlocked better reasoning capabilities for complex, long-horizon workflows. However, ensuring safety…