3 papers
cs.AI2026
Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization
Suyash Maniyar, Armaan Sandhu, Abhishek Mishra
When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization. Endpoin…
cs.CL2026
Enhancing Legal LLMs through Metadata-Enriched RAG Pipelines and Direct Preference Optimization
Suyash Maniyar, Deepali Singh, Rohith Reddy
Large Language Models (LLMs) perform well in short contexts but degrade on long legal documents, often producing hallucinations such as incorrect clauses or precedents. In the lega…
cs.CV2025
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
Suyash Maniyar, Vishvesh Trivedi, Ajoy Mondal +2
Lecture slide element detection and retrieval are key problems in slide understanding. Training effective models for these tasks often depends on extensive manual annotation. Howev…