collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

Yanlin Fei, Nazhou Liu, Xinmiao Yu +7

AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an…

cs.CL2026

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

Shaolong Chen, Madalina Ciobanu, Qingqing Mao +1

DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loop. Recent theoretical analys…

cs.CL2026

State-of-the-Art Arabic Language Modeling with Sparse MoE Fine-Tuning and Chain-of-Thought Distillation

Navan Preet Singh, Anurag Garikipati, Ahmed Abulkhair +6

This paper introduces Arabic-DeepSeek-R1, an application-driven open-source Arabic LLM that leverages a sparse MoE backbone to address the digital equity gap for under-represented…

cs.CL2026

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning

Navan Preet Singh, Xiaokun Wang, Anurag Garikipati +3

We present an innovative multi-stage optimization strategy combining reinforcement learning (RL) and supervised fine-tuning (SFT) to enhance the pedagogical knowledge of large lang…

cs.CL2024

OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models

Jenish Maharjan, Anurag Garikipati, Navan Preet Singh +7

LLMs have become increasingly capable at accomplishing a range of specialized-tasks and can be utilized to expand equitable access to medical knowledge. Most medical LLMs have invo…