1 citations · 1 across the 4 of their papers we have counts for
3 papers · 1 filter
Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation
Ting Fang Tan, Kabilan Elangovan, Andreas Pollreisz +13
Domain specific large language models are increasingly used to support patient education, triage, and clinical decision making in ophthalmology, making rigorous evaluation essentia…
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
Zhihan Liu, Lin Guan, Yixin Nie +6
Generalist LLM agents are often post-trained on a narrow set of environments but deployed across far broader, unseen domains. In this work, we investigate the challenge of agentic…
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
Ruijie Xu, Zhihan Liu, Yongfei Liu +4
We address the challenge of online Reinforcement Learning from Human Feedback (RLHF) with a focus on self-rewarding alignment methods. In online RLHF, obtaining feedback requires i…