3 papers
cs.CL2026
MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs
Saman Sarker Joy, Niloy Farhan
Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than pressured patient-facing con…
cs.CL2026
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali
Saman Sarker Joy, Swakkhar Shatabda
Large-scale multitask benchmarks have driven rapid progress in language modeling, yet most emphasize high-resource languages such as English, leaving Bengali underrepresented. We p…
cs.CV2025
Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation
Tanjim Islam Riju, Shuchismita Anwar, Saman Sarker Joy +2
Medical vision-language models still struggle to match radiologists' attention and to verbalize findings with explicit spatial grounding. We address this gap with a two-stage multi…