activity
20232026
most citedMediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models

2 citations · 10 across the 12 of their papers we have counts for

collaborators

15 papers

cs.LG2026

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi

In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function. Concretely, we consider a…

cs.CV2026

MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI

Paula Arguello, Berk Tinaz, Mohammad Shahab Sepehri +2

Deep learning underpins a wide range of applications in MRI, including reconstruction, artifact removal, and segmentation. However, progress has been driven largely by public datas…

cs.CV2026

ATHENA: Adaptive Test-Time Steering for Improving Count Fidelity in Diffusion Models

Mohammad Shahab Sepehri, Asal Mehradfar, Berk Tinaz +2

Text-to-image diffusion models achieve high visual fidelity but surprisingly exhibit systematic failures in numerical control when prompts specify explicit object counts. To addres…

cs.LG2026

Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards

Reinhard Heckel, Mahdi Soltanolkotabi, Christos Thramboulidis

Reinforcement learning with verifiable rewards has driven recent advances in LLM post-training, in particular for reasoning. Policy optimization algorithms generate a number of res…

stat.ML2026

Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning

Filip Kovačević, Hong Chang Ji, Denny Wu +2

It is folklore that reusing training data more than once can improve the statistical efficiency of gradient-based learning. While this phenomenon has been extensively studied in li…

cs.CV2025

ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization

Haosheng Gan, Berk Tinaz, Mohammad Shahab Sepehri +2

Current text-to-image (T2I) benchmarks evaluate models on rigid prompts, potentially underestimating true generative capabilities due to prompt sensitivity and creating biases that…