works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.IR2026

Can Argus Judge Them All? Comparing VLMs Across Domains

Harsh Joshi, Gautam Siddharth Kashyap, Rafiq Ali +5

The paper introduces ARGUS-EVAL, a framework that assesses vision-language models on both capability and reliability across domains, and uses it to compare several VLMs on retrieva…

cs.CL2026

Truth, Trust, and Trouble: Medical AI on the Edge

Mohammad Anas Azeez, Rafiq Ali, Ebad Shabbir +4

Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering. However, ensuring these models meet critical…

cs.CL2026

Are Large Language Models Economically Viable for Industry Deployment?

Abdullah Mohammad, Sushant Kumar Ray, Pushkar Arora +5

Generative AI-powered by Large Language Models (LLMs)-is increasingly deployed in industry across healthcare decision support, financial analytics, enterprise retrieval, and conver…

cs.CL2026

Are Aligned Large Language Models Still Misaligned?

Usman Naseem, Gautam Siddharth Kashyap, Rafiq Ali +4

Misalignment in Large Language Models (LLMs) arises when model behavior diverges from human expectations and fails to simultaneously satisfy safety, value, and cultural dimensions,…

cs.CL2026

Can Large Language Models Make Everyone Happy?

Usman Naseem, Gautam Siddharth Kashyap, Ebad Shabbir +3

Misalignment in Large Language Models (LLMs) refers to the failure to simultaneously satisfy safety, value, and cultural dimensions, leading to behaviors that diverge from human ex…

cs.CL2026

Do Large Language Models Reflect Demographic Pluralism in Safety?

Usman Naseem, Gautam Siddharth Kashyap, Sushant Kumar Ray +3

Large Language Model (LLM) safety is inherently pluralistic, reflecting variations in moral norms, cultural expectations, and demographic contexts. Yet, existing alignment datasets…