1 citations · 1 across the 3 of their papers we have counts for
13 papers
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun +64
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…
Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
Jinhao Pan, Chahat Raj, Anjishnu Mukherjee +4
Large language models (LLMs) exhibit social biases that reinforce harmful stereotypes, limiting their safe deployment. Most existing debiasing methods adopt a suppressive paradigm…
Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations
Mamnuya Rinki, Chahat Raj, Anjishnu Mukherjee +1
Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multilingual regions like South Asia. Thi…
Talent or Luck? Evaluating Attribution Bias in Large Language Models
Chahat Raj, Mahika Banerjee, Jinhao Pan +3
When a student fails an exam, do we tend to blame their effort or the test's difficulty? Attribution, defined as how reasons are assigned to event outcomes, shapes perceptions, rei…
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
Chahat Raj, Bowen Wei, Aylin Caliskan +2
While bias in large language models (LLMs) is well-studied, similar concerns in vision-language models (VLMs) have received comparatively less attention. Existing VLM bias studies…
Metadata Conditioned Large Language Models for Localization
Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos
Large language models are typically trained by treating text as a single global distribution, often resulting in geographically homogenized behavior. We study metadata conditioning…