1 citations · 1 across the 8 of their papers we have counts for
8 papers
Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs
Jingshen Zhang, Shaoyang Xu, Wenxuan Zhang
Cultural fine-tuning has become the de facto paradigm for building culture-aware large language models (LLMs), yet existing optimization exclusively for alignment scores provides a…
Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
Long P. Hoang, Hai V. Le, Shaoyang Xu +2
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In…
Multilingual Fine-Tuning via Localized Gradient Conflict Resolution
Long P. Hoang, Yiran Zhao, Wei Lu +1
The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems. However, fine-tuning these models frequently…
What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems
Chen Huang, Yuhao Wu, Wenxuan Zhang
Multi-agent systems (MAS) built on large language models are typically organized around roles, pipelines, and turn schedules, while the content that agents pass to one another is o…
Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems
Shaoyang Xu, Jingshen Zhang, Long P. Hoang +2
Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural backgrounds. Existing cultural e…
Process Rewards with Learned Reliability
Jinyuan Li, Langlin Huang, Chengsong Huang +5
Process Reward Models (PRMs) provide step-level feedback for reasoning, but current PRMs usually output only a single reward score for each step. Downstream methods must therefore…