most citedSafeWorld: Geo-Diverse Safety Alignment

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2025

MMGR: Multi-Modal Generative Reasoning

Zefan Cai, Haoyi Qiu, Tianyi Ma +9

Video foundation models generate visually realistic and temporally coherent content, but their reliability as world simulators depends on whether they capture physical, logical, an…

cs.CL2025

From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models

Zefan Cai, Haoyi Qiu, Haozhe Zhao +6

Recent advances in video diffusion models have significantly enhanced text-to-video generation, particularly through alignment tuning using reward models trained on human preferenc…

cs.CL2025

GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness

Kung-Hsiang Huang, Haoyi Qiu, Yutong Dai +2

Graphical user interface (GUI) agents built on vision-language models have emerged as a promising approach to automate human-computer workflows. However, they also face the ineffic…

cs.CL2025

Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies

Haoyi Qiu, Kung-Hsiang Huang, Ruichen Zheng +2

Large vision-language models (LVLMs) are increasingly deployed in globally distributed applications, such as tourism assistants, yet their ability to produce culturally appropriate…

cs.AI2025

Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding

Kung-Hsiang Huang, Can Qin, Haoyi Qiu +4

Vision Language Models (VLMs) have achieved remarkable progress in multimodal tasks, yet they often struggle with visual arithmetic, seemingly simple capabilities like object count…

cs.CL20241 cited

SafeWorld: Geo-Diverse Safety Alignment

Da Yin, Haoyi Qiu, Kung-Hsiang Huang +2

In the rapidly evolving field of Large Language Models (LLMs), ensuring safety is a crucial and widely discussed topic. However, existing works often overlook the geo-diversity of…