papers

Publications (12)

cs.CR2025

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Kun Wang, Guibin Zhang, Zhenhong Zhou +100

The remarkable success of Large Language Models (LLMs) has illuminated a promising pathway toward achieving Artificial General Intelligence for both academic and industrial communi…

cs.CL2026

DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognition

Hanjun Luo, Yingbin Jin, Xinfeng Li +6

The advancements of Large Language Models (LLMs) have spurred a growing interest in their application to Named Entity Recognition (NER) methods. However, existing datasets are prim…

cs.SD2026

AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

Kai Li, Can Shen, Yile Liu +31

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) demand rigorous evaluation of their trustworthiness. However, existing evaluation frameworks ar…

cs.AI2026

AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters

Hanjun Luo, Zhimu Huang, Sylvia Chung +6

Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet…

cs.CV2025

FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models

Hanjun Luo, Ziye Deng, Ruizhe Chen +1

The rapid development and reduced barriers to entry for Text-to-Image (T2I) models have raised concerns about the biases in their outputs, but existing research lacks a holistic de…

cs.SE2026

CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding

Hanjun Luo, Chiming Ni, Jiaheng Wen +9

LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to captur…

cs.HC2026

PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction

Jialin Li, Zhenhao Chen, Hanjun Luo +1

LLM-based agents can complete tasks correctly yet still frustrate users through poor interaction patterns, such as excessive confirmations, opaque reasoning, or misaligned pacing.…

cs.CV2024

VersusDebias: Universal Zero-Shot Debiasing for Text-to-Image Models via SLM-Based Prompt Engineering and Generative Adversary

Hanjun Luo, Ziye Deng, Haoyu Huang +3

With the rapid development of Text-to-Image (T2I) models, biases in human image generation against demographic social groups become a significant concern, impacting fairness and et…

cs.CY2026

BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models

Hanjun Luo, Zhimu Huang, Haoyu Huang +5

Text-to-Image (T2I) generative models have revolutionized content creation, yet they inherently risk amplifying societal biases. While sociological research provides systematic cla…

cs.CV2023

UniAP: Towards Universal Animal Perception in Vision via Few-shot Learning

Meiqi Sun, Zhonghan Zhao, Wenhao Chai +5

Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is…

cs.CV2025

BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models

Hanjun Luo, Haoyu Huang, Ziye Deng +6

Text-to-Image (T2I) generative models are becoming increasingly crucial due to their ability to generate high-quality images, but also raise concerns about social biases, particula…

cs.AI2026

AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents

Hanjun Luo, Shenyu Dai, Chiming Ni +5

Despite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators…