activity
20212026
most citedOpenAssistant Conversations -- Democratizing Large Language Model Alignment

135 citations · 145 across the 17 of their papers we have counts for

collaborators

17 papers

cs.AI2026

Joint Optimization of Tool Creation and Use for Large Language Model Agents

Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen +2

Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the…

cs.IR2026

ICICLE: Expanding Retrieval with In-Context Documents

Yu-Chen Den, Yung-Yu Shih, Zhi Rui Tam +4

Generative retrieval (GR) maps queries directly to document identifiers (docids) using parametric knowledge, However, this design makes corpus expansion costly: adding new document…

cs.GR2026

ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks

Samin Mahdizadeh Sani, Max Ku, Nima Jamali +23

Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet,…

cs.CR2026

Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs

Yen-Shan Chen, Zhi Rui Tam, Cheng-Kuang Wu +1

Current evaluations of LLM safety predominantly rely on severity-based taxonomies to assess the harmfulness of malicious queries. We argue that this formulation requires re-examina…

cs.CL2025

MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making

Zhi Rui Tam, Yun-Nung Chen

As large language models transition from text-based interfaces to audio interactions in clinical settings, they might introduce new vulnerabilities through paralinguistic cues in a…

cs.CL2025

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…