Publications (13)
Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
Jiaming Qu, Lucheng Fu, Yibo Hu
Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answe…
PACE: Two-Timescale Self-Evolution for Small Language Model Agents
Chen Ling, Pei Chen, Albert Guan +4
Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other components of the agent pipeline.…
Metaphors We Compute By: A Computational Audit of Cultural Translation vs. Thinking in LLMs
Yuan Chang, Jiaming Qu, Zhu Li
Large language models (LLMs) are often described as multilingual because they can understand and respond in many languages. However, speaking a language is not the same as reasonin…
AGWM: Affordance-Grounded World Models for Environments with Compositional Prerequisites
Qinshi Zhang, Weipeng Deng, Zhihan Jiang +4
In model-based learning, the agent learns behaviors by simulating trajectories based on world model predictions. Standard world models typically learn a stationary transition funct…
A Medical Literature Search System for Identifying Effective Treatments in Precision Medicine
Jiaming Qu, Yue Wang
The Precision Medicine Initiative states that treatments for a patient should take into account not only the patient's disease, but his/her specific genetic variation as well. The…
Social Pressure Breaks Majority Voting in LLM Safety Panels
Yibo Hu, Jiaming Qu
Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but th…
A Multi-Agent Framework for Democratizing XR Content Creation in K-12 Classrooms
Yuan Chang, Zhu Li, Jiaming Qu
Generative AI (GenAI) combined with Extended Reality (XR) offers potential for K-12 education, yet classroom adoption remains limited by the high technical barrier of XR content au…
Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
Yibo Hu, Jiaming Qu
LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even a…
Touching Space: Accessible Map Exploration Through Conversational Audio-Haptic Interaction
Li Liu, Jiaming Qu, Marc Jowell Bagaoisan +2
Most existing assistive navigation tools focus on providing real-time guidance for Blind and Low-Vision (BLV) people, but few support building a holistic spatial understanding of u…
Lighting Up or Dimming Down? Exploring Dark Patterns of LLMs in Co-Creativity
Zhu Li, Jiaming Qu, Yuan Chang
Large language models (LLMs) are increasingly acting as collaborative writing partners, raising questions about their impact on human agency. In this exploratory work, we investiga…
Why is "Problems" Predictive of Positive Sentiment? A Case Study of Explaining Unintuitive Features in Sentiment Classification
Jiaming Qu, Jaime Arguello, Yue Wang
Explainable AI (XAI) algorithms aim to help users understand how a machine learning model makes predictions. To this end, many approaches explain which input features are most pred…
Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text
Hongbo Du, Zixin Lu, Jiaming Qu
Large language models (LLMs) are increasingly used for clinical text tasks such as summarization and revision. While most studies evaluate the fluency and coherence of LLM-generate…
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
Jiaming Qu, Mengtian Guo, Yue Wang
Deceptive reviews mislead consumers, harm businesses, and undermine trust in online marketplaces. Machine learning classifiers can learn from large amounts of data to distinguish d…