activity
20242026
most citedRethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline

2 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.MA20262 cited

Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline

Jiawei Xu, Arief Koesdwiady, Sisong Bei +8

Recent advances in LLM-based multi-agent systems (MAS) show that workflows composed of multiple LLM agents with distinct roles, tools, and communication patterns can outperform sin…

cs.CL2025

Towards Effective Model Editing for LLM Personalization

Baixiang Huang, Limeng Cui, Jiapeng Liu +7

Personalization is becoming indispensable for LLMs to align with individual user preferences and needs. Yet current approaches are often computationally expensive, data-intensive,…

cs.LG2025

SAFE-D: A Spatiotemporal Detection Framework for Abnormal Driving Among Parkinson's Disease-like Drivers

Hangcheng Cao, Baixiang Huang, Longzhi Yuan +4

A driver's health state serves as a determinant factor in driving behavioral regulation. Subtle deviations from normalcy can lead to operational anomalies, posing risks to public t…

cs.AI2025

Who's Your Judge? On the Detectability of LLM-Generated Judgments

Dawei Li, Zhen Tan, Chengshuai Zhao +6

Large Language Model (LLM)-based judgments leverage powerful LLMs to efficiently evaluate candidate content and provide judgment scores. However, the inherent biases and vulnerabil…

cs.CL2025

Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm

Baixiang Huang, Zhen Tan, Haoran Wang +6

Agents based on Large Language Models (LLMs) have demonstrated strong capabilities across a wide range of tasks. However, deploying LLM-based agents in high-stakes domains comes wi…

cs.CL2024

Can Knowledge Editing Really Correct Hallucinations?

Baixiang Huang, Canyu Chen, Xiongxiao Xu +2

Large Language Models (LLMs) suffer from hallucinations, referring to the non-factual information in generated content, despite their superior capacities across tasks. Meanwhile, k…