11 papers
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps
Xiaofeng Lin, Yukai Yang, Daniel Guo +3
Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak defenses must reason about cr…
REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces
Xiaofeng Lin, Yingxu Wang, Tung Sum Thomas Kwok +4
Large language model (LLM) agents now solve complex tasks through long plan-and-execution traces, yet the ability to locate errors in a completed traces still lags far behind, espe…
ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement Learning
Xiaofeng Lin, Seungbae Kim, Zhuoya Li +3
Deep generative models can help with data scarcity and privacy by producing synthetic training data, but they struggle in low-data, imbalanced tabular settings to fully learn the c…
Finding Connections: Membership Inference Attacks for the Multi-Table Synthetic Data Setting
Joshua Ward, Chi-Hua Wang, Guang Cheng
Synthetic tabular data has gained attention for enabling privacy-preserving data sharing. While substantial progress has been made in single-table synthetic generation where data a…
When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation
Joshua Ward, Bochao Gu, Chi-Hua Wang +1
Large Language Models (LLMs) have recently demonstrated remarkable performance in generating high-quality tabular synthetic data. In practice, two primary approaches have emerged f…
MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs
Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin +8
Large Language Models exhibit strong reasoning and semantic understanding capabilities but often hallucinate in domains that require expert knowledge, among which fabrications, the…