collaborators

8 papers

cs.CR2025

The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search

Rongzhe Wei, Peizhi Niu, Xinjie Shen +7

Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Existing approaches overwhelmingly operate within the p…

cs.LG2025

Forecasting Fails: Unveiling Evasion Attacks in Weather Prediction Models

Huzaifa Arif, Pin-Yu Chen, Alex Gittens +2

With the increasing reliance on AI models for weather forecasting, it is imperative to evaluate their vulnerability to adversarial perturbations. This work introduces Weather Adapt…

cs.CR2025

Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization

Xurui Li, Kaisong Song, Rui Zhu +2

Large Language Models (LLMs) have developed rapidly in web services, delivering unprecedented capabilities while amplifying societal risks. Existing works tend to focus on either i…

cs.CL2025

ICX360: In-Context eXplainability 360 Toolkit

Dennis Wei, Ronny Luss, Xiaomeng Hu +6

Large Language Models (LLMs) have become ubiquitous in everyday life and are entering higher-stakes applications ranging from summarizing meeting transcripts to answering doctors'…

cs.LG2025

Enhancing Sentiment Classification with Machine Learning and Combinatorial Fusion

Sean Patten, Pin-Yu Chen, Christina Schweikert +1

This paper presents a novel approach to sentiment classification using the application of Combinatorial Fusion Analysis (CFA) to integrate an ensemble of diverse machine learning m…

cs.LG2025

CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning

Yung-Chen Tang, Pin-Yu Chen, Andrea Cavallaro

Allocating more computation during inference time (test-time scaling) improves language model performance, especially for reasoning tasks. However, popular methods like Best-of-