works on

From the 1 of 5 linked papers with an AI index.

most citedDecoupled Alignment for Robust Plug-and-Play Adaptation

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL20261 cited

Decoupled Alignment for Robust Plug-and-Play Adaptation

Haozheng Luo, Jiahao Yu, Wenxin Zhang +9

The paper proposes a training-free, plug-and-play method that uses knowledge distillation and model fusion to correct misaligned (shadow-aligned) large language models, improving s…

cs.AI2026

Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning

Jihyun Janice Ahn, Ryo Kamoi, Berk Atil +34

LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify…

cs.LG2025

GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models

Haozheng Luo, Chenghao Qiu, Yimin Wang +9

We propose the first unified adversarial attack benchmark for Genomic Foundation Models (GFMs), named GenoArmory. Unlike existing GFM benchmarks, GenoArmory offers the first compre…

cs.AI2025

Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries

Jiahao Yu, Haozheng Luo, Jerry Yao-Chieh Hu +3

Recent advances in Large Language Models (LLMs) have led to impressive alignment where models learn to distinguish harmful from harmless queries through supervised finetuning (SFT)…

cs.LG2025

Fast and Low-Cost Genomic Foundation Models via Outlier Removal

Haozheng Luo, Chenghao Qiu, Maojiang Su +5

To address the challenge of scarce computational resources in genomic modeling, we introduce GERM, a genomic foundation model with strong compression performance and fast adaptabil…