From the 1 of 5 linked papers with an AI index.
1 citations · 1 across the 1 of their papers we have counts for
5 papers
Decoupled Alignment for Robust Plug-and-Play Adaptation
Haozheng Luo, Jiahao Yu, Wenxin Zhang +9
The paper proposes a training-free, plug-and-play method that uses knowledge distillation and model fusion to correct misaligned (shadow-aligned) large language models, improving s…
Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
Jihyun Janice Ahn, Ryo Kamoi, Berk Atil +34
LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify…
GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models
Haozheng Luo, Chenghao Qiu, Yimin Wang +9
We propose the first unified adversarial attack benchmark for Genomic Foundation Models (GFMs), named GenoArmory. Unlike existing GFM benchmarks, GenoArmory offers the first compre…
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
Jiahao Yu, Haozheng Luo, Jerry Yao-Chieh Hu +3
Recent advances in Large Language Models (LLMs) have led to impressive alignment where models learn to distinguish harmful from harmless queries through supervised finetuning (SFT)…
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
Haozheng Luo, Chenghao Qiu, Maojiang Su +5
To address the challenge of scarce computational resources in genomic modeling, we introduce GERM, a genomic foundation model with strong compression performance and fast adaptabil…