activity
20232026
most citedDistract Large Language Models for Automatic Jailbreak Attack

1 citations · 3 across the 23 of their papers we have counts for

collaborators
Showing 2026Show all

10 papers · 1 filter

cs.CL2026

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

Peng Lai, He Zhu, Zhiwen Ruan +6

Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets a…

cs.CL2026

VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation

Yixia Li, Yaqing Shi, Zhiwen Ruan +6

Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high c…

cs.CL2026

Beyond Static Rules: Automated Discovery of Latent Vulnerabilities in Text-to-SQL

Hanqing Wang, Yongdong Chi, Jian Yang +4

While Large Language Models (LLMs) have achieved remarkable success in Text-to-SQL tasks, their deployment in real-world environments is hindered by latent reliability issues. Iden…

cs.CL2026

Bridging the Agent-World Gap: Text World Models for LLM-based Agents

Yixia Li, Hongru Wang, Peng Lai +13

Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet m…

cs.CL2026

GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models

Zhiwen Ruan, Yichao Du, Jianjie Zheng +6

A promising paradigm for adapting instruction-tuned language models is to learn task-specific updates on a pretrained base model and subsequently merge them into the instruction-tu…

cs.CL2026

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

Yutao Hou, Yihan Jiang, Yuhan Xie +5

Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal activities or unethical beha…