From the 1 of 36 linked papers with an AI index.
2 citations · 2 across the 11 of their papers we have counts for
7 papers · 1 filter
Do LLMs Know Their Vulnerable Scenarios?
Ziheng Peng, Huiqi Deng, Haoran Jing +5
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teami…
Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions
Yue Xu, Qian Chen, Zizhan Ma +5
Large language models have enabled agentic systems that reason, plan, and interact with tools and environments to accomplish complex tasks. As these agents operate over extended in…
Efficient and Stable Reinforcement Learning for Diffusion Language Models
Jiawei Liu, Xiting Wang, Yuanyuan Zhong +2
Reinforcement Learning (RL) is crucial for unlocking the complex reasoning capabilities of Diffusion-based Large Language Models (dLLMs). However, applying RL to dLLMs faces unique…
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
Qingyu Yin, Chak Tou Leong, Linyi Yang +7
Large reasoning models (LRMs) with multi-step reasoning capabilities have shown remarkable problem-solving abilities, yet they exhibit concerning safety vulnerabilities that remain…
Entropy-based Exploration Conduction for Multi-step Reasoning
Jinghan Zhang, Xiting Wang, Fengran Mo +3
Multi-step processes via large language models (LLMs) have proven effective for solving complex reasoning tasks. However, the depth of exploration of the reasoning procedure can si…
Large Language Models show both individual and collective creativity comparable to humans
Luning Sun, Yuzhuo Yuan, Yuan Yao +6
Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to hum…