Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
Himanshu Singh, Ziwei Xu, A. V. Subramanyam +1
Large Language Models (LLMs) are powerful text generators, yet they can produce toxic or harmful content even when given seemingly harmless prompts. This presents a serious safety…
cs.CL2025
Reasoning LLMs are Wandering Solution Explorers
Jiahao Lu, Ziwei Xu, Mohan Kankanhalli
Large Language Models (LLMs) have demonstrated impressive reasoning abilities through test-time computation (TTC) techniques such as chain-of-thought prompting and tree-based reaso…