From the 1 of 9 linked papers with an AI index.
9 papers
Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution
Zhengbo Jiao, Hongyu Xian, Qinglong Wang +5
The paper introduces Policy of Thoughts (PoT), a test‑time training framework that continuously updates a lightweight LoRA adapter using online policy optimization to improve large…
Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization
Zhenhua Xu, Dongsheng Chen, Jian Li +7
Building general-purpose role-playing agents that faithfully portray any character from a natural-language profile remains challenging. The dominant paradigm -- supervised fine-tun…
SRAF: Stealthy and Robust Adversarial Fingerprint for Copyright Verification of Large Language Models
Zhebo Wang, Zhenhua Xu, Maike Li +4
The protection of Intellectual Property (IP) for Large Language Models (LLMs) has become a critical concern as model theft and unauthorized commercialization escalate. While advers…
Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
Zhenhua Xu, Xubin Yue, Zhebo Wang +9
Copyright protection for large language models is of critical importance, given their substantial development costs, proprietary value, and potential for misuse. Existing surveys h…
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
Jin Li, Zhebo Wang, Tianliang Lu +3
Entropy-based inference methods have gained traction for improving the reliability of Large Language Models (LLMs). However, many existing approaches, such as entropy minimization…
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
Zhebo Wang, Xiaohu Mu, Zijie Zhou +4
Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, p…