From the 1 of 4 linked papers with an AI index.
4 papers
Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution
Zhengbo Jiao, Hongyu Xian, Qinglong Wang +5
The paper introduces Policy of Thoughts (PoT), a test‑time training framework that continuously updates a lightweight LoRA adapter using online policy optimization to improve large…
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
Zhebo Wang, Xiaohu Mu, Zijie Zhou +4
Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, p…
Scalpel-SAM: A Semi-Supervised Paradigm for Adapting SAM to Infrared Small Object Detection
Zihan Liu, Xiangning Ren, Dezhang Kong +2
Infrared small object detection urgently requires semi-supervised paradigms due to the high cost of annotation. However, existing methods like SAM face significant challenges of do…
LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources
Haichao Ji, Zibo Wang, Cheng Pan +4
Large Language Models (LLMs) have shown great promise in automating data analytics tasks by interpreting natural language queries and generating multi-operation execution plans. Ho…