4 papers
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
Dong Yan, Jian Liang, Dapeng Hu +4
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluati…
What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
Dong Yan, Jian Liang, Yanbo Wang +3
Test-Time Reinforcement Learning (TTRL) enables Large Language Models (LLMs) to enhance reasoning capabilities on unlabeled test streams by deriving pseudo-rewards from majority vo…
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
Dong Yan, Jian Liang, Ran He +1
Recent studies have shown that large language models (LLMs) can infer private user attributes (e.g., age, location, gender) from user-generated text shared online, enabling rapid a…
Mission Impossible: Feedback-Guided Dynamic Interactive Planning for Improving Reasoning on LLMs
Dong Yan, Gaochen Wu, Bowen Zhou
Recent advancements in language agents have led to significant improvements in multi-hop reasoning tasks. However, existing approaches often struggle with handling open-domain prob…