2 citations · 3 across the 4 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents
Hongxi Yan, Ziyue Huang, Shichao Fan +1
Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions. Existing critic-free methods propag…
cs.AI2026
EntroCut: Entropy-Guided Adaptive Truncation for Efficient Chain-of-Thought Reasoning in Small-scale Large Reasoning Models
Hongxi Yan, Qingjie Liu, Yunhong Wang
Large Reasoning Models (LRMs) excel at complex reasoning tasks through extended chain-of-thought generation, but their reliance on lengthy intermediate steps incurs substantial com…