11 papers · 1 filter
SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models
Enqiao Lu, Xingrui Yu, Yiwei Fu +7
Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models…
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang +5
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a rec…
Improving Generalization Robustness of Multimodal RLVR
Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng +11
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing th…
SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale
Tong Bai, Zhenglin Wan, Pengfei Zhou +3
As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend on, conflict with, specializ…
Agent-as-a-Router: Agentic Model Routing for Coding Tasks
Pengfei Zhou, Zhiwei Tang, Yixing Ma +8
Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct domains, yet none dominate all. Con…
Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents
Chubin Zhang, Zhenglin Wan, Xingrui Yu +5
Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would t…