3 papers
cs.AI2026
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
Lin Li, Jiawei Huang, Qihao Quan +7
In this paper, we propose the first VL gentic easoning framework for few-hot multimo…
cs.LG2026
Revisiting DAgger in the Era of LLM-Agents
Changhao Li, Rushi Qiang, Jiawei Huang +4
Long-horizon LM agents learn from multi-turn interaction, where a single early mistake can alter the subsequent state distribution and derail the whole trajectory. Existing recipes…
cs.LG2026
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
Jiawei Huang, Qingping Yang, Renjie Zheng +1
Despite recent progress in Large Language Model (LLM) Agents for Software Engineering (SWE) tasks, end-to-end fine-tuning typically relies on verifiable terminal rewards such as wh…