activity
20242026
most citedOSPC: Detecting Harmful Memes with Large Language Model as a Catalyst

3 citations · 12 across the 19 of their papers we have counts for

collaborators

20 papers

cs.LG2026

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

Zhiwei Zhang, Zechen Sun, Fei Zhao +6

On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supe…

cs.IR2026

World Model-Guided Reinforcement Learning via Counterfactual User Engagement Simulation

Ang Li, Xin Xu, Bin Liang +4

Reinforcement learning for user-centric agents is limited by the cost, latency, and risk of collecting online feedback, as well as by the lack of counterfactual comparisons under t…

cs.CL2026

Not All or None: Dynamic Construction of Target-aware Memory Graph for Conversational Stance Detection

Yifan Xiang, Bin Liang, Yuqi Huang +2

Stance detection is crucial for understanding the underlying attitude of an expression towards a target. Conversational stance detection is a more challenging stance detection task…

cs.CV2026

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

Bojia Zi, Xiaoyan Yang, Yu Zhou +7

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, t…

cs.CL2026

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

Kang Peng, Zhiwei Zhang, Yichen Zhang +7

Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following proc…

cs.AI2026

Expectation Alignment of Language Models for Real-World User Expectations

Miaomiao Li, Yang Wang, Bin Liang +3

Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing…