7 papers
CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark
Junzhao Zhang, Hsiu-Yuan Huang, Chenming Tang +2
Multimodal sarcasm detection has recently garnered significant attention. However, existing benchmarks suffer from coarse-grained annotations and limited cultural coverage, which h…
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
Kun Liang, Clive Bai, Xin Xu +5
Recent Large Reasoning Models (LRMs) achieve strong performance by leveraging long-form Chain-of-Thought (CoT) reasoning, but uniformly applying overlong reasoning at inference tim…
Think Outside the Policy: In-Context Steered Policy Optimization
Hsiu-Yuan Huang, Chenming Tang, Weijie Liu +3
Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods, such as Group Relative Policy Optimization (GRPO), have achieved remarkable progress in improving the reason…
Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
Chenming Tang, Hsiu-Yuan Huang, Weijie Liu +3
Reinforcement learning with verifiable rewards (RLVR) has significantly boosted the reasoning capability of language models (LMs). However, existing RLVR approaches train LMs based…
Aligning Language Models with Real-time Knowledge Editing
Chenming Tang, Yutong Yang, Kexue Wang +1
Knowledge editing aims to modify outdated knowledge in language models efficiently while retaining their original capabilities. Mainstream datasets for knowledge editing are predom…
Lost in the Passage: Passage-level In-context Learning Does Not Necessarily Need a "Passage"
Hao Sun, Chenming Tang, Gengyang Li +1
By simply incorporating demonstrations into the context, in-context learning (ICL) enables large language models (LLMs) to yield awesome performance on many tasks. In this study, w…