From the 1 of 6 linked papers with an AI index.
6 papers
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Bo-Wen Zhang, Junwei He, Wen Wang +5
The paper introduces CoRT, a method that uses counterfactual replay to assign token-level credit in rubric-guided reinforcement learning for language models, improving credit alloc…
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
Song-Lin Lv, Weiming Wu, Rui Zhu +2
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, to…
Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models
Rui Zhu, Song-Lin Lv, Zi-Kang Wang +1
Exploiting unlabeled data through semi-supervised learning (SSL) or leveraging pre-trained models via fine-tuning are two prevailing paradigms for addressing label-scarce scenarios…
Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models
Song-Lin Lv, Rui Zhu, Tong Wei +2
Semi-supervised learning (SSL) alleviates the cost of data labeling process by exploiting unlabeled data and has achieved promising results. Meanwhile, with the development of larg…
Shift-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
Song-Lin Lv, Yu-Yang Chen, Zhi Zhou +2
Vision-language models (VLMs), such as CLIP, adapt effectively to downstream tasks through prompt tuning, but fine-tuning can misalign predictive confidence and accuracy, particula…
BMIP: Bi-directional Modality Interaction Prompt Learning for VLM
Song-Lin Lv, Yu-Yang Chen, Zhi Zhou +2
Vision-language models (VLMs) have exhibited remarkable generalization capabilities, and prompt learning for VLMs has attracted great attention for the ability to adapt pre-trained…