collaborators

5 papers

cs.LG2026

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

Yunan Wang, Minghui Song, Zihan Zhang +6

Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL fra…

cs.CL2024

Context-DPO: Aligning Language Models for Context-Faithfulness

Baolong Bi, Shaohan Huang, Yiwei Wang +11

Reliable responses from large language models (LLMs) require adherence to user instructions and retrieved information. While alignment techniques help LLMs align with human intenti…

cs.CL2024

E5-V: Universal Embeddings with Multimodal Large Language Models

Ting Jiang, Minghui Song, Zihan Zhang +6

Multimodal large language models (MLLMs) have shown promising advancements in general visual and language understanding. However, the representation of multimodal information using…

cs.IR2024

ASI++: Towards Distributionally Balanced End-to-End Generative Retrieval

Yuxuan Liu, Tianchi Yang, Zihan Zhang +5

Generative retrieval, a promising new paradigm in information retrieval, employs a seq2seq model to encode document features into parameters and decode relevant document identifier…

cs.CL2024

MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning

Ting Jiang, Shaohan Huang, Shengyue Luo +8

Low-rank adaptation is a popular parameter-efficient fine-tuning method for large language models. In this paper, we analyze the impact of low-rank updating, as implemented in LoRA…