activity
20212025
most citedUsing Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks

43 citations · 43 across the 2 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

Length-Controlled Margin-Based Preference Optimization without Reference Model

Gengxu Li, Tingyu Xia, Yi Chang +1

Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF), designed to improve training simp…

cs.CL2025

A Survey of RWKV

Zhiyuan Li, Tingyu Xia, Yi Chang +1

The Receptance Weighted Key Value (RWKV) model offers a novel alternative to the Transformer architecture, merging the benefits of recurrent and attention-based systems. Unlike con…

cs.CL20241 cited

Large Language Model Evaluation via Matrix Nuclear-Norm

Yahan Li, Tingyu Xia, Yi Chang +1

As large language models (LLMs) continue to evolve, efficient evaluation metrics are vital for assessing their ability to compress information and reduce redundancy. While traditio…

cs.CL20242 cited

Rethinking Data Selection at Scale: Random Selection is Almost All You Need

Tingyu Xia, Bowen Yu, Kai Dang +5

Supervised fine-tuning (SFT) is crucial for aligning Large Language Models (LLMs) with human instructions. The primary goal during SFT is to select a small yet representative subse…

cs.CL2024

Language Models can Evaluate Themselves via Probability Discrepancy

Tingyu Xia, Bowen Yu, Yuan Wu +2

In this paper, we initiate our discussion by demonstrating how Large Language Models (LLMs), when tasked with responding to queries, display a more even probability distribution in…

cs.CL202143 cited

Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks

Tingyu Xia, Yue Wang, Yuan Tian +1

We study the problem of incorporating prior knowledge into a deep Transformer-based model,i.e.,Bidirectional Encoder Representations from Transformers (BERT), to enhance its perfor…