activity
20242026
most citedCite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models

1 citations · 1 across the 1 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

Identifying and Analyzing Performance-Critical Tokens in Large Language Models

Yu Bai, Heyan Huang, Cesare Spinoso-Di Piano +4

In-context learning (ICL) has emerged as an effective solution for few-shot learning with large language models (LLMs). However, how LLMs leverage demonstrations to specify a task…

cs.CL2025

Real-time Factuality Assessment from Adversarial Feedback

Sanxing Chen, Yukun Huang, Bhuwan Dhingra

We show that existing evaluations for assessing the factuality of news from conventional sources, such as claims on fact-checking websites, result in high accuracies over time for…

cs.CL2025

To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts

Yukun Huang, Sanxing Chen, Hongyi Cai +1

Large Language Models (LLMs) are often augmented with external contexts, such as those used in retrieval-augmented generation (RAG). However, these contexts can be inaccurate or in…

cs.CL2024

CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling

Yu Bai, Xiyuan Zou, Heyan Huang +4

Long sequence modeling has gained broad interest as large language models (LLMs) continue to advance. Recent research has identified that a large portion of hidden states within th…

cs.CL2024

Tailoring Vaccine Messaging with Common-Ground Opinions

Rickard Stureborg, Sanxing Chen, Ruoyu Xie +6

One way to personalize chatbot interactions is by establishing common ground with the intended reader. A domain where establishing mutual understanding could be particularly impact…

cs.CL2024

ChatShop: Interactive Information Seeking with Language Agents

Sanxing Chen, Sam Wiseman, Bhuwan Dhingra

The desire and ability to seek new information strategically are fundamental to human learning but often overlooked in current language agent evaluation. We analyze a popular web s…