papers

Publications (14)

cs.LG2025

When Bad Data Leads to Good Models

Kenneth Li, Yida Chen, Fernanda Viégas +1

In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- an…

cs.HC2024

An AI-Resilient Text Rendering Technique for Reading and Skimming Documents

Ziwei Gu, Ian Arawjo, Kenneth Li +2

Readers find text difficult to consume for many reasons. Summarization can address some of these difficulties, but introduce others, such as omitting, misrepresenting, or hallucina…

cs.CL2024

Designing a Dashboard for Transparency and Control of Conversational AI

Yida Chen, Aoyu Wu, Trevor DePodesta +9

Conversational LLMs function as black box systems, leaving users guessing about why they see the output they do. This lack of transparency is potentially problematic, especially gi…

cs.LG2024

Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task

Kenneth Li, Aspen K. Hopkins, David Bau +3

Language models show a surprising range of capabilities, but the source of their apparent competence is unclear. Do these networks just memorize a collection of surface statistics,…

cs.CV2021

Towards Tokenized Human Dynamics Representation

Kenneth Li, Xiao Sun, Zhirong Wu +2

For human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking. However, methods for u…

cs.CV2021

Do Time Constraints Re-Prioritize Attention to Shapes During Visual Photo Inspection?

Yiyuan Yang, Kenneth Li, Fernanda Eliott +1

People's visual experiences of the world are easy to carve up and examine along natural language boundaries, e.g., by category labels, attribute labels, etc. However, it is more di…