6 papers
Validation-Induced Shapley Shifts: How Validation Structure Distorts Data Valuation
Yinan Shen, Ziao Yang, Hongfu Liu
Shapley values are widely used to attribute value to training data based on their marginal contribution to performance on a validation set. Existing practice often assumes these va…
Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection
Yunzhe Zhang, Hongfu Liu, Pengyu Hong
Text-to-image diffusion models can synthesize high-quality images, yet the outcome is notoriously sensitive to the random seed: different initial seeds often yield large variations…
Revisit, Extend, and Enhance Hessian-Free Influence Functions
Ziao Yang, Han Yue, Jian Chen +1
Influence functions serve as crucial tools for assessing sample influence in model interpretation, subset training set selection, noisy label detection, and more. By employing the…
Recontextualizing Famous Quotes for Brand Slogan Generation
Ziao Yang, Zizhang Chen, Lei Zhang +1
Slogans are concise and memorable catchphrases that play a crucial role in advertising by conveying brand identity and shaping public perception. However, advertising fatigue reduc…
Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
Anshuman Chhabra, Bo Li, Jian Chen +2
A core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this t…
Layer-Aware Influence for Online Data Valuation Estimation
Ziao Yang, Longbo Huang, Hongfu Liu
Data-centric learning emphasizes curating high-quality training samples to boost performance rather than designing new architectures. A central problem is to estimate the influence…