8 papers
Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients
Zhihao Zhu, Hongyi Tang, Yi Yang +1
Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are particularly acute in healthcare,…
Assessing and Mitigating Miscalibration in LLM-Based Social Science Measurement
Jinyuan Wang, Ningyuan Deng, Yi Yang
Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables that can enter standard empirical…
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
Ho Hung Lim, Yi Yang
Visual RAG has offered an alternative to traditional RAG. It treats documents as images and uses vision encoders to obtain vision patch tokens. However, hundreds of patch tokens pe…
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
Jingzhou Jiang, Yi Yang, Kar Yan Tam
Hidden states change substantially across the layers of modern language models, but most layer-wise analyses focus on one aspect of that change. We propose Layer-wise Representatio…
DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction
Hongyi Tang, Zhihao Zhu, Yi Yang
Vision-language models (VLMs) are trained on large-scale image-text corpora that may contain private, copyrighted, or otherwise sensitive data, motivating membership inference as a…
Robust Predictive Modeling Under Unseen Data Distribution Shifts: A Methodological Commentary
Hanyu Duan, Yi Yang, Ahmed Abbasi +1
Most research designing novel predictive models, or employing existing ones, assumes that training and testing data are independent and identically distributed. In practice, the da…