4 papers
Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer
Tyler Ashoff, Jordan Rodu
Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly satu…
Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization
Tyler Ashoff, Jordan Rodu
Modern opaque AI models prize performance over interpretability, which makes testing difficult. However, formal statistical tests conducted on a model's embedding space can provide…
Synthetic Data, Information, and Prior Knowledge: Why Synthetic Data Augmentation to Boost Sample Doesn't Work for Statistical Inference
Reid Dale, Jordan Rodu, Mike Baiocchi
The use of synthetic data to deidentify data and to improve predictive models is well-attested to. The augmentation of datasets using synthetically generated data is an alluring pr…
Data Gluttony: Epistemic Risks, Dependent Testing and Data Reuse in Large Datasets
Reid Dale, Jordan Rodu, Maria E. Currie +1
Large-scale registries have collected vast amounts of data which has enabled investigators to efficiently conduct studies of observational data. Common practice is for investigator…