6 papers
When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models
Sina Alemohammad, Denghui Zhang, Bolong Tang +5
As language models move from drafting prose to running literature-search agents with tool calls, fabricated references are becoming easier to catch and constrain. The harder failur…
Not All Synthetic Data Is Yours to Learn From
Sina Alemohammad, Li Chen, Richard G. Baraniuk +1
Can a language model improve from plain text sampled from itself, with no prompts, no teacher, no verifier, and no reward model? Yes, but only when the synthetic corpus is compatib…
Minimizing Collateral Damage in Activation Steering
Tam Nguyen, Tu Anh Nguyen, Sina Alemohammad +1
Activation steering is a method for controlling Large Language Model (LLM) behavior by intervening in its internal representations to increase the alignment with a specific target…
Neon: Negative Extrapolation From Self-Training Improves Image Generation
Sina Alemohammad, Zhangyang Wang, Richard G. Baraniuk
Scaling generative AI models is bottlenecked by the scarcity of high-quality training data. The ease of synthesizing from a generative model suggests using (unverified) synthetic d…
WaLRUS: Wavelets for Long-range Representation Using SSMs
Hossein Babaei, Mel White, Sina Alemohammad +1
State-Space Models (SSMs) have proven to be powerful tools for modeling long-range dependencies in sequential data. While the recent method known as HiPPO has demonstrated strong p…
SaFARi: State-Space Models for Frame-Agnostic Representation
Hossein Babaei, Mel White, Sina Alemohammad +1
State-Space Models (SSMs) have re-emerged as a powerful tool for online function approximation, and as the backbone of machine learning models for long-range dependent data. Howeve…