3 papers
cs.LG2024
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
Youngseog Chung, Dhruv Malik, Jeff Schneider +2
The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small e…
cs.LG2024
What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
Sang Keun Choe, Hwijeen Ahn, Juhan Bae +11
Large language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attr…
physics.plasm-ph2024
Full Shot Predictions for the DIII-D Tokamak via Deep Recurrent Networks
Ian Char, Youngseog Chung, Joseph Abbate +2
Although tokamaks are one of the most promising devices for realizing nuclear fusion as an energy source, there are still key obstacles when it comes to understanding the dynamics…