From the 1 of 20 linked papers with an AI index.
20 papers
SceneBind: Binding What and Where Across Vision, Audio and Language
Mingfei Chen, Zijun Cui, Ruoke Zhang +2
SceneBind introduces an omni‑modal representation that jointly encodes what objects are and where they are in 3D space across vision, audio, and language, enabling cross‑modal scen…
The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning
James Hazelden, Laura Driscoll, Eli Shlizerman +1
In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal state variables. We define…
Advantages of Broadband Metalenses for Generalizable Image Classification
Yubo Zhang, Johannes Fröch, Jinlin Xiang +6
Optical neural networks (ONNs) are gaining increasing attention to accelerate machine learning tasks. In particular, static meta-optical encoders designed for task-specific pre-pro…
DiffuMask: Diffusion Language Model for Token-level Prompt Pruning
Caleb Zheng, Jyotika Singh, Fang Tu +6
In-Context Learning and Chain-of-Thought prompting improve reasoning in large language models (LLMs). These typically come at the cost of longer, more expensive prompts that may co…
RPNT: Robust Pre-trained Neural Transformer -- A Pathway for Generalized Motor Decoding
Hao Fang, Ryan A. Canfield, Tomohiro Ouchi +3
Brain motor decoding aims to interpret and translate neural activity into behaviors. Decoding models should generalize across variations, such as recordings from different brain si…
2ndMatch: Finetuning Pruned Diffusion Models via Second-Order Jacobian Matching
Caleb Zheng, Eli Shlizerman
Diffusion models achieve remarkable performance across diverse generative tasks in computer vision, but their high computational cost remains a major barrier to deployment. Model p…