works on

From the 1 of 20 linked papers with an AI index.

activity
20242026
collaborators

20 papers

cs.CV2026

SceneBind: Binding What and Where Across Vision, Audio and Language

Mingfei Chen, Zijun Cui, Ruoke Zhang +2

SceneBind introduces an omni‑modal representation that jointly encodes what objects are and where they are in 3D space across vision, audio, and language, enabling cross‑modal scen…

cs.LG2026

The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning

James Hazelden, Laura Driscoll, Eli Shlizerman +1

In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal state variables. We define…

physics.optics2026

Advantages of Broadband Metalenses for Generalizable Image Classification

Yubo Zhang, Johannes Fröch, Jinlin Xiang +6

Optical neural networks (ONNs) are gaining increasing attention to accelerate machine learning tasks. In particular, static meta-optical encoders designed for task-specific pre-pro…

cs.CL2026

DiffuMask: Diffusion Language Model for Token-level Prompt Pruning

Caleb Zheng, Jyotika Singh, Fang Tu +6

In-Context Learning and Chain-of-Thought prompting improve reasoning in large language models (LLMs). These typically come at the cost of longer, more expensive prompts that may co…

cs.LG2026

RPNT: Robust Pre-trained Neural Transformer -- A Pathway for Generalized Motor Decoding

Hao Fang, Ryan A. Canfield, Tomohiro Ouchi +3

Brain motor decoding aims to interpret and translate neural activity into behaviors. Decoding models should generalize across variations, such as recordings from different brain si…

cs.GR2026

2ndMatch: Finetuning Pruned Diffusion Models via Second-Order Jacobian Matching

Caleb Zheng, Eli Shlizerman

Diffusion models achieve remarkable performance across diverse generative tasks in computer vision, but their high computational cost remains a major barrier to deployment. Model p…