papers

Publications (36)

cs.LG2024

B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory

Luca Zancato, Arjun Seshadri, Yonatan Dukler +6

We describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resou…

cs.LG2024

Function Space and Critical Points of Linear Convolutional Networks

Kathlén Kohn, Guido Montúfar, Vahid Shahverdi +1

We study the geometry of linear networks with one-dimensional convolutional layers. The function spaces of these networks can be identified with semi-algebraic families of polynomi…

math.AG2017

Multigraded Cayley-Chow forms

Brian Osserman, Matthew Trager

We introduce a theory of multigraded Cayley-Chow forms associated to subvarieties of products of projective spaces. Two new phenomena arise: first, the construction turns out to re…

cs.CL2025

PICASO: Permutation-Invariant Context Composition with State Space Models

Tian Yu Liu, Alessandro Achille, Matthew Trager +3

Providing Large Language Models with relevant contextual knowledge at inference time has been shown to greatly improve the quality of their generations. This is often achieved by p…

cs.CV2017

General models for rational cameras and the case of two-slit projections

Matthew Trager, Bernd Sturmfels, John Canny +2

The rational camera model recently introduced in [19] provides a general methodology for studying abstract nonlinear imaging systems and their multi-view geometry. This paper build…

cs.CG2018

Consistent sets of lines with no colorful incidence

Boris Bukh, Xavier Goaoc, Alfredo Hubard +1

We consider incidences among colored sets of lines in and examine whether the existence of certain concurrences between lines of colors force the existence of at…

cs.CV2025

Descriminative-Generative Custom Tokens for Vision-Language Models

Pramuditha Perera, Matthew Trager, Luca Zancato +2

This paper explores the possibility of learning custom tokens for representing new concepts in Vision-Language Models (VLMs). Our aim is to learn tokens that can be effective for b…

cs.LG2026

EvoMAS: Evolutionary Generation of Multi-Agent Systems

Yuntong Hu, Yuting Zhang, Matthew Trager +4

Large language model (LLM)-based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but designing effective MAS architectures…

cs.LG2025

Geometry and Optimization of Shallow Polynomial Networks

Yossi Arjevani, Joan Bruna, Joe Kileel +2

We study shallow neural networks with monomial activations and output dimension one. The function space for these models can be identified with a set of symmetric tensors with boun…

cs.LG2024

Linear Spaces of Meanings: Compositional Structures in Vision-Language Models

Matthew Trager, Pramuditha Perera, Luca Zancato +3

We investigate compositional structures in data embeddings from pre-trained vision-language models (VLMs). Traditionally, compositionality has been associated with algebraic operat…

cs.LG2023

À-la-carte Prompt Tuning (APT): Combining Distinct Data Via Composable Prompting

Benjamin Bowman, Alessandro Achille, Luca Zancato +4

We introduce À-la-carte Prompt Tuning (APT), a transformer-based scheme to tune prompts on distinct data so that they can be arbitrarily composed at inference time. The individual…

cs.LG2022

Geometry of Linear Convolutional Networks

Kathlén Kohn, Thomas Merkh, Guido Montúfar +1

We study the family of functions that are represented by a linear convolutional neural network (LCN). These functions form a semi-algebraic subset of the set of linear maps from in…

cs.LG2025

Algebra Unveils Deep Learning -- An Invitation to Neuroalgebraic Geometry

Giovanni Luca Marchetti, Vahid Shahverdi, Stefano Mereta +2

In this position paper, we promote the study of function spaces parameterized by machine learning models through the lens of algebraic geometry. To this end, we focus on algebraic…

cs.CV2024

Multi-Modal Hallucination Control by Visual Information Grounding

Alessandro Favero, Luca Zancato, Matthew Trager +5

Generative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phe…

cs.AI2025

Experience-Guided Adaptation of Inference-Time Reasoning Strategies

Adam Stein, Matthew Trager, Benjamin Bowman +4

Enabling agentic AI systems to adapt their problem-solving approaches based on post-training interactions remains a fundamental challenge. While systems that update and maintain a…

math.AG2017

Changing Views on Curves and Surfaces

Kathlén Kohn, Bernd Sturmfels, Matthew Trager

Visual events in computer vision are studied from the perspective of algebraic geometry. Given a sufficiently general curve or surface in 3-space, we consider the image or contour…

cs.CV2023

Prompt Algebra for Task Composition

Pramuditha Perera, Matthew Trager, Luca Zancato +2

We investigate whether prompts learned independently for different tasks can be later combined through prompt algebra to obtain a model that supports composition of tasks. We consi…

cs.CV2018

On the Solvability of Viewing Graphs

Matthew Trager, Brian Osserman, Jean Ponce

A set of fundamental matrices relating pairs of cameras in some configuration can be represented as edges of a "viewing graph". Whether or not these fundamental matrices are generi…

math.OC2023

Symmetry Breaking in Symmetric Tensor Decomposition

Yossi Arjevani, Joan Bruna, Michael Field +3

In this note, we consider the highly nonconvex optimization problem associated with computing the rank decomposition of symmetric tensors. We formulate the invariance properties of…

cs.CV2021

Neural Splines: Fitting 3D Surfaces with Infinitely-Wide Neural Networks

Francis Williams, Matthew Trager, Joan Bruna +1

We present Neural Splines, a technique for 3D surface reconstruction that is based on random feature kernels arising from infinitely-wide shallow ReLU networks. Our method achieves…

cs.AI2026

How LLMs Might Think

Joseph Gottlieb, Ethan Kemp, Matthew Trager

The paper examines whether large language models can be said to think, arguing that while they likely do not engage in rational thought, they may exhibit a form of arational, purel…

#large language models#philosophy of mind#rationality#associative cognition
cs.CV2023

Train/Test-Time Adaptation with Retrieval

Luca Zancato, Alessandro Achille, Tian Yu Liu +3

We introduce Train/Test-Time Adaptation with Retrieval (), a method to adapt models both at train and test time by means of a retrieval module and a searchable pool of…

cs.LG2024

Compositional Structures in Neural Embedding and Interaction Decompositions

Matthew Trager, Alessandro Achille, Pramuditha Perera +2

We describe a basic correspondence between linear algebraic structures within vector embeddings in artificial neural networks and conditional independence constraints on the probab…

cs.LG2024

The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation

Lawrence Stewart, Matthew Trager, Sujan Kumar Gonugondla +1

Speculative decoding aims to speed up autoregressive generation of a language model by verifying in parallel the tokens generated by a smaller draft model.In this work, we explore…

cs.AI2025

LATTS: Locally Adaptive Test-Time Scaling

Theo Uscidda, Matthew Trager, Michael Kleinman +3

One common strategy for improving the performance of Large Language Models (LLMs) on downstream tasks involves using a \emph{verifier model} to either select the best answer from a…

cs.LG2019

Gradient Dynamics of Shallow Univariate ReLU Networks

Francis Williams, Matthew Trager, Claudio Silva +3

We present a theoretical and empirical study of the gradient dynamics of overparameterized shallow ReLU networks with one-dimensional input, solving least-squares interpolation. We…

cs.CV2024

NeRF-Insert: 3D Local Editing with Multimodal Control Signals

Benet Oriol Sabat, Alessandro Achille, Matthew Trager +1

We propose NeRF-Insert, a NeRF editing framework that allows users to make high-quality local edits with a flexible level of control. Unlike previous work that relied on image-to-i…

cs.CL2023

Meaning Representations from Trajectories in Autoregressive Models

Tian Yu Liu, Matthew Trager, Alessandro Achille +3

We propose to extract meaning representations from autoregressive language models by considering the distribution of all possible trajectories extending an input text. This strateg…

cs.AI2025

e1: Learning Adaptive Control of Reasoning Effort

Michael Kleinman, Matthew Trager, Alessandro Achille +2

Increasing the thinking budget of AI models can significantly improve accuracy, but not all questions warrant the same amount of reasoning. Users may prefer to allocate different a…

cs.LG2019

On the Expressive Power of Deep Polynomial Neural Networks

Joe Kileel, Matthew Trager, Joan Bruna

We study deep neural networks with polynomial activations, particularly their expressive power. For a fixed architecture and activation degree, a polynomial neural network defines…

cs.CV2024

Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-Encoding

Alessandro Achille, Greg Ver Steeg, Tian Yu Liu +3

Quantifying the degree of similarity between images is a key copyright issue for image-based machine learning. In legal doctrine however, determining the degree of similarity betwe…

cs.LG2020

Pure and Spurious Critical Points: a Geometric Study of Linear Networks

Matthew Trager, Kathlén Kohn, Joan Bruna

The critical locus of the loss function of a neural network is determined by the geometry of the functional space and by the parameterization of this space by the network's weights…

cs.CV2023

Towards Visual Foundational Models of Physical Scenes

Chethan Parameshwara, Alessandro Achille, Matthew Trager +7

We describe a first step towards learning general-purpose visual representations of physical scenes using only image prediction as a training criterion. To do so, we first define "…

cs.CL2026

Learning When to Attend: Conditional Memory Access for Long-Context LLMs

Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4

Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued pretraining on long-context data can help but is…

cs.CL2025

Maximally-Informative Retrieval for State Space Model Generation

Evan Becker, Benjamin Bowman, Matthew Trager +4

Given a query and dataset, the optimal way of answering the query is to make use all the information available. Modern LLMs exhibit impressive ability to memorize training data, bu…

math.AG2016

Congruences and Concurrent Lines in Multi-View Geometry

Jean Ponce, Bernd Sturmfels, Matthew Trager

We present a new framework for multi-view geometry in computer vision. A camera is a mapping between and a line congruence. This model, which ignores image planes an…