most citedTransformers Discover Molecular Structure Without Graph Priors

2 citations · 2 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CV2025

Pillar-0: A New Frontier for Radiology Foundation Models

Kumar Krishna Agrawal, Longchao Liu, Long Lian +11

Radiology plays an integral role in modern medicine, yet rising imaging volumes have far outpaced workforce growth. Foundation models offer a path toward assisting with the full sp…

cs.CV2025

GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation

Karim Elmaaroufi, Liheng Lai, Justin Svegliato +3

Vision Language Models (VLMs) achieve strong performance on many vision-language tasks but often struggle with spatial reasoning$\unicode{x2014}$a prerequisite for many application…

cs.LG20252 cited

Transformers Discover Molecular Structure Without Graph Priors

Tobias Kreiman, Yutong Bai, Fadi Atieh +3

Graph Neural Networks (GNNs) are the dominant architecture for molecular machine learning, particularly for molecular property prediction and machine learning interatomic potential…

cs.CV2025

Whole-Body Conditioned Egocentric Video Prediction

Yutong Bai, Danny Tran, Amir Bar +3

We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic po…

cs.LG2025

REOrdering Patches Improves Vision Models

Declan Kutscher, David M. Chan, Yutong Bai +2

Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (ra…

cs.CV2025

Vector Quantized Feature Fields for Fast 3D Semantic Lifting

George Tang, Aditya Agarwal, Weiqiao Han +2

We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding mult…