7 papers · 1 filter
Align-RAG: Alignment Is All You Need for TSFM In-Context Learning
Mohammad Asadi, Soheil Hor, Bardiya Akhbari +6
Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusi…
Ordered Action Tokens for Visuomotor Policy Learning
Chaoqi Liu, Yue Zhao, Haonan Chen +4
Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on…
A vision foundation model for single-cell biology via spatial gene cartography
Ridvan Yesiloglu, Sakib Mostafa, James Zou +5
Most single-cell foundation models are adapted from language models, representing each cell as a sequence of gene tokens. This discards the relationships among genes and often the…
HumanScore: Benchmarking Human Motions in Generated Videos
Yusu Fang, Tiange Xiang, Tian Tan +4
Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method syste…
Neuro-Symbolic Decoding of Neural Activity
Yanchen Wang, Joy Hsu, Ehsan Adeli +1
We propose NEURONA, a neuro-symbolic framework for fMRI decoding and concept grounding in neural activity. Leveraging image- and video-based fMRI question-answering datasets, NEURO…
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
Alan Baade, Eric Ryan Chan, Kyle Sargent +4
Latent diffusion models excel at generating high-quality images but lose the benefits of end-to-end modeling. They discard information during image encoding, require a separately t…