13 papers
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
Ziqi Gao, Weikai Huang, Jieyu Zhang +2
Recent advances in text-to-vision generation excel in visual fidelity but struggle with compositional generalization and semantic alignment. Existing datasets are noisy and weakly…
Neuro-Parametric Spectral Classification of Black Hole and Neutron Star X-ray Binary Systems
Akash Garg, Aman Kumar, Ajit Kembhavi +5
We perform the classification of black hole and neutron star X-ray binary systems using deep neural networks applied to archival RXTE X-ray spectral data. We first construct two ne…
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
Arijit Ray, Jiafei Duan, Ellis Brown +9
Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal lang…
Towards Embodiment Scaling Laws in Robot Locomotion
Bo Ai, Liu Dai, Nico Bohlinger +7
Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…
One Diffusion to Generate Them All
Duong H. Le, Tuan Pham, Sangho Lee +5
We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables condit…
The One RING: a Robotic Indoor Navigation Generalist
Ainaz Eftekhar, Rose Hendrix, Luca Weihs +11
Modern robots vary significantly in shape, size, and sensor configurations used to perceive and interact with their environments. However, most navigation policies are embodiment-s…