activity
20242026
collaborators

13 papers

cs.CV2026

Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training

Ziqi Gao, Weikai Huang, Jieyu Zhang +2

Recent advances in text-to-vision generation excel in visual fidelity but struggle with compositional generalization and semantic alignment. Existing datasets are noisy and weakly…

astro-ph.HE2026

Neuro-Parametric Spectral Classification of Black Hole and Neutron Star X-ray Binary Systems

Akash Garg, Aman Kumar, Ajit Kembhavi +5

We perform the classification of black hole and neutron star X-ray binary systems using deep neural networks applied to archival RXTE X-ray spectral data. We first construct two ne…

cs.CV2025

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

Arijit Ray, Jiafei Duan, Ellis Brown +9

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal lang…

cs.RO2025

Towards Embodiment Scaling Laws in Robot Locomotion

Bo Ai, Liu Dai, Nico Bohlinger +7

Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…

cs.CV2025

One Diffusion to Generate Them All

Duong H. Le, Tuan Pham, Sangho Lee +5

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables condit…

cs.RO2025

The One RING: a Robotic Indoor Navigation Generalist

Ainaz Eftekhar, Rose Hendrix, Luca Weihs +11

Modern robots vary significantly in shape, size, and sensor configurations used to perceive and interact with their environments. However, most navigation policies are embodiment-s…