8 papers
Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers
Yutian Chen, Yuheng Qiu, Ruogu Li +4
We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a l…
World2Rules: A Neuro-Symbolic Framework for Learning World-Governing Safety Rules for Aviation
Haichuan Wang, Jay Patrikar, Sebastian Scherer
Many real-world safety-critical systems are governed by explicit rules that define unsafe world configurations and constrain agent interactions. In practice, these rules are comple…
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
Haoyang He, Jay Patrikar, Dong-Ki Kim +5
Recent advances in video world modeling have enabled large-scale generative models to simulate embodied environments with high visual fidelity, providing strong priors for predicti…
Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
Jason Jabbour, Dong-Ki Kim, Max Smith +6
Vision-Language-Action (VLA) models have advanced robotic capabilities but remain challenging to deploy on resource-limited hardware. Pruning has enabled efficient compression of l…
Amelia: A Large Dataset and Benchmark for Airport Surface Movement Forecasting
Ingrid Navarro, Pablo Ortega-Kral, Jay Patrikar +6
Demand for air travel is rising, straining existing aviation infrastructure. In the US, more than 90% of airport control towers are understaffed, falling short of FAA and union sta…
The Case for Negative Data: From Crash Reports to Counterfactuals for Reasonable Driving
Jay Patrikar, Apoorva Sharma, Sushant Veer +3
Learning-based autonomous driving systems are trained mostly on incident-free data, offering little guidance near safety-performance boundaries. Real crash reports contain precisel…