27 papers
CAPRI: Contract-Aware Proof Repair for Isabelle
Jim Woodcock, Gabriel Leite, Augusto Sampaio +1
We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an LLM change…
MobileWan: Closing the Quality Gap for Mobile Video Diffusion
Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv +9
Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coheren…
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
Hyojin Park, Yi Li, Janghoon Cho +8
Despite decades of work, surveillance still struggles in searching and reasoning about specific targets across long, multi-camera videos. Existing methods - tracking, retrieval, an…
Rethinking RAG in Long Videos: What to Retrieve and How to Use It?
Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim +5
Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-relevant chunks across multiple modalities and temporal granularit…
Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs
Sudhanshu Agrawal, Risheek Garrepalli, Raghavv Goel +3
Diffusion LLMs (dLLMs) have recently emerged as a powerful alternative to autoregressive LLMs (AR-LLMs) with the potential to operate at significantly higher token-generation rates…
MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
Rajeev Yasarla, Deepti Hegde, Hsin-Pai Cheng +9
Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to being trained under traditional im…