27 papers
CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography
Tomer Erez, Moshe Kimhi, Chaim Baskin +1
Large chest radiography archives are difficult to search because most studies are paired only with free-text reports rather than structured clinical annotations. Vision-language mo…
DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception
Eli Goldenshluger, Barak Pinkovich, Chaim Baskin
DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-modal 3D detection for autonomous mobility by learning a sparse acquisiti…
FLASH: Flexible Learning of Adaptive Sampling from History in Temporal Graph Neural Networks
Or Feldman, Krishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb +2
Aggregating temporal signals from historic interactions is a key step in future link prediction on dynamic graphs. However, incorporating long histories is resource-intensive. Henc…
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
Dan Ben-Ami, Gabriele Serussi, Kobi Cohen +1
Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bo…
Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency
Itay Elam, Eliron Rahimi, Avi Mendelson +1
Pipeline parallelism is essential for training large neural networks, but existing schedules trade off throughput, memory, and optimization consistency. Synchronous pipelines prese…
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
Eliron Rahimi, Elad Hirshel, Rom Himelstein +3
Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding, competitive generation quality, and i…