2 papers
cs.CV2026
Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration
Qi Ming, Yuyang Wang, Mingjing Zhao +7
Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rel…
cs.GR2026
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
Chuanzhi Xu, Huiqi Liang, Bang Shi +7
Long video generation requires high-fidelity synthesis, coherent narrative structure, and user control over extended time spans. Existing text-to-video methods often rely on a sing…