3 papers
cs.LG2026
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
cs.CV2026
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
Guanfeng Tang, Hongbo Zhao, Ziwei Long +5
Inspired by the human visual system, which operates on two parallel yet interactive streams for contextual and spatial understanding, this article presents Two Interactive Streams…
cs.CV2024
Can video generation replace cinematographers? Research on the cinematic language of generated video
Xiaozhe Li, Kai WU, Siyi Yang +12
Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing…