2 papers
cs.CV2026
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation
Haojie Zhang, Di Wu, Bingyan Liu +5
While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained…
cs.MM2024
An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture
Tiancheng Shi, Yuanchen Wei, John R. Kender
We demonstrate the efficiencies and explanatory abilities of extensions to the common tools of Autoencoders and LLM interpreters, in the novel context of comparing different cultur…