3 papers
cs.CV2026
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
Junfu Pu, Yuxin Chen, Teng Wang +1
Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detail…
cs.CL2026
MemEvolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
Zihao Cheng, Zeming Liu, Yingyu Shan +7
While large language model--powered agents can self-evolve by accumulating experience or by dynamically creating new assets (i.e., tools or expert agents), existing frameworks typi…
cs.CV2024
Story3D-Agent: Exploring 3D Storytelling Visualization with Large Language Models
Yuzhou Huang, Yiran Qin, Shunlin Lu +4
Traditional visual storytelling is complex, requiring specialized knowledge and substantial resources, yet often constrained by human creativity and creation precision. While Large…