4 papers
MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation
Debashish Chakraborty, Dengjia Zhang, Jialiang Jin +7
Retrieval-augmented generation from videos requires systems to retrieve relevant audiovisual evidence from large corpora and synthesize it into coherent, attributed text. Current a…
Captain Safari: A World Engine with Pose-Aligned 3D Memory
Yu-Cheng Chou, Xingrui Wang, Yitong Li +5
World engines aim to synthesize long, 3D-consistent videos that support interactive exploration of a scene under user-controlled camera motion. However, existing systems struggle u…
Parameter Competition Balancing for Model Merging
Guodong Du, Junlin Lee, Jing Li +8
While fine-tuning pretrained models has become common practice, these models often underperform outside their specific domains. Recently developed model merging techniques enable t…
Knowledge Fusion By Evolving Weights of Language Models
Guodong Du, Jing Li, Hanting Liu +5
Fine-tuning pre-trained language models, particularly large language models, demands extensive computing resources and can result in varying performance outcomes across different d…