12 citations · 23 across the 7 of their papers we have counts for
8 papers
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Yicheng Xiao, Lin Song, Rui Yang +6
With the advancement of language models, unified multimodal understanding and generation have made significant strides, with model architectures evolving from separated components…
Generating 360° Video is What You Need For a 3D Scene
Zhaoyang Zhang, Yannick Hold-Geoffroy, Miloš Hašan +4
Generating 3D scenes is still a challenging task due to the lack of readily available scene data. Most existing methods only produce partial scenes and provide limited navigational…
A dual contrastive framework
Yuan Sun, Zhao Zhang, Jorge Ortiz
In current multimodal tasks, models typically freeze the encoder and decoder while adapting intermediate layers to task-specific goals, such as region captioning. Region-level visu…
Artistic Neural Style Transfer Algorithms with Activation Smoothing
Xiangtian Li, Han Cao, Zhaoyang Zhang +3
The works of Gatys et al. demonstrated the capability of Convolutional Neural Networks (CNNs) in creating artistic style images. This process of transferring content images in diff…
Mitigating Knowledge Conflicts in Language Model-Driven Question Answering
Han Cao, Zhaoyang Zhang, Xiangtian Li +3
In the context of knowledge-driven seq-to-seq generation tasks, such as document-based question answering and document summarization systems, two fundamental knowledge sources play…
Research on Key Technologies for Cross-Cloud Federated Training of Large Language Models
Haowei Yang, Mingxiu Sui, Shaobo Liu +3
With the rapid development of natural language processing technology, large language models have demonstrated exceptional performance in various application scenarios. However, tra…