24 citations · 26 across the 7 of their papers we have counts for
7 papers · 1 filter
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
Huijie Liu, Bingcan Wang, Jie Hu +2
Dish images play a crucial role in the digital era, with the demand for culturally distinctive dish images continuously increasing due to the digitization of the food industry and…
Faster Multi-GPU Training with PPLL: A Pipeline Parallelism Framework Leveraging Local Learning
Xiuyuan Guo, Chengqi Xu, Guinan Guo +6
Currently, training large-scale deep learning models is typically achieved through parallel training across multiple GPUs. However, due to the inherent communication overhead and s…
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Jiajun Liu, Yibing Wang, Hanghang Ma +6
Rapid advancements have been made in extending Large Language Models (LLMs) to Large Multi-modal Models (LMMs). However, extending input modality of LLMs to video data remains a ch…
Fine-gained Zero-shot Video Sampling
Dengsheng Chen, Jie Hu, Xiaoming Wei +1
Incorporating a temporal dimension into pretrained image diffusion models for video generation is a prevalent approach. However, this method is computationally demanding and necess…
EfficientRep:An Efficient Repvgg-style ConvNets with Hardware-aware Neural Network Design
Kaiheng Weng, Xiangxiang Chu, Xiaoming Xu +2
We present a hardware-efficient architecture of convolutional neural network, which has a repvgg-like architecture. Flops or parameters are traditional metrics to evaluate the effi…
PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative Grounding
Zihan Ding, Zi-han Ding, Tianrui Hui +4
Panoptic Narrative Grounding (PNG) is an emerging task whose goal is to segment visual objects of things and stuff categories described by dense narrative captions of a still image…