5 papers
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
Xianghao Zang, Zijian Jiang, Jiarong Cheng +8
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Mod…
OV-DEIM: Real-time DETR-Style Open-Vocabulary Object Detection with GridSynthetic Augmentation
Leilei Wang, Longfei Liu, Xi Shen +4
Real-time open-vocabulary object detection (OVOD) is essential for practical deployment in dynamic environments, where models must recognize a large and evolving set of categories…
Flexora: Flexible Low Rank Adaptation for Large Language Models
Chenxing Wei, Yao Shu, Ying Tiffany He +1
Large Language Models (LLMs) are driving advancements in artificial intelligence by increasing the scale of model parameters, which has significantly enhanced generalization abilit…
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
Nianbo Zeng, Haowen Hou, Fei Richard Yu +2
Despite recent advances in retrieval-augmented generation (RAG) for video understanding, effectively understanding long-form video content remains underexplored due to the vast sca…
A Review of Human Emotion Synthesis Based on Generative Technology
Fei Ma, Yukan Li, Yifan Xie +8
Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the…