4 papers
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
Chenglin Li, Qianglong Chen, Zhi Li +2
Recent advancements in Large Video-Language Models (LVLMs) have led to promising results in multimodal video understanding. However, it remains unclear whether these models possess…
InstructAttribute: Fine-grained Object Attributes editing with Instruction
Xingxi Yin, Jingfeng Zhang, Yue Deng +3
Text-to-image (T2I) diffusion models are widely used in image editing due to their powerful generative capabilities. However, achieving fine-grained control over specific object at…
ColorEdit: Training-free Image-Guided Color editing with diffusion model
Xingxi Yin, Zhi Li, Jingfeng Zhang +2
Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to a…
Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search
Chenglin Li, Qianglong Chen, Zhi Li +5
Instruction tuning is a crucial technique for aligning language models with humans' actual goals in the real world. Extensive research has highlighted the quality of instruction da…