most citedVision-guided and Mask-enhanced Adaptive Denoising for Prompt-based Image Editing

1 citations · 1 across the 9 of their papers we have counts for

collaborators

14 papers

cs.CL2025

Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Haokun Wen +3

Textual response generation is pivotal for multimodal \mbox{task-oriented} dialog systems, which aims to generate proper textual responses based on the multimodal context. While ex…

cs.CV2025

Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025

Qiaohui Chu, Haoyu Zhang, Yisen Feng +4

In this report, we present a novel three-stage framework developed for the Ego4D Long-Term Action Anticipation (LTA) task. Inspired by recent advances in foundation models, our met…

cs.LG2025

SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting

Haomiao Qiu, Miao Zhang, Ziyue Qiao +3

Continual Learning requires a model to learn multiple tasks in sequence while maintaining both stability:preserving knowledge from previously learned tasks, and plasticity:effectiv…

cs.CV2025

OSGNet @ Ego4D Episodic Memory Challenge 2025

Yisen Feng, Haoyu Zhang, Qiaohui Chu +4

In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise…

cs.LG2025

Boost Post-Training Quantization via Null Space Optimization for Large Language Models

Jiaqi Zhao, Miao Zhang, Deng Xiang +3

Existing post-training quantization methods for large language models (LLMs) offer remarkable success. However, the increasingly marginal performance gains suggest that existing qu…

cs.CV2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

Haoyu Zhang, Meng Liu, Zaijing Li +4

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied in…