5 papers
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving
Shihao Ji, HongXi Li, Zihui Song +1
Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically…
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
Hongxi Li, Tong Wang, Chengjing Wu +6
Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background in…
Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum
Zihui Song, Shihao Ji, Hongxi Li +2
We investigate the hypothesis that real-data scaling laws are governed by progressive coverage of a latent predictive contribution spectrum rather than by token-frequency tails alo…
Video Summarization using Denoising Diffusion Probabilistic Model
Zirui Shang, Yubo Zhu, Hongxi Li +2
Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative…
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
Yayun Qi, Hongxi Li, Yiqi Song +2
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelli…