papers

Publications (30)

cs.CL2025

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

Qixin Sun, Ziqin Wang, Hengyuan Zhao +6

Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating…

cs.RO2025

SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps

Haojun Xu, Jiaqi Xiang, Wu Wei +5

A typical human strategy for giving navigation guidance is to sketch route maps based on the environmental layout. Inspired by this, we introduce Sketch map-based visual Navigation…

cs.CV2022

Actor and Action Modular Network for Text-based Video Segmentation

Jianhua Yang, Yan Huang, Kai Niu +3

Text-based video segmentation aims to segment an actor in video sequences by specifying the actor and its performing action with a textual query. Previous methods fail to explicitl…

cs.CV2026

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

Pengyiang Liu, Zhongyue Shi, Hongye Hao +7

Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across m…

cs.CV2025

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation

Ruipu Wu, Yige Zhang, Jinyu Chen +5

Aerial Vision-and-Language Navigation (VLN) is an emerging task that enables Unmanned Aerial Vehicles (UAVs) to navigate outdoor environments using natural language instructions an…

cs.CV2025

Group Critical-token Policy Optimization for Autoregressive Image Generation

Guohui Zhang, Hu Yu, Xiaoxiao Ma +6

Recent studies have extended Reinforcement Learning with Verifiable Rewards (RLVR) to autoregressive (AR) visual generation and achieved promising progress. However, existing metho…

cs.RO2025

"Hi AirStar, Guide Me to the Badminton Court."

Ziqin Wang, Jinyu Chen, Xiangyi Zheng +3

Unmanned Aerial Vehicles, operating in environments with relatively few obstacles, offer high maneuverability and full three-dimensional mobility. This allows them to rapidly appro…

cs.CV2025

MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning

Weikang Shi, Aldrich Yu, Rongyao Fang +11

While Large Language Models (LLMs) have excelled in textual reasoning, they struggle with mathematical domains like geometry that intrinsically rely on visual aids. Existing approa…

cs.CV2024

GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance

Jingqiu Zhou, Lue Fan, Xuesong Chen +3

In this paper, we present GaussianPainter, the first method to paint a point cloud into 3D Gaussians given a reference image. GaussianPainter introduces an innovative feed-forward…

cs.RO2026

CR-Solver: GPU-Accelerated Kinematics Solver for Tendon-driven Continuum Robots

Heqing Yang, Yang Yi, Linqing Zhong +2

The paper introduces CR-Solver, a GPU‑accelerated optimization framework that solves inverse kinematics, path following, and trajectory planning for tendon‑driven continuum robots…

#continuum robots#tendon-driven actuation#kinematics solving#gpu acceleration
cs.CV2025

Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation

Hang Xu, Linjiang Huang, Feng Zhao

Test-time scaling (TTS) aims to achieve better results by increasing random sampling and evaluating samples based on rules and metrics. However, in text-to-image(T2I) diffusion mod…

cs.CV2025

InfoScale: Unleashing Training-free Variable-scaled Image Generation via Effective Utilization of Information

Guohui Zhang, Jiangtong Tan, Linjiang Huang +4

Diffusion models (DMs) have become dominant in visual generation but suffer performance drop when tested on resolutions that differ from the training scale, whether lower or higher…

cs.CV2025

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

Hongyu Li, Manyuan Zhang, Dian Zheng +11

Instruction-based image editing has emerged as a prominent research area, which, benefiting from image generation foundation models, have achieved high aesthetic quality, making in…

cs.CV2021

Foreground-Action Consistency Network for Weakly Supervised Temporal Action Localization

Linjiang Huang, Liang Wang, Hongsheng Li

As a challenging task of high-level video understanding, weakly supervised temporal action localization has been attracting increasing attention. With only video annotations, most…

cs.CV2022

Teach-DETR: Better Training DETR with Teachers

Linjiang Huang, Kaixin Lu, Guanglu Song +4

In this paper, we present a novel training scheme, namely Teach-DETR, to learn better DETR-based detectors from versatile teacher detectors. We show that the predicted boxes from t…

cs.CV2025

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Rongyao Fang, Aldrich Yu, Chengqi Duan +7

The advancement of open-source text-to-image (T2I) models has been hindered by the absence of large-scale, reasoning-focused datasets and comprehensive evaluation benchmarks, resul…

cs.CV2024

FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis

Linjiang Huang, Rongyao Fang, Aiping Zhang +4

In this study, we delve into the generation of high-resolution images from pre-trained diffusion models, addressing persistent challenges, such as repetitive patterns and structura…

cs.LG2025

AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert

Yuting Gao, Wang Lan, Hengyuan Zhao +3

Multimodal Mixture-of-Experts (MoE) models offer a promising path toward scalable and efficient large vision-language systems. However, existing approaches rely on rigid routing st…

cs.CV2022

Weakly Supervised Temporal Action Localization via Representative Snippet Knowledge Propagation

Linjiang Huang, Liang Wang, Hongsheng Li

Weakly supervised temporal action localization aims to localize temporal boundaries of actions and simultaneously identify their categories with only video-level category labels. M…

cs.CV2025

FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment

Hang Xu, Jie Huang, Linjiang Huang +3

Domain Adaptation(DA) for dense prediction tasks is an important topic, which enhances the dense prediction model's performance when tested on its unseen domain. Recently, with the…

cs.CV2025

SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving

Xuesong Chen, Linjiang Huang, Tao Ma +3

The integration of Vision-Language Models (VLMs) into autonomous driving systems has shown promise in addressing key challenges such as learning complexity, interpretability, and c…

cs.CV2025

FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal

Hang Xu, Linjiang Huang, Feng Zhao

Test-time scaling (TTS) has become a prevalent technique in image generation, significantly boosting output quality by expanding the number of parallel samples and filtering them u…

cs.CV2026

Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models

Siming Fu, Haojun Xu, Ruizhe He +9

Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfull…

cs.RO2026

EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots

Heqing Yang, Yang Yi, Liyao Wang +8

We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulation policies on real robots. Sitting between a policy server and…

cs.CV2026

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

Liyao Wang, Ruipu Wu, Haojun Xu +3

Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing…

cs.CV2025

FlexDrive: Toward Trajectory Flexibility in Driving Scene Reconstruction and Rendering

Jingqiu Zhou, Lue Fan, Linjiang Huang +4

Driving scene reconstruction and rendering have advanced significantly using the 3D Gaussian Splatting. However, most prior research has focused on the rendering quality along a pr…

cs.CV2025

GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Rongyao Fang, Chengqi Duan, Kun Wang +9

Current image generation and editing methods primarily process textual prompts as direct inputs without reasoning about visual composition and explicit operations. We present Gener…

cs.CV2023

Improving Weakly Supervised Temporal Action Localization by Bridging Train-Test Gap in Pseudo Labels

Jingqiu Zhou, Linjiang Huang, Liang Wang +2

The task of weakly supervised temporal action localization targets at generating temporal boundaries for actions of interest, meanwhile the action category should also be classifie…

cs.CV2026

GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Chengqi Duan, Rongyao Fang, Yuqing Wang +5

Visual generation models have made remarkable progress in creating realistic images from text prompts, yet struggle with complex prompts that specify multiple objects with precise…

cs.CV2024

FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction

Runze He, Kai Ma, Linjiang Huang +6

Introducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose F…