papers

Publications (41)

cs.CV2026

UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations

Yaqi Zhao, Wang Lin, Zijian Zhang +5

Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic informa…

physics.acc-ph2013

Progress on the Construction of the 100 MeV / 100 kW Electron Linac for the NSC KIPT Neutron Source

Chi Yun-Long, Pei Shi-Lun, Pei Guo-Xi +28

IHEP, China is constructing a 100 MeV / 100 kW electron Linac for NSC KIPT, Ukraine. This linac will be used as the driver of a neutron source based on a subcritical assembly. In 2…

cs.CV2024

Instruction Tuning-free Visual Token Complement for Multimodal LLMs

Dongsheng Wang, Jiequan Cui, Miaoge Li +3

As the open community of large language models (LLMs) matures, multimodal LLMs (MLLMs) have promised an elegant bridge between vision and language. However, current research is inh…

cs.CV2023

MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition

Xize Cheng, Linjun Li, Tao Jin +7

Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and a…

cs.SC2013

Domain-of-Attraction Estimation for Uncertain Non-polynomial Systems

Min Wu, Zhengfeng Yang, Wang Lin

In this paper, we consider the problem of computing estimates of the domain-of-attraction for non-polynomial systems. A polynomial approximation technique, based on multivariate po…

physics.acc-ph2014

Impedance budget and instability estimation of the HLS-II storage ring

Zhang Qingkun, Wang Lin, Li Weimin +1

The upgrade project of Hefei Light Source storage ring is under way. In this paper, the wake fields of new designed vacuum chambers have been simulated by CST code, and then broadb…

cs.CV2026

Proact-VL: A Proactive VideoLLM for Real-Time AI Companions

Weicai Yan, Yuhong Dai, Qi Ran +6

Proactive and real-time interactive experiences are essential for human-like AI companions, yet face three key challenges: (1) achieving low-latency inference under continuous stre…

hep-ph2025

Semileptonic Decays of and from Light-Cone Sum Rules

Wang Lin, Xiao-En Huang, Shan Cheng +1

We investigate the semileptonic decays of charmed mesons to light vector mesons within the framework of light-cone sum rules. Our calculation is performed at leading order in QCD c…

cs.CV2025

Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision

Weicai Yan, Wang Lin, Zirun Guo +5

Prompt learning has demonstrated promising results in fine-tuning pre-trained multimodal models. However, the performance improvement is limited when applied to more complex and fi…

cs.CV2026

Imagine Before You Draw: Visual Prompt Engineering for Image Generation

Liyu Jia, Fengda Zhang, Jiachun Pan +7

Incorporating visual semantic representations as an intermediate step before image generation can reduce the modeling difficulty between text and images, thereby improving generati…

cs.CV2025

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Zhiqi Ge, Juncheng Li, Xinglei Pang +7

Digital agents are increasingly employed to automate tasks in interactive digital environments such as web pages, software applications, and operating systems. While text-based age…

cs.SC2013

Exact Safety Verification of Interval Hybrid Systems Based on Symbolic-Numeric Computation

Zhengfeng Yang, Min Wu, Wang Lin

In this paper, we address the problem of safety verification of interval hybrid systems in which the coefficients are intervals instead of explicit numbers. A hybrid symbolic-numer…

physics.optics2025

Nanoplasmonic Optical Fiber Sensing of SARS-CoV-2 Nucleocapsid Protein Using an Aptamer-DNA Tetrahedron Interface

Xu Pin, Cui Jingyu, Cheng Zhi +8

Optical fiber sensing carries a number of potential advantages for diagnostics and biomarker detection and monitoring, yet particular challenges persist in linking molecular recogn…

cs.CV2026

WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark

Wang Lin, Feng Wang, Majun Zhang +7

Recent advances in image editing models have demonstrated remarkable capabilities in executing explicit instructions, such as attribute manipulation, style transfer, and pose synth…

cs.CV2025

Low-rank Prompt Interaction for Continual Vision-Language Retrieval

Weicai Yan, Ye Wang, Wang Lin +3

Research on continual learning in multi-modal tasks has been receiving increasing attention. However, most existing work overlooks the explicit cross-modal and cross-task interacti…

cs.SE2012

Exact Safety Verification of Hybrid Systems Based on Bilinear SOS Representation

Zhengfeng Yang, Min Wu, Wang Lin

In this paper, we address the problem of safety verification of nonlinear hybrid systems. A hybrid symbolic-numeric method is presented to compute exact inequality invariants of hy…

cs.CL2026

ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment

Zhipeng Bian, Jieming Zhu, Qijiong Liu +6

Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, personalized cover image generat…

cs.CV2025

IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

Hanting Wang, Tao Jin, Wang Lin +4

Bridge models in image restoration construct a diffusion process from degraded to clear images. However, existing methods typically require training a bridge model from scratch for…

cs.RO2025

Online Controller Synthesis for Robot Collision Avoidance: A Case Study

Yuheng Fan, Wang Lin

The inherent uncertainty of dynamic environments poses significant challenges for modeling robot behavior, particularly in tasks such as collision avoidance. This paper presents an…

cs.CL2023

OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

Xize Cheng, Tao Jin, Linjun Li +3

Speech Recognition builds a bridge between the multimedia streaming (audio-only, visual-only or audio-visual) and the corresponding text transcription. However, when training the s…

cs.IR2024

EAGER: Two-Stream Generative Recommender with Behavior-Semantic Collaboration

Ye Wang, Jiahao Xun, Minjie Hong +8

Generative retrieval has recently emerged as a promising approach to sequential recommendation, framing candidate item retrieval as an autoregressive sequence generation problem. H…

cs.CL2025

Cognitive-Level Adaptive Generation via Capability-Aware Retrieval and Style Adaptation

Qingsong Wang, Tao Wu, Wang Lin +4

Large Language Models (LLMs) have demonstrated strong performance in open-ended generation tasks. However, they often struggle to adapt content to users with differing cognitive ca…

cs.LG2025

Bridging the Gap for Test-Time Multimodal Sentiment Analysis

Zirun Guo, Tao Jin, Wenlong Xu +2

Multimodal sentiment analysis (MSA) is an emerging research topic that aims to understand and recognize human sentiment or emotions through multiple modalities. However, in real-wo…

cs.CV2025

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Bohan Wang, Zhongqi Yue, Fengda Zhang +15

We completely discard the conventional spatial prior in image representation and introduce a novel discrete visual tokenizer: Self-consistency Tokenizer (Selftok). At its design co…

physics.acc-ph2013

Longitudinal Single Bunch Instability Study on BEPCII

Wang Dou, Li Yong, Duan Zhe +4

In order to study the single bunch longitudinal instability in BEPCII, experiments on the positron ring (BPR) for the bunch lengthening phenomenon were made. By analyzing the exper…

cs.CV2026

Text-Guided Multi-Scale Frequency Representation Adaptation

Weicai Yan, Xinhua Ma, Wang Lin +1

Parameter-efficient fine-tuning methods introduce a small number of training parameters, enabling pre-trained models to adapt rapidly to new data distributions. While these methods…

cs.SE2011

Exact Safety Verification of Hybrid Systems Using Sums-Of-Squares Representation

Wang Lin, Min Wu, Zhengfeng Yang +1

In this paper we discuss how to generate inductive invariants for safety verification of hybrid systems. A hybrid symbolic-numeric method is presented to compute inequality inducti…

cs.LG2024

AutoGeo: Automating Geometric Image Dataset Creation for Enhanced Geometry Understanding

Zihan Huang, Tao Wu, Wang Lin +3

With the rapid advancement of large language models, there has been a growing interest in their capabilities in mathematical reasoning. However, existing research has primarily foc…

cs.LG2025

Efficient Prompting for Continual Adaptation to Missing Modalities

Zirun Guo, Shulei Wang, Wang Lin +3

Missing modality issues are common in real-world applications, arising from factors such as equipment failures and privacy concerns. When fine-tuning pre-trained models on downstre…

cs.AI2025

Contrastive Cross-Course Knowledge Tracing via Concept Graph Guided Knowledge Transfer

Wenkang Han, Wang Lin, Liya Hu +6

Knowledge tracing (KT) aims to predict learners' future performance based on historical learning interactions. However, existing KT models predominantly focus on data from a single…

cs.LG2025

Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents

Tao Wu, Jingyuan Chen, Wang Lin +5

Large language models (LLMs) are revolutionizing education, with LLM-based agents playing a key role in simulating student behavior. A major challenge in student simulation is mode…

cs.CV2025

Towards Transformer-Based Aligned Generation with Self-Coherence Guidance

Shulei Wang, Wang Lin, Hai Huang +8

We introduce a novel, training-free approach for enhancing alignment in Transformer-based Text-Guided Diffusion Models (TGDMs). Existing TGDMs often struggle to generate semantical…

cs.CV2024

FlowDreamer: Exploring High Fidelity Text-to-3D Generation via Rectified Flow

Hangyu Li, Xiangxiang Chu, Dingyuan Shi +1

Recent advances in text-to-3D generation have made significant progress. In particular, with the pretrained diffusion models, existing methods predominantly use Score Distillation…

cs.CV2025

Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration

Wenkang Han, Wang Lin, Yiyun Zhou +4

Face Video Restoration (FVR) aims to recover high-quality face videos from degraded versions. Traditional methods struggle to preserve fine-grained, identity-specific features when…

cs.CL2026

Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning Reconstruction

Tao Wu, Jingyuan Chen, Wang Lin +6

Distractors-incorrect yet plausible answer choices in multiple-choice questions (MCQs)-are vital in educational assessments, as they help identify student misconceptions by present…

cs.CV2025

Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning

Wang Lin, Liyu Jia, Wentao Hu +6

Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapol…

cs.AI2025

From Noisy to Native: LLM-driven Graph Restoration for Test-Time Graph Domain Adaptation

Xiangwei Lv, JinLuan Yang, Wang Lin +2

Graph domain adaptation (GDA) has achieved great attention due to its effectiveness in addressing the domain shift between train and test data. A significant bottleneck in existing…

cs.CV2024

Non-confusing Generation of Customized Concepts in Diffusion Models

Wang Lin, Jingyuan Chen, Jiaxin Shi +8

We tackle the common challenge of inter-concept visual confusion in compositional concept generation using text-guided diffusion models (TGDMs). It becomes even more pronounced in…

physics.acc-ph2014

Simulation of beam gas coulomb scattering in HALS

Yu Lu-Xin, Gao Wei-Wei, Wang Lin +1

In conventional research on the beam gas coulomb scattering (BGCS), only the related beam lifetime using the analytical method is studied. In this paper, using the PIC-MCC method,…

cs.CV2024

Semantic Alignment for Multimodal Large Language Models

Tao Wu, Mengze Li, Jingyuan Chen +6

Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly…

cs.CV2025

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Kaihang Pan, Wang Lin, Zhongqi Yue +6

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each ta…