papers

Publications (42)

cs.CV2021

TransZero: Attribute-guided Transformer for Zero-Shot Learning

Shiming Chen, Ziming Hong, Yang Liu +6

Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descripti…

eess.SP2021

Adversarial Energy Disaggregation for Non-intrusive Load Monitoring

Zhekai Du, Jingjing Li, Lei Zhu +2

Energy disaggregation, also known as non-intrusive load monitoring (NILM), challenges the problem of separating the whole-home electricity usage into appliance-specific individual…

cond-mat.mtrl-sci2024

One Pot Synthesis of Cubic Gauche Polymeric Nitrogen

Runteng Chen, Jun Zhang, Zelong Wang +5

The long sought cubic gauche polymeric nitrogen (cg-N) consisting of N-N single bonds has been synthesized by a simple route using sodium azide as a precursor at ambient conditions…

cond-mat.supr-con2023

Superconductivity above 30 K achieved in dense scandium

Xin He, Changling Zhang, Zhiwen Li +15

Superconductivity is one of most intriguing quantum phenomena, and the quest for elemental superconductors with high critical temperature (Tc) is of great scientific significance d…

cs.LG2025

LoCA: Location-Aware Cosine Adaptation for Parameter-Efficient Fine-Tuning

Zhekai Du, Yinjie Min, Jingjing Li +5

Low-rank adaptation (LoRA) has become a prevalent method for adapting pre-trained large language models to downstream tasks. However, the simple low-rank decomposition form may con…

cs.CV2017

Learning Correspondence Structures for Person Re-identification

Weiyao Lin, Yang Shen, Junchi Yan +4

This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-ba…

cs.CV2026

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

Xiuyuan Zhu, Ke Lu, Hao Wu +4

Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given…

cond-mat.mtrl-sci2024

A facile route to synthesize cubic gauche polymeric nitrogen

Runteng Chen, Jun Zhang, Zelong Wang +8

In this work, the long-sought cg-N with N-N single bond has been synthesized for the first time by a thermal-driven-only chemical route at ambient conditions. The successful synthe…

cs.CV2022

Markerless Body Motion Capturing for 3D Character Animation based on Multi-view Cameras

Jinbao Wang, Ke Lu, Jian Xue

This paper proposes a novel application system for the generation of three-dimensional (3D) character animation driven by markerless human body motion capturing. The entire pipelin…

cs.CV2021

Domain Adaptive Semantic Segmentation without Source Data

Fuming You, Jingjing Li, Lei Zhu +3

Domain adaptive semantic segmentation is recognized as a promising technique to alleviate the domain shift between the labeled source domain and the unlabeled target domain in many…

cs.CV2019

From Zero-Shot Learning to Cold-Start Recommendation

Jingjing Li, Mengmeng Jing, Ke Lu +3

Zero-shot learning (ZSL) and cold-start recommendation (CSR) are two challenging problems in computer vision and recommender system, respectively. In general, they are independentl…

cs.AI2024

Domain-Agnostic Mutual Prompting for Unsupervised Domain Adaptation

Zhekai Du, Xinyao Li, Fengling Li +3

Conventional Unsupervised Domain Adaptation (UDA) strives to minimize distribution discrepancy between domains, which neglects to harness rich semantics from data and struggles to…

cs.CV2019

Alleviating Feature Confusion for Generative Zero-shot Learning

Jingjing Li, Mengmeng Jing, Ke Lu +3

Lately, generative adversarial networks (GANs) have been successfully applied to zero-shot learning (ZSL) and achieved state-of-the-art performance. By synthesizing virtual unseen…

cs.CV2024

ExpLLM: Towards Chain of Thought for Facial Expression Recognition

Xing Lan, Jian Xue, Ji Qi +3

Facial expression recognition (FER) is a critical task in multimedia with significant implications across various domains. However, analyzing the causes of facial expressions is es…

cs.CV2019

Leveraging the Invariant Side of Generative Zero-Shot Learning

Jingjing Li, Mengmeng Jin, Ke Lu +3

Conventional zero-shot learning (ZSL) methods generally learn an embedding, e.g., visual-semantic mapping, to handle the unseen visual samples via an indirect manner. In this paper…

cs.CV2026

Decoupled Hierarchical Distillation for Multimodal Emotion Recognition

Yong Li, Yuanzhi Wang, Yi Ding +3

Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches…

cs.CV2026

Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA

Zibo Xu, Qiang Li, Ke Lu +3

Medical Visual Question Answering (MedVQA) aims to generate clinically reliable answers conditioned on complex medical images and questions. However, existing methods often overfit…

cs.CV2017

Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion

Weiyao Lin, Yang Mi, Jianxin Wu +2

Action recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recogn…

cs.CE2025

Enhancing Black-Litterman Portfolio via Hybrid Forecasting Model Combining Multivariate Decomposition and Noise Reduction

Ziye Yang, Ke Lu, Yang Wang +1

Modern portfolio construction demands robust methods for integrating data-driven insights into asset allocation. The Black-Litterman model offers a powerful Bayesian approach to ad…

eess.SY2018

Computation Load Balancing Real-Time Model Predictive Control in Urban Traffic Networks

Qiming Zou, Ke Lu, Yu Li

Owing to the rapid growth number of vehicles, urban traffic congestion has become more and more severe in the last decades. As an effective approach, Model Predictive Control (MPC)…

cs.CV2021

FEAFA+: An Extended Well-Annotated Dataset for Facial Expression Analysis and 3D Facial Animation

Wei Gan, Jian Xue, Ke Lu +3

Nearly all existing Facial Action Coding System-based datasets that include facial action unit (AU) intensity information annotate the intensity values hierarchically using A--E le…

cs.CV2019

Cycle-consistent Conditional Adversarial Transfer Networks

Jingjing Li, Erpeng Chen, Zhengming Ding +3

Domain adaptation investigates the problem of cross-domain knowledge transfer where the labeled source domain and unlabeled target domain have distinctive data distributions. Recen…

cs.LG2026

Clinically Interpretable Sepsis Early Warning via LLM-Guided Simulation of Temporal Physiological Dynamics

Weizhi Nie, Zhen Qu, Weijie Wang +4

Timely and interpretable early warning of sepsis remains a major clinical challenge due to the complex temporal dynamics of physiological deterioration. Traditional data-driven mod…

cs.CV2024

Agile Multi-Source-Free Domain Adaptation

Xinyao Li, Jingjing Li, Fengling Li +2

Efficiently utilizing rich knowledge in pretrained models has become a critical topic in the era of large models. This work focuses on adaptively utilizing knowledge from multiple…

cs.CV2026

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

Xiuyuan Zhu, Ke Lu, Kun Dong +6

Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify thi…

cs.CV2023

Neighborhood Contrastive Transformer for Change Captioning

Yunbin Tu, Liang Li, Li Su +2

Change captioning is to describe the semantic change between a pair of similar images in natural language. It is more challenging than general image captioning, because it requires…

cs.HC2026

DiffLens: A Visualization System to Explore Local Differences in Graph Sampling

Zhiguang Zhou, Yong Zhang, Yuming Ma +8

DiffLens is an interactive visualization tool that lets users examine how sampled graphs differ from the original graphs in terms of local node neighborhoods, paths, and overall st…

#graph sampling#network visualization#interactive analytics#difference analysis
cs.CE2026

Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives

Tengyue Xu, Zhuoyang Qian, Gaoge Liu +16

Autonomous scientific discovery with large language model (LLM)-based agents has recently made substantial progress, demonstrating the ability to automate end-to-end research workf…

cs.LG2019

FEAFA: A Well-Annotated Dataset for Facial Expression Analysis and 3D Facial Animation

Yanfu Yan, Ke Lu, Jian Xue +2

Facial expression analysis based on machine learning requires large number of well-annotated data to reflect different changes in facial motion. Publicly available datasets truly h…

cs.CV2026

QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval

Xiuyuan Zhu, Ke Lu, Zijie Yang +3

Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis. Existing approaches predomi…

eess.SY2024

Safe Reinforcement Learning-Based Eco-Driving Control for Mixed Traffic Flows With Disturbances

Ke Lu, Dongjun Li, Qun Wang +3

This paper presents a safe learning-based eco-driving framework tailored for mixed traffic flows, which aims to optimize energy efficiency while guaranteeing safety during real-sys…

cs.CV2017

Challenge of Multi-Camera Tracking

Yong Wang, Ke Lu

Multi-camera tracking is quite different from single camera tracking, and it faces new technology and system architecture challenges. By analyzing the corresponding characteristics…

cond-mat.supr-con2023

Superconductivity above 70 K observed in lutetium polyhydrides

Zhiwen Li, Xin He, Changling Zhang +17

The binary polyhydrides of heavy rare earth lutetium that shares a similar valence electron configuration to lanthanum have been experimentally discovered to be superconductive. Th…

cs.CV2023

G-Rep: Gaussian Representation for Arbitrary-Oriented Object Detection

Liping Hou, Ke Lu, Xue Yang +2

Typical representations for arbitrary-oriented object detection tasks include oriented bounding box (OBB), quadrilateral bounding box (QBB), and point set (PointSet). Each represen…

cs.CV2020

Characters as Graphs: Recognizing Online Handwritten Chinese Characters via Spatial Graph Convolutional Network

Ji Gan, Weiqiang Wang, Ke Lu

Chinese is one of the most widely used languages in the world, yet online handwritten Chinese character recognition (OLHCCR) remains challenging. To recognize Chinese characters, o…

cs.LG2018

Sample-Efficient Policy Learning based on Completely Behavior Cloning

Qiming Zou, Ling Wang, Ke Lu +1

Direct policy search is one of the most important algorithm of reinforcement learning. However, learning from scratch needs a large amount of experience data and can be easily pron…

cs.CV2021

Cross-Domain Gradient Discrepancy Minimization for Unsupervised Domain Adaptation

Zhekai Du, Jingjing Li, Hongzu Su +2

Unsupervised Domain Adaptation (UDA) aims to generalize the knowledge learned from a well-labeled source domain to an unlabeled target domain. Recently, adversarial domain adaptati…

cs.CV2024

MambaDETR: Query-based Temporal Modeling using State Space Model for Multi-View 3D Object Detection

Tong Ning, Ke Lu, Xirui Jiang +1

Utilizing temporal information to improve the performance of 3D detection has made great progress recently in the field of autonomous driving. Traditional transformer-based tempora…

cs.CV2024

Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation

Xinyao Li, Yuke Li, Zhekai Du +3

Large vision-language models (VLMs) like CLIP have demonstrated good zero-shot learning performance in the unsupervised domain adaptation task. Yet, most transfer approaches for VL…

cs.CV2019

Agile Domain Adaptation

Jingjing Li, Mengmeng Jing, Yue Xie +2

Domain adaptation investigates the problem of leveraging knowledge from a well-labeled source domain to an unlabeled target domain, where the two domains are drawn from different d…

cs.CV2014

Multiview Hessian regularized logistic regression for action recognition

W. Liu, H. Liu, D. Tao +2

With the rapid development of social media sharing, people often need to manage the growing volume of multimedia data such as large scale video classification and annotation, espec…

cs.CV2024

MVLLaVA: An Intelligent Agent for Unified and Flexible Novel View Synthesis

Hanyu Jiang, Jian Xue, Xing Lan +2

This paper introduces MVLLaVA, an intelligent agent designed for novel view synthesis tasks. MVLLaVA integrates multiple multi-view diffusion models with a large multimodal model,…