Publications (98)
KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning
Peng Yu, Cheng Deng, Beiya Dai +2
Autoregressive large language models (LLMs) pre-trained by next token prediction are inherently proficient in generative tasks. However, their performance on knowledge-driven tasks…
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
Jiahua Li, Zhanhe Zhang, Chenghao Xu +4
Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model…
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
Cheng Deng, Luoyang Sun, Jiwen Jiang +10
While scaling laws have been continuously validated in large language models (LLMs) with increasing model parameters, the inherent tension between the inference demands of LLMs and…
DSCD-Nav: Dual-Stance Cooperative Debate for Object Navigation
Weitao An, Qi Liu, Chenghao Xu +4
Adaptive navigation in unfamiliar indoor environments is crucial for household service robots. Despite advances in zero-shot perception and reasoning from vision-language models, e…
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
Siting Wang, Minnan Pei, Luoyang Sun +6
Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual informat…
Faster Stochastic Quasi-Newton Methods
Qingsong Zhang, Feihu Huang, Cheng Deng +1
Stochastic optimization methods have become a class of popular optimization tools in machine learning. Especially, stochastic gradient descent (SGD) has been widely used for machin…
Semantic Evidence Regulation via Relational Bias for Zero-Shot Object Navigation
Weitao An, Chenghao Xu, Xu Yang +1
Object navigation requires an embodied agent to locate a target object in an unknown environment through visual observations. Existing zero-shot methods typically leverage open-voc…
AStF: Motion Style Transfer via Adaptive Statistics Fusor
Hanmo Chen, Chenghao Xu, Jiexi Yan +1
Human motion style transfer allows characters to appear less rigidity and more realism with specific style. Traditional arbitrary image style transfer typically process mean and va…
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Haoyang Li, Zhanchao Xu, Yiming Li +9
Multi-turn dialogues are essential in many real-world applications of large language models, such as chatbots and virtual assistants. As conversation histories become longer, exist…
Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus
Tianhang Zhang, Lin Qiu, Qipeng Guo +6
Large Language Models (LLMs) have gained significant popularity for their impressive performance across diverse fields. However, LLMs are prone to hallucinate untruthful or nonsens…
Deep Spectral Clustering using Dual Autoencoder Network
Xu Yang, Cheng Deng, Feng Zheng +2
The clustering methods have recently absorbed even-increasing attention in learning and vision. Deep clustering combines embedding and clustering together to obtain optimal embeddi…
A Tale of Two Experts: Cooperative Learning for Source-Free Unsupervised Domain Adaptation
Jiaping Yu, Muli Yang, Jiapeng Ji +2
Source-Free Unsupervised Domain Adaptation (SFUDA) addresses the realistic challenge of adapting a source-trained model to a target domain without access to the source data, driven…
Theoretic Analysis and Extremely Easy Algorithms for Domain Adaptive Feature Learning
Wenhao Jiang, Cheng Deng, Wei Liu +3
Domain adaptation problems arise in a variety of applications, where a training dataset from the \textit{source} domain and a test dataset from the \textit{target} domain typically…
Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
Guangtao Lyu, Xinyi Cheng, Chenghao Xu +7
Large Vision-Language Models (LVLMs) have shown remarkable capabilities, yet hallucinations remain a persistent challenge. This work presents a systematic analysis of the internal…
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
Hanmo Chen, Chenghao Xu, Xu Yang +2
Video generation is pivotal to digital media creation, and recent advances in autoregressive video generation have markedly enhanced the efficiency of real-time video synthesis. Ho…
Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering
Guangtao Lyu, Xinyi Cheng, Qi Liu +5
LVLMs achieve remarkable multimodal understanding and generation but remain susceptible to hallucinations. Existing mitigation methods predominantly focus on output-level adjustmen…
Fewer is More: A Deep Graph Metric Learning Perspective Using Fewer Proxies
Yuehua Zhu, Muli Yang, Cheng Deng +1
Deep metric learning plays a key role in various machine learning tasks. Most of the previous works have been confined to sampling from a mini-batch, which cannot precisely charact…
Secure Bilevel Asynchronous Vertical Federated Learning with Backward Updating
Qingsong Zhang, Bin Gu, Cheng Deng +1
Vertical federated learning (VFL) attracts increasing attention due to the emerging demands of multi-party collaborative modeling and concerns of privacy leakage. In the real VFL a…
Desirable Companion for Vertical Federated Learning: New Zeroth-Order Gradient Based Algorithm
Qingsong Zhang, Bin Gu, Zhiyuan Dang +2
Vertical federated learning (VFL) attracts increasing attention due to the emerging demands of multi-party collaborative modeling and concerns of privacy leakage. A complete list o…
Not All Pixels Are Equal: Pixel-wise Meta-Learning for Medical Segmentation with Noisy Labels
Chenyu Mu, Guihai Chen, Xun Yang +2
Medical image segmentation is crucial for clinical applications, but it is frequently disrupted by noisy annotations and ambiguous anatomical boundaries, limiting its application i…
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
Luoyang Sun, Jiwen Jiang, Yifeng Ding +9
Vision-Language-Action Models (VLAs) have emerged as a key paradigm of Physical AI and are increasingly deployed in autonomous vehicles, robots, and smart spaces. In these resource…
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
Zhe Xu, Kun Wei, Xu Yang +1
Human dance generation (HDG) aims to synthesize realistic videos from images and sequences of driving poses. Despite great success, existing methods are limited to generating video…
Semantic Adversarial Network with Multi-scale Pyramid Attention for Video Classification
De Xie, Cheng Deng, Hao Wang +2
Two-stream architecture have shown strong performance in video classification task. The key idea is to learn spatio-temporal features by fusing convolutional networks spatially and…
Deep Multi-scale Discriminative Networks for Double JPEG Compression Forensics
Cheng Deng, Zhao Li, Xinbo Gao +1
As JPEG is the most widely used image format, the importance of tampering detection for JPEG images in blind forensics is self-evident. In this area, extracting effective statistic…
Fully-Featured Attribute Transfer
De Xie, Muli Yang, Cheng Deng +2
Image attribute transfer aims to change an input image to a target one with expected attributes, which has received significant attention in recent years. However, most of the exis…
Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph Generation
Yukuan Min, Aming Wu, Cheng Deng
The scene graph generation (SGG) task is designed to identify the predicates based on the subject-object pairs.However,existing datasets generally include two imbalance cases: one…
K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization
Cheng Deng, Tianhang Zhang, Zhongmou He +9
Large language models (LLMs) have achieved great success in general domains of natural language processing. In this paper, we bring LLMs to the realm of geoscience with the objecti…
AsySQN: Faster Vertical Federated Learning Algorithms with Better Computation Resource Utilization
Qingsong Zhang, Bin Gu, Cheng Deng +4
Vertical federated learning (VFL) is an effective paradigm of training the emerging cross-organizational (e.g., different corporations, companies and organizations) collaborative l…
NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens
Cunxiang Wang, Ruoxi Ning, Boqi Pan +8
Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of…
AceMap: Knowledge Discovery through Academic Graph
Xinbing Wang, Luoyi Fu, Xiaoying Gan +23
The exponential growth of scientific literature requires effective management and extraction of valuable insights. While existing scientific search engines excel at delivering sear…
Shared Predictive Cross-Modal Deep Quantization
Erkun Yang, Cheng Deng, Chao Li +3
With explosive growth of data volume and ever-increasing diversity of data modalities, cross-modal similarity search, which conducts nearest neighbor search across different modali…
Group Contrastive Self-Supervised Learning on Graphs
Xinyi Xu, Cheng Deng, Yaochen Xie +1
We study self-supervised learning on graphs using contrastive methods. A general scheme of prior methods is to optimize two-view representations of input graphs. In many studies, a…
Label Learning Method Based on Tensor Projection
Jing Li, Quanxue Gao, Qianqian Wang +2
Multi-view clustering method based on anchor graph has been widely concerned due to its high efficiency and effectiveness. In order to avoid post-processing, most of the existing a…
MSR: Making Self-supervised learning Robust to Aggressive Augmentations
Yingbin Bai, Erkun Yang, Zhaoqing Wang +5
Most recent self-supervised learning methods learn visual representation by contrasting different augmented views of images. Compared with supervised learning, more aggressive augm…
PK-Chat: Pointer Network Guided Knowledge Driven Generative Dialogue Model
Cheng Deng, Bo Tong, Luoyi Fu +4
In the research of end-to-end dialogue systems, using real-world knowledge to generate natural, fluent, and human-like utterances with correct answers is crucial. However, domain-s…
Covidia: COVID-19 Interdisciplinary Academic Knowledge Graph
Cheng Deng, Jiaxin Ding, Luoyi Fu +3
The pandemic of COVID-19 has inspired extensive works across different research fields. Existing literature and knowledge platforms on COVID-19 only focus on collecting papers on b…
Towards Arbitrary Motion Completing via Hierarchical Continuous Representation
Chenghao Xu, Guangtao Lyu, Qi Liu +3
Physical motions are inherently continuous, and higher camera frame rates typically contribute to improved smoothness and temporal coherence. For the first time, we explore continu…
Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
Xiang Zhang, Kun Wei, Xu Yang +3
As Large Language Models (LLMs) become increasingly prevalent, their security vulnerabilities have already drawn attention. Machine unlearning is introduced to seek to mitigate the…
Self-Supervised Graph Embedding Clustering
Fangfang Li, Quanxue Gao, Cheng Deng +1
The K-means one-step dimensionality reduction clustering method has made some progress in addressing the curse of dimensionality in clustering tasks. However, it combines the K-mea…
MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework
Qirui Mi, Mengyue Yang, Xiangning Yu +6
Simulating collective decision-making involves more than aggregating individual behaviors; it emerges from dynamic interactions among individuals. While large language models (LLMs…
AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing
Huawei Ji, Cheng Deng, Bo Xue +6
With the development of data-centric AI, the focus has shifted from model-driven approaches to improving data quality. Academic literature, as one of the crucial types, is predomin…
Towards Visual Feature Translation
Jie Hu, Rongrong Ji, Hong Liu +3
Most existing visual search systems are deployed based upon fixed kinds of visual features, which prohibits the feature reusing across different systems or when upgrading systems w…
Invisible Backdoor Attack with Dynamic Triggers against Person Re-identification
Wenli Sun, Xinyang Jiang, Shuguang Dou +4
In recent years, person Re-identification (ReID) has rapidly progressed with wide real-world applications, but also poses significant risks of adversarial attacks. In this paper, w…
ContextPilot: Fast Long-Context Inference via Context Reuse
Yinsicheng Jiang, Yeqi Huang, Liang Cheng +3
AI applications increasingly depend on long-context inference, where LLMs consume substantial context to support stronger reasoning. Common examples include retrieval-augmented gen…
Adaptive Hierarchical Similarity Metric Learning with Noisy Labels
Jiexi Yan, Lei Luo, Cheng Deng +1
Deep Metric Learning (DML) plays a critical role in various machine learning tasks. However, most existing deep metric learning methods with binary similarity are sensitive to nois…
Domain-Smoothing Network for Zero-Shot Sketch-Based Image Retrieval
Zhipeng Wang, Hao Wang, Jiexi Yan +2
Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a novel cross-modal retrieval task, where abstract sketches are used as queries to retrieve natural images under zero-shot scena…
Good Idea or Not, Representation of LLM Could Tell
Yi Xu, Bo Xue, Shuqian Sheng +6
In the ever-expanding landscape of academic research, the proliferation of ideas presents a significant challenge for researchers: discerning valuable ideas from the less impactful…
Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning
Hanmo Chen, Guangtao Lyu, Chenghao Xu +3
As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intu…
The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation
Chenyu Mu, Xin He, Qu Yang +13
Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, these models struggle to generate long-f…
DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning
Siyuan Guo, Cheng Deng, Ying Wen +3
In this work, we investigate the potential of large language models (LLMs) based agents to automate data science tasks, with the goal of comprehending task requirements, then build…
Incremental Embedding Learning via Zero-Shot Translation
Kun Wei, Cheng Deng, Xu Yang +1
Modern deep learning methods have achieved great success in machine learning and computer vision fields by learning a set of pre-defined datasets. Howerver, these methods perform u…
Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning
Xiangyu Li, Xu Yang, Kun Wei +2
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen compositions formed from seen state and object during training. Since the same state may be various in the visual a…
Interpretable Multi-View Clustering Based on Anchor Graph Tensor Factorization
Rui Wang, Jing Li, Quanxue Gao +1
The clustering method based on the anchor graph has gained significant attention due to its exceptional clustering performance and ability to process large-scale data. One common a…
Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models
Mengxin Qin, Xiang Zhang, Kun Wei +2
Class-incremental learning aims to continuously acquire new knowledge while preserving previously learned information, thereby mitigating catastrophic forgetting. Existing methods…
GeoGalactica: A Scientific Large Language Model in Geoscience
Zhouhan Lin, Cheng Deng, Le Zhou +18
Large language models (LLMs) have achieved huge success for their general knowledge and ability to solve a wide spectrum of tasks in natural language processing (NLP). Due to their…
FMGNN: Fused Manifold Graph Neural Network
Cheng Deng, Fan Xu, Jiaxing Ding +3
Graph representation learning has been widely studied and demonstrated effectiveness in various graph tasks. Most existing works embed graph data in the Euclidean space, while rece…
Towards Improved and Interpretable Deep Metric Learning via Attentive Grouping
Xinyi Xu, Zhengyang Wang, Cheng Deng +2
Grouping has been commonly used in deep metric learning for computing diverse features. However, current methods are prone to overfitting and lack interpretability. In this work, w…
Compensating Visual Insufficiency with Stratified Language Guidance for Long-Tail Class Incremental Learning
Xi Wang, Xu Yang, Donghao Sun +1
Long-tail class incremental learning (LT CIL) remains highly challenging because the scarcity of samples in tail classes not only hampers their learning but also exacerbates catast…
Revitalizing the Beginning: Avoiding Storage Dependency for Model Merging in Continual Learning
Xi Wang, Cheng Deng
Model merging provides a compelling paradigm for integrating specialized expertise into a unified multi-task model, a goal that aligns naturally with the sequential knowledge acqui…
A Turn Toward Better Alignment: Few-Shot Generative Adaptation with Equivariant Feature Rotation
Chenghao Xu, Qi Liu, Jiexi Yan +2
Few-shot image generation aims to effectively adapt a source generative model to a target domain using very few training images. Most existing approaches introduce consistency cons…
One-Step Multi-View Clustering Based on Transition Probability
Wenhui Zhao, Quanxue Gao, Guangfei Li +2
The large-scale multi-view clustering algorithms, based on the anchor graph, have shown promising performance and efficiency and have been extensively explored in recent years. Des…
DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models
Mengxin Qin, Xiang Zhang, Xi Wang +3
Continual learning enables vision-language models to accumulate knowledge and adapt to evolving tasks without retraining from scratch. However, in multi-domain task-incremental lea…
Projection & Probability-Driven Black-Box Attack
Jie Li, Rongrong Ji, Hong Liu +4
Generating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer f…
Stacked Semantic-Guided Network for Zero-Shot Sketch-Based Image Retrieval
Hao Wang, Cheng Deng, Xinxu Xu +3
Zero-shot sketch-based image retrieval (ZS-SBIR) is a task of cross-domain image retrieval from a natural image gallery with free-hand sketch under a zero-shot scenario. Previous w…
High-Discriminative Attribute Feature Learning for Generalized Zero-Shot Learning
Yu Lei, Guoshuai Sheng, Fangfang Li +3
Zero-shot learning(ZSL) aims to recognize new classes without prior exposure to their samples, relying on semantic knowledge from observed classes. However, current attention-based…
Pyramidal Person Re-IDentification via Multi-Loss Dynamic Training
Feng Zheng, Cheng Deng, Xing Sun +5
Most existing Re-IDentification (Re-ID) methods are highly dependent on precise bounding boxes that enable images to be aligned with each other. However, due to the challenging pra…
GTA: Grouped-head latenT Attention
Luoyang Sun, Cheng Deng, Jiwen Jiang +5
Attention mechanisms underpin the success of large language models (LLMs), yet their substantial computational and memory overhead poses challenges for optimizing efficiency and pe…
-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
Siting Wang, Xiaofeng Wang, Zheng Zhu +7
Flow-based vision-language-action (VLA) models excel in embodied control but suffer from intractable likelihoods during multi-step sampling, hindering online reinforcement learning…
Multimodal Fusion on Low-quality Data: A Comprehensive Survey
Qingyang Zhang, Yake Wei, Zongbo Han +8
Multimodal fusion focuses on integrating information from multiple modalities with the goal of more accurate prediction, which has achieved remarkable progress in a wide range of s…
Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
Gen Luo, Yiyi Zhou, Xiaoshuai Sun +4
Referring expression comprehension (REC) and segmentation (RES) are two highly-related tasks, which both aim at identifying the referent according to a natural language expression.…
Learning Stateful Predictive Knowledge From Experience
Yan Song, Xidong Feng, Bo Liu +7
As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predicti…
Active Transfer Learning Network: A Unified Deep Joint Spectral-Spatial Feature Learning Model For Hyperspectral Image Classification
Cheng Deng, Yumeng Xue, Xianglong Liu +2
Deep learning has recently attracted significant attention in the field of hyperspectral images (HSIs) classification. However, the construction of an efficient deep neural network…
3D Molecular Geometry Analysis with 2D Graphs
Zhao Xu, Yaochen Xie, Youzhi Luo +7
Ground-state 3D geometries of molecules are essential for many molecular analysis tasks. Modern quantum mechanical methods can compute accurate 3D geometries but are computationall…
Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs
Guangtao Lyu, Qi Liu, Chenghao Xu +5
LVLMs have achieved strong multimodal reasoning capabilities but remain prone to hallucinations, producing outputs inconsistent with visual inputs or user instructions. Existing tr…
Active Multi-Kernel Domain Adaptation for Hyperspectral Image Classification
Cheng Deng, Xianglong Liu, Chao Li +1
Recent years have witnessed the quick progress of the hyperspectral images (HSI) classification. Most of existing studies either heavily rely on the expensive label information usi…
Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
Shaohui Lin, Rongrong Ji, Yuchao Li +2
The success of convolutional neural networks (CNNs) in computer vision applications has been accompanied by a significant increase of computation and memory costs, which prohibits…
Coupled CycleGAN: Unsupervised Hashing Network for Cross-Modal Retrieval
Chao Li, Cheng Deng, Lei Wang +2
In recent years, hashing has attracted more and more attention owing to its superior capacity of low storage cost and high query efficiency in large-scale cross-modal retrieval. Be…
Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation
Guangtao Lyu, Chenghao Xu, Qi Liu +4
Music to 3D dance generation aims to synthesize realistic and rhythmically synchronized human dance from music. While existing methods often rely on additional genre labels to furt…
Privacy-Preserving Asynchronous Federated Learning Algorithms for Multi-Party Vertically Collaborative Learning
Bin Gu, An Xu, Zhouyuan Huo +2
The privacy-preserving federated learning for vertically partitioned data has shown promising results as the solution of the emerging multi-party joint modeling application, in whi…
Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy Minimization
Kamran Ghasedi Dizaji, Amirhossein Herandi, Cheng Deng +2
Image clustering is one of the most important computer vision applications, which has been extensively studied in literature. However, current clustering methods mostly suffer from…
Similarity Distribution based Membership Inference Attack on Person Re-identification
Junyao Gao, Xinyang Jiang, Huishuai Zhang +6
While person Re-identification (Re-ID) has progressed rapidly due to its wide real-world applications, it also causes severe risks of leaking personal information from training dat…
DistillHash: Unsupervised Deep Hashing by Distilling Data Pairs
Erkun Yang, Tongliang Liu, Cheng Deng +2
Due to the high storage and search efficiency, hashing has become prevalent for large-scale similarity search. Particularly, deep hashing methods have greatly improved the search p…
Query-Adaptive Hash Code Ranking for Large-Scale Multi-View Visual Search
Xianglong Liu, Lei Huang, Cheng Deng +2
Hash based nearest neighbor search has become attractive in many applications. However, the quantization in hashing usually degenerates the discriminative power when using Hamming…
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
Zibo Su, Kun Wei, Jiahua Li +3
Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-Engl…
Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval
Chao Li, Cheng Deng, Ning Li +3
Thanks to the success of deep learning, cross-modal retrieval has made significant progress recently. However, there still remains a crucial bottleneck: how to bridge the modality…
Semantic Adversarial Network for Zero-Shot Sketch-Based Image Retrieval
Xinxun Xu, Hao Wang, Leida Li +1
Zero-shot sketch-based image retrieval (ZS-SBIR) is a specific cross-modal retrieval task for retrieving natural images with free-hand sketches under zero-shot scenario. Previous w…
CAMS: Towards Compositional Zero-Shot Learning via Gated Cross-Attention and Multi-Space Disentanglement
Pan Yang, Cheng Deng, Jing Yang +5
Compositional zero-shot learning (CZSL) aims to learn the concepts of attributes and objects in seen compositions and to recognize their unseen compositions. Most Contrastive Langu…
CerfGAN: A Compact, Effective, Robust, and Fast Model for Unsupervised Multi-Domain Image-to-Image Translation
Xiao Liu, Shengchuan Zhang, Hong Liu +3
In this paper, we aim at solving the multi-domain image-to-image translation problem with a unified model in an unsupervised manner. The most successful work in this area refers to…
Taming Transformer for Emotion-Controllable Talking Face Generation
Ziqi Zhang, Cheng Deng
Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talki…
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
Qingfa Xiao, Jiachuan Wang, Haoyang Li +6
Recent advances in large language models (LLMs) have showcased exceptional performance in long-context tasks, while facing significant inference efficiency challenges with limited…
Triplet-Based Deep Hashing Network for Cross-Modal Retrieval
Cheng Deng, Zhaojia Chen, Xianglong Liu +2
Given the benefits of its low storage requirements and high retrieval efficiency, hashing has recently received increasing attention. In particular,cross-modal hashing has been wid…
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
Muning Wen, Junwei Liao, Cheng Deng +3
Large Language Models (LLMs) have shown promise as intelligent agents in interactive decision-making tasks. Traditional approaches often depend on meticulously designed prompts, hi…
Molecule3D: A Benchmark for Predicting 3D Geometries from Molecular Graphs
Zhao Xu, Youzhi Luo, Xuan Zhang +7
Graph neural networks are emerging as promising methods for modeling molecular graphs, in which nodes and edges correspond to atoms and chemical bonds, respectively. Recent studies…
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
Zhen Bi, Xueshu Chen, Luoyang Sun +4
The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characterization on resource-constrained edge ha…
VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models
Huawei Ji, Yuanhao Sun, Yuan Jin +4
Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vision-language models (VLMs).…
Doubly Contrastive Deep Clustering
Zhiyuan Dang, Cheng Deng, Xu Yang +1
Deep clustering successfully provides more effective features than conventional ones and thus becomes an important technique in current unsupervised learning. However, most deep cl…
Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection
Hao Wang, Cheng Deng, Zhidong Zhao
Recent generative models demonstrate impressive performance on synthesizing photographic images, which makes humans hardly to distinguish them from pristine ones, especially on rea…
Logical Structure as Knowledge: Enhancing LLM Reasoning via Structured Logical Knowledge Density Estimation
Zhen Bi, Zhenlin Hu, Xueshu Chen +7
The reasoning capabilities of Large Language Models (LLMs) are increasingly attributed to training data quality rather than mere parameter scaling. However, existing data-centric p…