papers

Publications (34)

cond-mat.mes-hall2024

Electrically tunable, rapid spin-orbit torque induced modulation of colossal magnetoresistance in MnSiTe nanoflakes

Cheng Tan, Mingxun Deng, Yuanjun Yang +13

As a quasi-layered ferrimagnetic material, MnSiTe nanoflakes exhibit magnetoresistance behaviour that is fundamentally different from their bulk crystal counterparts. T…

cs.CV2022

Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-shot Learning

Yangji He, Weihan Liang, Dongyang Zhao +4

This paper presents new hierarchically cascaded transformers that can improve data efficiency through attribute surrogates learning and spectral tokens pooling. Vision transformers…

cs.CV2019

Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification from the Bottom Up

Weifeng Ge, Xiangru Lin, Yizhou Yu

Given a training dataset composed of images and corresponding category labels, deep convolutional neural networks show a strong ability in mining discriminative parts for image cla…

cs.CV2017

Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-tuning

Weifeng Ge, Yizhou Yu

Deep neural networks require a large amount of labeled training data during supervised learning. However, collecting and labeling so much data might be infeasible in many cases. In…

cond-mat.str-el2025

A language-inspired machine learning approach for solving strongly correlated problems with dynamical mean-field theory

Hovan Lee, Zelong Zhao, George Booth +2

We present SCALINN -- Strongly Correlated Approach with Language Inspired Neural Network -- as a method for solving the Anderson impurity model and reducing the computational cost…

cs.CV2018

Image Super-Resolution via Deterministic-Stochastic Synthesis and Local Statistical Rectification

Weifeng Ge, Bingchen Gong, Yizhou Yu

Single image superresolution has been a popular research topic in the last two decades and has recently received a new wave of interest due to deep neural networks. In this paper,…

cs.CL2026

Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions

Yutao Hou, Yajing Luo, Zhiwen Ruan +4

Large language models (LLMs) demonstrate remarkable performance across various tasks, prompting researchers to develop diverse evaluation benchmarks. However, most benchmarks typic…

cs.CV2022

FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos

Yan Wang, Yixuan Sun, Yiwen Huang +5

Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whet…

cs.CV2026

Adaptive Attention Distillation for Robust Few-Shot Segmentation under Environmental Perturbations

Qianyu Guo, Jingrong Wu, Jieji Ren +2

Few-shot segmentation (FSS) aims to rapidly learn novel class concepts from limited examples to segment specific targets in unseen images, and has been widely applied in areas such…

cs.CL2023

Improving Empathetic Dialogue Generation by Dynamically Infusing Commonsense Knowledge

Hua Cai, Xuli Shen, Qing Xu +5

In empathetic conversations, individuals express their empathy towards others. Previous work has mainly focused on generating empathetic responses by utilizing the speaker's emotio…

cs.CV2025

StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant

Haibo Wang, Bo Feng, Zhengfeng Lai +6

We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in ad…

cs.CV2024

Reading Relevant Feature from Global Representation Memory for Visual Object Tracking

Xinyu Zhou, Pinxue Guo, Lingyi Hong +4

Reference features from a template or historical frames are crucial for visual object tracking. Prior works utilize all features from a fixed template or memory for visual object t…

cs.CV2025

GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy Prediction

Tianhao Li, Yang Li, Mengtian Li +2

Accurately perceiving dynamic environments is a fundamental task for autonomous driving and robotic systems. Existing methods inadequately utilize temporal information, relying mai…

cs.CV2025

Boosting Salient Object Detection with Knowledge Distillated from Large Foundation Models

Miaoyang He, Shuyong Gao, Tsui Qin Mok +2

Salient Object Detection (SOD) aims to identify and segment prominent regions within a scene. Traditional models rely on manually annotated pseudo labels with precise pixel-level a…

cs.CV2022

GraphFPN: Graph Feature Pyramid Network for Object Detection

Gangming Zhao, Weifeng Ge, Yizhou Yu

Feature pyramids have been proven powerful in image understanding tasks that require multi-scale features. State-of-the-art methods for multi-scale feature learning focus on perfor…

cs.CV2025

Enhancing Environmental Robustness in Few-shot Learning via Conditional Representation Learning

Qianyu Guo, Jingrong Wu, Tianxing Wu +3

Few-shot learning (FSL) has recently been extensively utilized to overcome the scarcity of training data in domain-specific visual recognition. In real-world scenarios, environment…

cs.RO2025

Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation

Xiangyi Wei, Haotian Zhang, Xinyi Cao +4

The Vision-Language-Action models (VLA) have achieved significant advances in robotic manipulation recently. However, vision-only VLA models create fundamental limitations, particu…

cs.CV2024

Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge

Haibo Wang, Weifeng Ge

With the breakthrough of multi-modal large language models, answering complex visual questions that demand advanced reasoning abilities and world knowledge has become a much more i…

cs.CV2024

Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection

Jinglun Li, Xinyu Zhou, Pinxue Guo +4

Detecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling sche…

cs.MM2022

A Systematic Review on Affective Computing: Emotion Models, Databases, and Recent Advances

Yan Wang, Wei Song, Wei Tao +8

Affective computing plays a key role in human-computer interactions, entertainment, teaching, safe driving, and multimedia integration. Major breakthroughs have been made recently…

cs.CV2025

DeTrack: In-model Latent Denoising Learning for Visual Object Tracking

Xinyu Zhou, Jinglun Li, Lingyi Hong +4

Previous visual object tracking methods employ image-feature regression models or coordinate autoregression models for bounding box prediction. Image-feature regression methods hea…

cs.LG2026

Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization

Boya Xiong, Shuo Wang, Weifeng Ge +2

Supervised Fine-Tuning (SFT) empowers Large Language Models (LLMs) with exceptional performance on specialized tasks, but it yields dense, high-dimensional delta parameters that po…

cs.CL2026

AI Can Learn Scientific Taste

Jingqi Tong, Mingzhe Li, Hangcheng Li +20

The paper introduces a reinforcement‑learning framework that uses citation‑based community feedback to train models that can judge the impact of scientific papers and generate high…

#scientific discovery#reinforcement learning#large language models#citation prediction
cs.CV2024

Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering

Haibo Wang, Chenghang Lai, Yixuan Sun +1

Video Question Answering (VideoQA) aims to answer natural language questions based on the information observed in videos. Despite the recent success of Large Multimodal Models (LMM…

cs.CV2025

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Haibo Wang, Zhiyang Xu, Yu Cheng +6

Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal groundin…

cs.CV2022

ColoristaNet for Photorealistic Video Style Transfer

Xiaowen Qiu, Ruize Xu, Boan He +3

Photorealistic style transfer aims to transfer the artistic style of an image onto an input image or video while keeping photorealism. In this paper, we think it's the summary stat…

cs.CV2020

Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance Segmentation

Weifeng Ge, Sheng Guo, Weilin Huang +1

Weakly-supervised instance segmentation aims to detect and segment object instances precisely, given imagelevel labels only. Unlike previous methods which are composed of multiple…

cs.CV2018

Multi-Evidence Filtering and Fusion for Multi-Label Classification, Object Detection and Semantic Segmentation Based on Weakly Supervised Learning

Weifeng Ge, Sibei Yang, Yizhou Yu

Supervised object detection and semantic segmentation require object or even pixel level annotations. When there exist image level labels only, it is challenging for weakly supervi…

cs.CV2022

RankDNN: Learning to Rank for Few-shot Learning

Qianyu Guo, Hongtong Gong, Xujun Wei +4

This paper introduces a new few-shot learning pipeline that casts relevance ranking for image retrieval as binary ranking relation classification. In comparison to image classifica…

cs.CV2025

Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection

Jinglun Li, Kaixun Jiang, Zhaoyu Chen +4

Pre-trained vision-language models have exhibited remarkable abilities in detecting out-of-distribution (OOD) samples. However, some challenging OOD samples, which lie close to in-…

cs.CL2025

Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning

Jingqi Tong, Jixin Tang, Hangcheng Li +21

Vision-language reinforcement learning (RL) has primarily focused on narrow domains (e.g. geometry or chart reasoning). This leaves broader training scenarios and resources underex…

cs.CV2024

TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning

Jinglun Li, Xinyu Zhou, Kaixun Jiang +5

Multimodal fusion, leveraging data like vision and language, is rapidly gaining traction. This enriched data representation improves performance across various tasks. Existing meth…

cs.CV2021

Multi-scale Matching Networks for Semantic Correspondence

Dongyang Zhao, Ziyang Song, Zhenghao Ji +3

Deep features have been proven powerful in building accurate dense semantic correspondences in various previous works. However, the multi-scale and pyramidal hierarchy of convoluti…

cs.CV2018

Deep Metric Learning with Hierarchical Triplet Loss

Weifeng Ge, Weilin Huang, Dengke Dong +1

We present a novel hierarchical triplet loss (HTL) capable of automatically collecting informative training samples (triplets) via a defined hierarchical tree that encodes global c…