papers

Publications (106)

cs.LG2025

Efficient Quantification of Multimodal Interaction at Sample Level

Zequn Yang, Hongfa Wang, Di Hu

Interactions between modalities -- redundancy, uniqueness, and synergy -- collectively determine the composition of multimodal information. Understanding these interactions is cruc…

cs.CV2020

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

Di Hu, Rui Qian, Minyue Jiang +5

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a…

cs.SD2023

Multi-Scale Attention for Audio Question Answering

Guangyao Li, Yixin Xu, Di Hu

Audio question answering (AQA), acting as a widely used proxy task to explore scene understanding, has got more attention. The AQA is challenging for it requires comprehensive temp…

physics.ins-det2017

Measurement of frequency sweep nonlinearity using atomic absorption spectroscopy

Ningfang Song, Xiangxiang Lu, Xiaobin Xu +4

A low cost scheme to determine the frequency sweep nonlinearity using atomic saturated absorption spectroscopy is demonstrated. The frequency modulation rate is determined by direc…

cs.LG2022

Towards Accurate Knowledge Transfer via Target-awareness Representation Disentanglement

Xingjian Li, Di Hu, Xuhong Li +5

Fine-tuning deep neural networks pre-trained on large scale datasets is one of the most practical transfer learning paradigm given limited quantity of training samples. To obtain b…

cs.SE2021

Parsing Data Formats of the Inputs and Outputs of Geographic Models with Code Analysis

Xinghua Cheng, Di Hu, Handong He +2

Model web services provide an approach for implementing and facilitating the sharing of geographic models. The description and acquisition of inputs and outputs (IO) of geographic…

physics.plasm-ph2018

Energy spectrum of tearing mode turbulence in sheared background field

Di Hu, Amitava Bhattacharjee, Yi-Min Huang

The energy spectrum of tearing mode turbulence in a sheared background magnetic field is studied in this work. We consider the scenario where the nonlinear interaction of overlappi…

cs.CV2024

Diagnosing and Re-learning for Balanced Multimodal Learning

Yake Wei, Siwei Li, Ruoxuan Feng +1

To overcome the imbalanced multimodal learning problem, where models prefer the training of specific modalities, existing methods propose to control the training of uni-modal encod…

cs.MM2022

SeCo: Separating Unknown Musical Visual Sounds with Consistency Guidance

Xinchi Zhou, Dongzhan Zhou, Wanli Ouyang +3

Recent years have witnessed the success of deep learning on the visual sound separation task. However, existing works follow similar settings where the training and testing dataset…

cs.LG2025

Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition

Chengxiang Huang, Yake Wei, Zequn Yang +1

Sensory training during the early ages is vital for human development. Inspired by this cognitive phenomenon, we observe that the early training stage is also important for the mul…

cs.CV2024

SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model

Tao Wu, Xuewei Li, Zhongang Qi +4

Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains.However, it remains a challenging task due to the inherent sph…

cs.CV2021

Class-aware Sounding Objects Localization via Audiovisual Correspondence

Di Hu, Yake Wei, Rui Qian +3

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achie…

cs.CY2024

An Interactive Web Application for School-Based Physical Fitness Testing in California: Geospatial Analysis and Custom Mapping

Yawen Guo, Kaiyuan Hu, Di Hu +2

Physical activity is essential for children's healthy growth and development. In the US, most states, including California, adhere to physical education standards and have implemen…

cs.CV2023

Progressive Spatio-temporal Perception for Audio-Visual Question Answering

Guangyao Li, Wenxuan Hou, Di Hu

Audio-Visual Question Answering (AVQA) task aims to answer questions about different visual objects, sounds, and their associations in videos. Such naturally multi-modal videos are…

cs.CV2026

MIBench: Evaluating LMMs on Multimodal Interaction

Yu Miao, Zequn Yang, Yake Wei +5

In different multimodal scenarios, it needs to integrate and utilize information across modalities in a specific way based on the demands of the task. Different integration ways be…

cs.CV2018

Dense Multimodal Fusion for Hierarchically Joint Representation

Di Hu, Feiping Nie, Xuelong Li

Multiple modalities can provide more valuable information than single one by describing the same contents in various ways. Hence, it is highly expected to learn effective joint rep…

cs.AI2025

Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI

Sha Zhang, Suorong Yang, Tong Xie +18

Scientific discovery has long been constrained by human limitations in expertise, physical capability, and sleep cycles. The recent rise of AI scientists and automated laboratories…

cs.LG2022

Not All Knowledge Is Created Equal: Mutual Distillation of Confident Knowledge

Ziyun Li, Xinshao Wang, Di Hu +4

Mutual knowledge distillation (MKD) improves a model by distilling knowledge from another model. However, \textit{not all knowledge is certain and correct}, especially under advers…

astro-ph.SR2020

The Unusual Eruption of the Extragalactic Classical Nova M31N 2017-09a

Christopher Lloyd, Lewis M. Cook, Seiichiro Kiyota +10

M31N 2017-09a is a classical nova and was observed for some 160 days following its initial eruption, during which time it underwent a number of bright secondary outbursts. The ligh…

cs.CV2026

APPO: Attention-guided Perception Policy Optimization for Video Reasoning

Henghui Du, Chang Zhou, Xi Chen +1

Complex video reasoning, actually, relies excessively on fine-grained perception rather than on expert (e.g., Ph.D, Science)-level reasoning. Through extensive empirical observatio…

cs.CV2024

Boosting Audio Visual Question Answering via Key Semantic-Aware Cues

Guangyao Li, Henghui Du, Di Hu

The Audio Visual Question Answering (AVQA) task aims to answer questions related to various visual objects, sounds, and their interactions in videos. Such naturally multimodal vide…

cs.MM2024

Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning

Meng Shen, Yake Wei, Jianxiong Yin +3

Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on s…

cs.CV2024

Unveiling and Mitigating Bias in Audio Visual Segmentation

Peiwen Sun, Honggang Zhang, Di Hu

Community researchers have developed a range of advanced audio-visual segmentation models aimed at improving the quality of sounding objects' masks. While masks created by these mo…

physics.app-ph2024

Facile synthesis of fine-grained CoFeO anchored on porous carbon for simultaneous removal of tetracycline and arsenite

Yuwen Chen, Ke Zhu, Yizhe Huang +6

The coexistence of tetracycline (TC) and arsenite (As(III)) in livestock wastewater threatens public health, and the heterogeneous Fenton-like system is a practical approach for th…

cs.LG2026

Information-Theoretic Decomposition for Multimodal Interaction Learning

Zequn Yang, Yake Wei, Haotian Ni +2

Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet unde…

cs.CV2024

Quantifying and Enhancing Multi-modal Robustness with Modality Preference

Zequn Yang, Yake Wei, Ce Liang +1

Multi-modal models have shown a promising capability to effectively integrate information from various sources, yet meanwhile, they are found vulnerable to pervasive perturbations,…

cs.CV2025

Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Henghui Du, Guangyao Li, Chang Zhou +3

In recent years, numerous tasks have been proposed to encourage model to develop specified capability in understanding audio-visual scene, primarily categorized into temporal local…

cs.CV2023

Towards Inadequately Pre-trained Models in Transfer Learning

Andong Deng, Xingjian Li, Di Hu +3

Pre-training has been a popular learning paradigm in deep learning era, especially in annotation-insufficient scenario. Better ImageNet pre-trained models have been demonstrated, f…

cs.CL2025

Understanding Stigmatizing Language Lexicons: A Comparative Analysis in Clinical Contexts

Yiliang Zhou, Di Hu, Tianchu Lyu +4

Stigmatizing language results in healthcare inequities, yet there is no universally accepted or standardized lexicon defining which words, terms, or phrases constitute stigmatizing…

physics.plasm-ph2023

Drift surface solver for runaway electron current dominant equilibria during the Current Quench

Lu Yuan, Di Hu

Runaway electron current generated during the Current Quench phase of tokamak disruptions could result in severe damage to future high performance devices. To control and mitigate…

physics.plasm-ph2016

On the inward drift of runaway electrons during the plateau phase of runaway current

Di Hu, Hong Qin

The well observed inward drift of current carrying runaway electrons during runaway plateau regime after disruption is studied by considering the phase space dynamic of runaways in…

cs.CV2020

Ambient Sound Helps: Audiovisual Crowd Counting in Extreme Conditions

Di Hu, Lichao Mou, Qingzhong Wang +4

Visual crowd counting has been recently studied as a way to enable people counting in crowd scenes from images. Albeit successful, vision-based crowd counting approaches could fail…

cs.CV2024

MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance

Yake Wei, Di Hu

Multimodal learning methods with targeted unimodal learning objectives have exhibited their superior efficacy in alleviating the imbalanced multimodal learning problem. However, in…

cs.RO2024

Learning Manipulation by Predicting Interaction

Jia Zeng, Qingwen Bu, Bangjun Wang +12

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage larg…

cs.CV2024

Can Textual Semantics Mitigate Sounding Object Segmentation Preference?

Yaoting Wang, Peiwen Sun, Yuanchao Li +2

The Audio-Visual Segmentation (AVS) task aims to segment sounding objects in the visual space using audio cues. However, in this work, it is recognized that previous AVS methods sh…

cs.HC2025

Supporting Patients in Managing Electronic Health Records and Biospecimens Consent for Research: Insights from a Mixed-Methods Usability Evaluation of the iAGREE Portal

Di Hu, Xi Lu, Yunan Chen +6

De-identified health data are frequently used in research. As AI advances heighten the risk of re-identification, it is important to respond to concerns about transparency, data pr…

cs.CV2022

Balanced Multimodal Learning via On-the-fly Gradient Modulation

Xiaokang Peng, Yake Wei, Andong Deng +2

Multimodal learning helps to comprehensively understand the world, by integrating different senses. Accordingly, multiple input modalities are expected to boost model performance,…

cs.HC2024

Investigating the effects of housing instability on depression, anxiety, and mental health treatment in childhood and adolescence

Rachael Zehrung, Di Hu, Yawen Guo +2

Housing instability is a widespread phenomenon in the United States. In combination with other social determinants of health, housing instability affects children's overall health…

cs.CV2026

Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning

Yake Wei, Yuan Wang, Fengyun Rao +2

Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with i…

cs.CL2024

YuLan: An Open-source Large Language Model

Yutao Zhu, Kun Zhou, Kelong Mao +35

Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many o…

cs.CY2025

When AI Writes Back: Ethical Considerations by Physicians on AI-Drafted Patient Message Replies

Di Hu, Yawen Guo, Ha Na Cho +7

The increasing burden of responding to large volumes of patient messages has become a key factor contributing to physician burnout. Generative AI (GenAI) shows great promise to all…

cs.CV2020

Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition

Di Hu, Xuhong Li, Lichao Mou +5

Aerial scene recognition is a fundamental task in remote sensing and has recently received increased interest. While the visual information from overhead images with powerful model…

cond-mat.mes-hall2023

Modulation of skyrmionic magnetic textures in two-dimensional vdW materials and their heterostructures

Xiaoyan Yao, Di Hu, Shuai Dong

The intrinsic magnetism observed in two-dimensional (2D) van der Waals (vdW) materials provides a unique opportunity for exploring the 2D topological magnetic textures, in particul…

cs.CV2020

Curriculum Audiovisual Learning

Di Hu, Zheng Wang, Haoyi Xiong +3

Associating sound and its producer in complex audiovisual scene is a challenging task, especially when we are lack of annotated training data. In this paper, we present a flexible…

cs.RO2024

KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance

Jingxian Lu, Wenke Xia, Dong Wang +4

Online Imitation Learning struggles with the gap between extensive online exploration space and limited expert trajectories, hindering efficient exploration due to inaccurate rewar…

cs.CV2022

Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Guangyao Li, Yake Wei, Yapeng Tian +3

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in vid…

cs.CV2025

Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

Ruotian Peng, Haiying He, Yake Wei +2

High-quality image captions play a crucial role in improving the performance of cross-modal applications such as text-to-image generation, text-to-video generation, and text-image…

cs.RO2026

AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception

Ruoxuan Feng, Yuxuan Zhou, Siyu Mei +6

Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties as well as force…

physics.plasm-ph2024

Stability impacts from the current and pressure profile modifications within finite sized island

Yuxiang Sun, Di Hu

The stability (or instability) of finite sized magnetic island could play a significant role in disruption avoidance or disruption mitigation dynamics. Especially, various current…

cs.LG2023

Balanced Audiovisual Dataset for Imbalance Analysis

Wenke Xia, Xu Zhao, Xincheng Pang +2

The imbalance problem is widespread in the field of machine learning, which also exists in multimodal learning areas caused by the intrinsic discrepancy between modalities of sampl…

cs.LG2025

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

Haotian Ni, Yake Wei, Hang Liu +4

Multimodal learning faces challenges in effectively fusing information from diverse modalities, especially when modality quality varies across samples. Dynamic fusion strategies, s…

cs.SE2021

Integrating Structural Description of Data Format Information into Programming to Auto-generate File Reading Programs

Xinghua Cheng, Erjie Hu, Di Hu

File reading is the basis for data sharing and scientific computing. However, manual programming for file reading is labour-intensive and time-consuming, as data formats are hetero…

cond-mat.supr-con2014

Analysis of fields in an air-cored superconducting synchronous motor with an HTS racetrack field winding

Di Hu, Jin Zou, Tim J. Flack +3

High temperature superconducting (HTS) synchronous motors can offer significant weight and size reductions, as well as improved efficiency, over conventional copper-wound machines…

cs.CV2024

Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Yaoting Wang, Weisong Liu, Guangyao Li +3

Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on th…

cs.CV2022

Dual Domain-Adversarial Learning for Audio-Visual Saliency Prediction

Yingzi Fan, Longfei Han, Yue Zhang +3

Both visual and auditory information are valuable to determine the salient regions in videos. Deep convolution neural networks (CNN) showcase strong capacity in coping with the aud…

cs.LG2025

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

Ruoxuan Feng, Jiangyu Hu, Wenke Xia +5

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tacti…

cs.CV2024

Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation

Juncheng Ma, Peiwen Sun, Yaoting Wang +1

Audio-Visual Segmentation (AVS) aims to achieve pixel-level localization of sound sources in videos, while Audio-Visual Semantic Segmentation (AVSS), as an extension of AVS, furthe…

cs.HC2026

We Need Granular Sharing of De-Identified Data-But Will Patients Engage? Investigating Health System Leaders' and Patients' Perspectives on A Patient-Controlled Data-Sharing Platform

Xi Lu, Di Hu, An T. Nguyen +6

Patient-controlled data-sharing systems are increasingly promoted as a way to empower patients with greater autonomy over their health data. Yet it remains unclear how different st…

cs.RO2026

When would Vision-Proprioception Policies Fail in Robotic Manipulation?

Jingxian Lu, Wenke Xia, Yuxuan Wu +2

Proprioceptive information is critical for precise servo control by providing real-time robotic states. Its collaboration with vision is highly expected to enhance performances of…

cs.CV2020

Temporal Relational Modeling with Self-Supervision for Action Segmentation

Dong Wang, Di Hu, Xingjian Li +1

Temporal relational modeling in video is essential for human action understanding, such as action recognition and action segmentation. Although Graph Convolution Networks (GCNs) ha…

cond-mat.str-el2023

Two-orbital spin-fermion model study of ferromagnetism in honeycomb lattice

Kaidi Xu, Di Hu, Jun Chen +4

The spin-fermion model was previously successful to describe the complex phase diagrams of colossal magnetoresistive manganites and iron-based superconductors. In recent years, two…

cs.MM2023

Towards Long Form Audio-visual Video Understanding

Wenxuan Hou, Guangyao Li, Yapeng Tian +1

We live in a world filled with never-ending streams of multimodal information. As a more natural recording of the real scenario, long form audio-visual videos are expected as an im…

cs.RO2025

Human-assisted Robotic Policy Refinement via Action Preference Optimization

Wenke Xia, Yichu Yang, Hongtao Wu +3

Establishing a reliable and iteratively refined robotic system is essential for deploying real-world applications. While Vision-Language-Action (VLA) models are widely recognized a…

cs.CV2017

Deep Binary Reconstruction for Cross-modal Hashing

Xuelong Li, Di Hu, Feiping Nie

With the increasing demand of massive multimodal data storage and organization, cross-modal retrieval based on hashing technique has drawn much attention nowadays. It takes the bin…

cs.MM2025

Object-AVEdit: An Object-level Audio-Visual Editing Model

Youquan Fu, Ruiyang Si, Hongfa Wang +6

There is a high demand for audio-visual editing in video post-production and the film making field. While numerous models have explored audio and video editing, they struggle with…

cs.CV2020

Multiple Sound Sources Localization from Coarse to Fine

Rui Qian, Di Hu, Heinrich Dinkel +3

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this proble…

cs.CV2017

Image2song: Song Retrieval via Bridging Image Content and Lyric Words

Xuelong Li, Di Hu, Xiaoqiang Lu

Image is usually taken for expressing some kinds of emotions or purposes, such as love, celebrating Christmas. There is another better way that combines the image and relevant song…

cs.CV2026

GeoISF: Instance Semantic Forest Inspired Large-Scale Cross-View Geo-Localization via Ground LiDAR-to-Satellite Image

Di Hu, Xia Yuan, Chunxia Zhao

The problem of localization on a large-scale satellite image given a frame of query ground view point clouds remains challenging. Existing LiDAR-to-image cross-view localization me…

cs.RO2026

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Shihan Wu, Xuecheng Liu, Shaoxuan Xie +81

Despite the critical role of bimanual manipulation in endowing robots with human-like dexterity, large-scale and diverse datasets remain scarce due to the significant hardware hete…

cs.CV2026

Crab: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Dongnuan Cai, Henghui Du, Chang Zhou +5

Developing Audio-Visual Large Language Models (AV-LLMs) for unified scene understanding is pivotal in multimodal intelligence. While instruction tuning enables pre-trained models w…

cs.CV2023

Revisiting Pre-training in Audio-Visual Learning

Ruoxuan Feng, Wenke Xia, Di Hu

Pre-training technique has gained tremendous success in enhancing model performance on various tasks, but found to perform worse than training from scratch in some uni-modal situat…

cond-mat.supr-con2015

DC Characterization of a Circular, Epoxy-Impregnated High Temperature Superconducting (HTS) Coil

Di Hu, Mark D. Ainslie, Jordan P. Rush +2

Direct current (DC) characterization of high temperature superconducting (HTS) coils is important for HTS applications, such as electric machines, superconducting magnetic energy s…

cs.CV2024

Self-supervised Audiovisual Representation Learning for Remote Sensing Data

Konrad Heidler, Lichao Mou, Di Hu +5

Many current deep learning approaches make extensive use of backbone networks pre-trained on large datasets like ImageNet, which are then fine-tuned to perform a certain task. In r…

cs.RO2024

Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection

Xincheng Pang, Wenke Xia, Zhigang Wang +4

3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-bas…

cs.CV2022

Learning in Audio-visual Context: A Review, Analysis, and New Perspective

Yake Wei, Di Hu, Yapeng Tian +1

Sight and hearing are two senses that play a vital role in human communication and scene understanding. To mimic human perception ability, audio-visual learning, aimed at developin…

cs.LG2024

Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Qingyang Zhang, Yake Wei, Zongbo Han +8

Multimodal fusion focuses on integrating information from multiple modalities with the goal of more accurate prediction, which has achieved remarkable progress in a wide range of s…

cs.RO2024

Play to the Score: Stage-Guided Dynamic Multi-Sensory Fusion for Robotic Manipulation

Ruoxuan Feng, Di Hu, Wenke Ma +1

Humans possess a remarkable talent for flexibly alternating to different senses when interacting with the environment. Picture a chef skillfully gauging the timing of ingredient ad…

cs.LG2023

Supervised Knowledge May Hurt Novel Class Discovery Performance

Ziyun Li, Jona Otholt, Ben Dai +3

Novel class discovery (NCD) aims to infer novel categories in an unlabeled dataset by leveraging prior knowledge of a labeled set comprising disjoint but related classes. Given tha…

cs.CV2021

Unsupervised Multi-Source Domain Adaptation for Person Re-Identification

Zechen Bai, Zhigang Wang, Jian Wang +2

Unsupervised domain adaptation (UDA) methods for person re-identification (re-ID) aim at transferring re-ID knowledge from labeled source data to unlabeled target data. Although ac…

cs.CV2023

Robust Cross-Modal Knowledge Distillation for Unconstrained Videos

Wenke Xia, Xingjian Li, Andong Deng +3

Cross-modal distillation has been widely used to transfer knowledge across different modalities, enriching the representation of the target unimodal one. Recent studies highly rela…

cs.RO2024

Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs

Wenke Xia, Dong Wang, Xincheng Pang +4

Generalizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in si…

cond-mat.mes-hall2024

Two-dimensional 5d multiferroic W3Cl8: breathing Kagome lattice and tunable magneto-optical Kerr effect

Di Hu, Haoshen Ye, Ning Ding +4

Owing to the strong spin-orbit coupling and the related fascinating physical properties, heavy 5d transition-metals exhibit desirable application prospects. However, up to now, the…

cs.HC2025

Ambient Listening in Clinical Practice: Evaluating EPIC Signal Data Before and After Implementation and Its Impact on Physician Workload

Yawen Guo, Di Hu, Jiayuan Wang +4

The widespread adoption of EHRs following the HITECH Act has increased the clinician documentation burden, contributing to burnout. Emerging technologies, such as ambient listening…

cs.CV2019

Deep Multimodal Clustering for Unsupervised Audiovisual Learning

Di Hu, Feiping Nie, Xuelong Li

The seen birds twitter, the running cars accompany with noise, etc. These naturally audiovisual correspondences provide the possibilities to explore and understand the outside worl…

cs.RO2025

Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction

Wenke Xia, Ruoxuan Feng, Dong Wang +1

Building a generalizable self-correction system is crucial for robots to recover from failures. Despite advancements in Multimodal Large Language Models (MLLMs) that empower robots…

cs.RO2026

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

Wenke Xia, Pei Ren, Wenbo Yu +10

Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems…

cs.CV2019

Listen to the Image

Di Hu, Dong Wang, Xuelong Li +2

Visual-to-auditory sensory substitution devices can assist the blind in sensing the visual environment by translating the visual information into a sound pattern. To improve the tr…

cs.CV2026

Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos

Henghui Du, Chunjie Zhang, Xi Chen +2

Long Video Question-Answering (LVQA) presents a significant challenge for Multi-modal Large Language Models (MLLMs) due to immense context and overloaded information, which could a…

physics.plasm-ph2022

Hot-tail electrons' impact on assimilation and injection penetration of D2 Shattered Pellet Injections

Di Hu, Chang Liu

The fragment ablation rate plays significant roles in the mitigation efficiency of Shattered Pellet Injection (SPI) as a Disruption Mitigation System (DMS). Current mainstream 3D M…

physics.chem-ph2024

Highly dispersed Ru nanoparticles anchored on NiAl layered double oxides catalyst for selective hydrodeoxygenation of vanillin

Yongjian Zeng, Lu Lin, Di Hu +5

The hydrodeoxygenation (HDO) of lignin-derived feedstocks into value-added chemicals with high efficiency and selectivity is desirable for the utilization of biomass resource. The…

cond-mat.supr-con2014

Design and Performance Analysis of a 2.5 MW-Class HTS Synchronous Motor for Ship Propulsion

Jin Zou, Di Hu, Tim J. Flack +3

The development of cryogenic technology and high temperature superconducting (HTS) materials has seen continued interest worldwide in the development of HTS machines since the late…

cs.RO2026

GeCo-SRT: Geometry-aware Continual Adaptation for Robotic Cross-Task Sim-to-Real Transfer

Wenbo Yu, Wenke Xia, Weitao Zhang +1

Bridging the sim-to-real gap is important for applying low-cost simulation data to real-world robotic systems. However, previous methods are severely limited by treating each trans…

cs.CL2024

Towards Effective and Efficient Continual Pre-training of Large Language Models

Jie Chen, Zhipeng Chen, Jiapeng Wang +16

Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents…

cs.CL2023

TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real World

Hongpeng Lin, Ludan Ruan, Wenke Xia +8

To facilitate the research on intelligent and human-like chatbots with multi-modal context, we introduce a new video-based multi-modal dialogue dataset, called TikTalk. We collect…

cs.CV2022

Visual Sound Localization in the Wild by Cross-Modal Interference Erasing

Xian Liu, Rui Qian, Hang Zhou +5

The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios ar…

cs.LG2026

When Molecular Similarity Works: Property Cliffs Reveal Hidden Errors

Di Hu, Kun Li, Haojie Rao +6

Accurate prediction of molecular properties underpins drug discovery and material design, yet even state-of-the-art models remain vulnerable to localized failure modes that aggrega…

cs.CV2024

On-the-fly Modulation for Balanced Multimodal Learning

Yake Wei, Di Hu, Henghui Du +1

Multimodal learning is expected to boost model performance by integrating information from different modalities. However, its potential is not fully exploited because the widely-us…

cs.CV2021

Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation

Yapeng Tian, Di Hu, Chenliang Xu

There are rich synchronized audio and visual events in our daily life. Inside the events, audio scenes are associated with the corresponding visual objects; meanwhile, sounding obj…

cs.CV2018

Deep LDA Hashing

Di Hu, Feiping Nie, Xuelong Li

The conventional supervised hashing methods based on classification do not entirely meet the requirements of hashing technique, but Linear Discriminant Analysis (LDA) does. In this…

cs.CV2024

Enhancing multimodal cooperation via sample-level modality valuation

Yake Wei, Ruoxuan Feng, Zihe Wang +1

One primary topic of multimodal learning is to jointly incorporate heterogeneous information from different modalities. However most models often suffer from unsatisfactory multimo…