papers

Publications (11)

cs.CV2023

Artificial-Spiking Hierarchical Networks for Vision-Language Representation Learning

Yeming Chen, Siyu Zhang, Yaoru Sun +2

With the success of self-supervised learning, multimodal foundation models have rapidly adapted a wide range of downstream tasks driven by vision and language (VL) pretraining. Sta…

cs.LG2024

Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout

Haoran Wang, Zeshen Tang, Leya Yang +4

Goal-conditioned hierarchical reinforcement learning (HRL) presents a promising approach for enabling effective exploration in complex, long-horizon reinforcement learning (RL) tas…

cs.CV2023

Pixel Difference Convolutional Network for RGB-D Semantic Segmentation

Jun Yang, Lizhi Bai, Yaoru Sun +3

RGB-D semantic segmentation can be advanced with convolutional neural networks due to the availability of Depth data. Although objects cannot be easily discriminated by just the 2D…

cs.CL2026

TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity

Zheyuan Yang, Liqiang Shang, Junjie Chen +6

We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,0…

cs.LG2025

HG2P: Hippocampus-inspired High-reward Graph and Model-Free Q-Gradient Penalty for Path Planning and Motion Control

Haoran Wang, Yaoru Sun, Zeshen Tang +2

Goal-conditioned hierarchical reinforcement learning (HRL) decomposes complex reaching tasks into a sequence of simple subgoal-conditioned tasks, showing significant promise for ad…

cs.CV2025

Compress image to patches for Vision Transformer

Xinfeng Zhao, Yaoru Sun

The Vision Transformer (ViT) has made significant strides in the field of computer vision. However, as the depth of the model and the resolution of the input images increase, the c…

cs.CV2023

LOIS: Looking Out of Instance Semantics for Visual Question Answering

Siyu Zhang, Yeming Chen, Yaoru Sun +3

Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts h…

cs.CV2024

Superpixel Semantics Representation and Pre-training for Vision-Language Task

Siyu Zhang, Yeming Chen, Yaoru Sun +4

The key to integrating visual language tasks is to establish a good alignment strategy. Recently, visual semantic representation has achieved fine-grained visual understanding by d…

eess.IV2022

DCANet: Differential Convolution Attention Network for RGB-D Semantic Segmentation

Lizhi Bai, Jun Yang, Chunqi Tian +4

Combining RGB images and the corresponding depth maps in semantic segmentation proves the effectiveness in the past few years. Existing RGB-D modal fusion methods either lack the n…

cs.CL2023

Task-oriented Memory-efficient Pruning-Adapter

Guorun Wang, Jun Yang, Yaoru Sun

The Outstanding performance and growing size of Large Language Models has led to increased attention in parameter efficient learning. The two predominant approaches are Adapters an…

physics.soc-ph2006

Structure of Peer-to-Peer Social Networks

Fang Wang, Yamir Moreno, Yaoru Sun

This paper presents a statistical analysis of the structure of Peer-to-Peer (P2P) social networks that captures social associations of distributed peers in resource sharing. Peer s…