papers

Publications (92)

math.AP2025

Robin Problems of Elliptic Equations on Rough Domains: Hölder Regularity, Green's Functions, and Harmonic Measures

Jiayi Wang, Dachun Yang, Sibei Yang

Let and . Assume that is a one-sided bounded non-tangentially accessible domain with -Ahlfors regular boundary and is the su…

cs.CV2023

PCRLv2: A Unified Visual Information Preservation Framework for Self-supervised Pre-training in Medical Image Analysis

Hong-Yu Zhou, Chixiang Lu, Chaoqi Chen +2

Recent advances in self-supervised learning (SSL) in computer vision are primarily comparative, whose goal is to preserve invariant and discriminative semantics in latent represent…

cs.CV2026

Vision Transformers Need More Than Registers

Cheng Shi, Yizhou Yu, Sibei Yang

Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, artifacts in ViTs are widely observ…

cs.CR2025

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures

Yukai Zhou, Sibei Yang, Wenjie Wang

Large language models (LLMs) are increasingly deployed in real-world applications, raising concerns about their security. While jailbreak attacks highlight failures under overtly h…

cs.CV2023

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

Hanzhuo Huang, Yufan Feng, Cheng Shi +3

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text pr…

cs.CV2025

Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability

Yingdong Shi, Changming Li, Yifan Wang +5

Diffusion models have demonstrated impressive capabilities in synthesizing diverse content. However, despite their high-quality outputs, these models often perpetuate social biases…

cs.CV2024

Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation

Qiyuan Dai, Sibei Yang

Referring image segmentation (RIS) aims to precisely segment referents in images through corresponding natural language expressions, yet relying on cost-intensive mask annotations.…

cs.CV2025

MVTokenFlow: High-quality 4D Content Generation using Multiview Token Flow

Hanzhuo Huang, Yuan Liu, Ge Zheng +3

In this paper, we present MVTokenFlow for high-quality 4D content creation from monocular videos. Recent advancements in generative models such as video diffusion models and multiv…

math.AP2016

Gradient Estimates via Rearrangements for Solutions of Some Schrödinger Equations

Sibei Yang, Der-Chen Chang, Dachun Yang +1

In this article, by applying the well known method for dealing with -Laplace type elliptic boundary value problems, the authors establish a sharp estimate for the decreasing rea…

cs.CV2024

OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers

Han Liang, Jiacheng Bao, Ruichi Zhang +6

We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods often fail or produce implausible motions with unseen text inputs, which…

cs.CV2025

Vision Function Layer in Multimodal LLMs

Cheng Shi, Yizhou Yu, Sibei Yang

This study identifies that visual-related functional decoding is distributed across different decoder layers in Multimodal Large Language Models (MLLMs). Typically, each function,…

cs.CV2025

TransXNet: Learning Both Global and Local Dynamics with a Dual Dynamic Token Mixer for Visual Recognition

Meng Lou, Shu Zhang, Hong-Yu Zhou +3

Recent studies have integrated convolutions into transformers to introduce inductive bias and improve generalization performance. However, the static nature of conventional convolu…

math.NT2026

Lattice point counting in Cygan--Korányi balls on Heisenberg groups

Sheng-Chen Mao, Sibei Yang

Lattice point counting in gauge balls on the Heisenberg group is a non-commutative analogue of the Euclidean multidimensional sphere problem, initiated by Garg, Nevo…

math.AP2026

Sobolev--Morrey Spaces and Divergence-Form Degenerate Second-Order Elliptic Equations on Domains with Higher Co-Dimensional Boundaries

Weiyi Kong, Yoshihiro Sawano, Dachun Yang +2

In this article, we study the weighted homogeneous Sobolev--Morrey spaces on domains in with higher co-dimensional boundaries. Precisely, we systematically establish…

math.AP2024

Some remarks on Riesz transform on exterior Lipschitz domains

Renjin Jiang, Sibei Yang

Let and be an elliptic operator on . Given an exterior Lipschitz domain , let be the elliptic op…

cs.CV2024

The devil is in the object boundary: towards annotation-free instance segmentation using Foundation Models

Cheng Shi, Sibei Yang

Foundation models, pre-trained on a large amount of data have demonstrated impressive zero-shot capabilities in various downstream tasks. However, in object detection and instance…

cs.CV2025

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Jiajin Tang, Zhengxuan Wei, Ge Zheng +1

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process…

cs.CV2024

A Survey on Graph Neural Networks and Graph Transformers in Computer Vision: A Task-Oriented Perspective

Chaoqi Chen, Yushuang Wu, Qiyuan Dai +5

Graph Neural Networks (GNNs) have gained momentum in graph representation learning and boosted the state of the art in a variety of areas, such as data mining (\emph{e.g.,} social…

cs.CV2026

Chart Deep Research in LVLMs via Parallel Relative Policy Optimization

Jiajin Tang, Gaoyang, Wenjie Wang +2

With the rapid advancement of data science, charts have evolved from simple numerical presentation tools to essential instruments for insight discovery and decision-making support.…

cs.CV2023

DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

Ge Zheng, Bin Yang, Jiajin Tang +2

A long-standing goal of AI systems is to perform complex multimodal reasoning like humans. Recently, large language models (LLMs) have made remarkable strides in such multi-step re…

math.CA2018

Atomic and Maximal Function Characterizations of Musielak-Orlicz-Hardy Spaces Associated to Non-negative Self-adjoint Operators on Spaces of Homogeneous Type

Sibei Yang, Dachun Yang

Let be a metric space with doubling measure and a non-negative self-adjoint operator on whose heat kernels satisfy the Gaussian upper bound est…

cs.CV2025

VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction

Zijian He, Yuwei Ning, Yipeng Qin +4

Virtual Try-On (VTON) is a transformative technology in e-commerce and fashion design, enabling realistic digital visualization of clothing on individuals. In this work, we propose…

cs.CV2022

Preservational Learning Improves Self-supervised Medical Image Models by Reconstructing Diverse Contexts

Hong-Yu Zhou, Chixiang Lu, Sibei Yang +2

Preserving maximal information is one of principles of designing self-supervised learning methodologies. To reach this goal, contrastive learning adopts an implicit way which is co…

math.CA2014

Riesz Transform Characterizations of Musielak-Orlicz-Hardy Spaces

Jun Cao, Der-Chen Chang, Dachun Yang +1

Let be a Musielak-Orlicz function satisfying that, for any , belongs to the Muckenhoupt weight class $A_\infty (\math…

math.AP2025

Equivalent Characterizations and Their Applications of Solvability of Poisson--Robin(-Regularity) Problems on Rough Domains

Xuelian Fu, Dachun Yang, Sibei Yang

Let , be a bounded one-sided chord arc domain, and . In this article, we study the (weak) Poisson--Robin(-regularity) problem f…

cs.CV2026

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization

Yu Pan, Andi Zhang, Yi Wang +2

Diffusion Vision-Language Models (dVLMs), built upon the non-causal foundations of Diffusion Large Language Models (dLLMs), have demonstrated remarkable efficacy in multimodal task…

math.FA2023

Hardy Spaces Associated with Non-Negative Self-Adjoint Operators and Ball Quasi-Banach Function Spaces on Doubling Metric Measure Spaces and Their Applications

Xiaosheng Lin, Dachun Yang, Sibei Yang +1

Let be a doubling metric measure space in the sense of R. R. Coifman and G. Weiss, a non-negative self-adjoint operator on satisfying th…

cs.CV2025

Augmenting Moment Retrieval: Zero-Dependency Two-Stage Learning

Zhengxuan Wei, Jiajin Tang, Sibei Yang

Existing Moment Retrieval methods face three critical bottlenecks: (1) data scarcity forces models into shallow keyword-feature associations; (2) boundary ambiguity in transition r…

cs.CV2024

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

Zhenxiang Lin, Xidong Peng, Peishan Cong +6

We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images…

cs.CV2023

Temporal Collection and Distribution for Referring Video Object Segmentation

Jiajin Tang, Ge Zheng, Sibei Yang

Referring video object segmentation aims to segment a referent throughout a video sequence according to a natural language expression. It requires aligning the natural language exp…

math.CA2014

Musielak-Orlicz BMO-Type Spaces Associated with Generalized Approximations to the Identity

Shaoxiong Hou, Dachun Yang, Sibei Yang

Let be a space of homogenous type and a growth function such that is a Muckenhoupt weight uniformly in…

math.CA2019

Applications of Hardy Spaces Associated with Ball Quasi-Banach Function Spaces

Fan Wang, Dachun Yang, Sibei Yang

Let be a ball quasi-Banach function space satisfying some minor assumptions. In this article, the authors establish the characterizations of , the Hardy spac…

math.AP2026

A Counterexample to Kenig's Interpolation Problem for Sobolev Spaces with Zero Boundary Conditions

Xiaosheng Lin, Dachun Yang, Sibei Yang +2

Let . In this article, we show that there exists a bounded domain such that, for any given $s\in(1,2)\setminus\{\frac32\…

cs.CV2026

RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation

Hanzhuo Huang, Qingyang Bao, Zekai Gu +4

In this paper, we propose a 3D asset-referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference-based ima…

cs.CV2018

Multi-Evidence Filtering and Fusion for Multi-Label Classification, Object Detection and Semantic Segmentation Based on Weakly Supervised Learning

Weifeng Ge, Sibei Yang, Yizhou Yu

Supervised object detection and semantic segmentation require object or even pixel level annotations. When there exist image level labels only, it is challenging for weakly supervi…

cs.CV2020

Graph-Structured Referring Expression Reasoning in The Wild

Sibei Yang, Guanbin Li, Yizhou Yu

Grounding referring expressions aims to locate in an image an object referred to by a natural language expression. The linguistic structure of a referring expression provides a lay…

cs.CV2025

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

Ruifei Zhang, Wei Zhang, Xiao Tan +4

Recent advancements in language-grounded autonomous driving have been significantly promoted by the sophisticated cognition and reasoning capabilities of large language models (LLM…

math.CA2025

Heat kernel estimates, fractional Riesz transforms and applications on exterior domains

Renjin Jiang, Tianjun Shen, Sibei Yang +1

In this paper, we derive sharp two side heat kernel estimate on exterior domains in the plane, and sharp upper heat kernel bound on exterior domains…

cs.CV2026

Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats

Jiaye Qian, Ge Zheng, Yuchen Zhu +1

Despite their impressive performance across a wide range of tasks, Large Vision-Language Models (LVLMs) remain prone to hallucination. In this study, we propose a comprehensive int…

cs.CV2025

Rethinking Query-based Transformer for Continual Image Segmentation

Yuchen Zhu, Cheng Shi, Dingyou Wang +5

Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-…

math.CA2013

Musielak-Orlicz-Hardy Spaces Associated with Operators Satisfying Reinforced Off-Diagonal Estimates

The Anh Bui, Jun Cao, Luong Dang Ky +2

Let be a metric space with doubling measure and a one-to-one operator of type having a bounded -functional calculus in satisfyin…

cs.CV2025

No More Sibling Rivalry: Debiasing Human-Object Interaction Detection

Bin Yang, Yulin Zhang, Hong-Yu Zhou +1

Detection transformers have been applied to human-object interaction (HOI) detection, enhancing the localization and recognition of human-action-object triplets in images. Despite…

cs.LG2014

Introduction to Clustering Algorithms and Applications

Sibei Yang, Liangde Tao, Bingchen Gong

Data clustering is the process of identifying natural groupings or clusters within multidimensional data based on some similarity measure. Clustering is a fundamental process in ma…

cs.CV2026

Self-Prophetic Decoding to Unlock Visual Search in LVLMs

Zhendong He, Qiyuan Dai, Guanbin Li +2

Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of the thinking-with-images par…

cs.CL2025

Auto-Search and Refinement: An Automated Framework for Gender Bias Mitigation in Large Language Models

Yue Xu, Chengyan Fu, Li Xiong +2

Pre-training large language models (LLMs) on vast text corpora enhances natural language processing capabilities but risks encoding social biases, particularly gender bias. While p…

cs.CV2025

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

Chunlin Yu, Hanqing Wang, Ye Shi +4

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single…

cs.CR2025

DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing

Yi Wang, Fenghua Weng, Sibei Yang +3

Large Language Models (LLMs) are widely applied in decision making, but their deployment is threatened by jailbreak attacks, where adversarial users manipulate model behavior to by…

cs.CV2019

Dynamic Graph Attention for Referring Expression Comprehension

Sibei Yang, Guanbin Li, Yizhou Yu

Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently re…

cs.CV2025

Penalizing Boundary Activation for Object Completeness in Diffusion Models

Haoyang Xu, Tianhao Zhao, Sibei Yang +1

Diffusion models have emerged as a powerful technique for text-to-image (T2I) generation, creating high-quality, diverse images across various domains. However, a common limitation…

math.AP2021

A Two-Weight Boundedness Criterion and Its Applications

Sibei Yang, Zhenyu Yang

In this article, the authors establish a general (two-weight) boundedness criterion for a pair of functions, , on in the scale of weighted Lebesgue spaces, we…

cs.CV2023

EdaDet: Open-Vocabulary Object Detection Using Early Dense Alignment

Cheng Shi, Sibei Yang

Vision-language models such as CLIP have boosted the performance of open-vocabulary object detection, where the detector is trained on base categories but required to detect novel…

math.AP2022

Global Gradient Estimates for Dirichlet Problems of Elliptic Operators with a BMO Anti-Symmetric Part

Sibei Yang, Dachun Yang, Wen Yuan

Let and be a bounded NTA domain. In this article, the authors investigate (weighted) global gradient estimates for Dirichlet boundary value problems…

math.CA2012

Endpoint Boundedness of Riesz Transforms on Hardy Spaces Associated with Operators

Jun Cao, Dachun Yang, Sibei Yang

Let be a nonnegative self-adjoint operator in satisfying the Davies-Gaffney estimates and a second order divergence form elliptic operator with com…

math.AP2026

A Counterexample to the Necessity of the Vanishing Carleson Condition for VMO Poisson Kernels

Xiaosheng Lin, Dachun Yang, Sibei Yang +2

In [Problem 3.2.23, CBMS Regional Conference Series in Mathematics 83, 1994], Kenig asked whether the vanishing Carleson condition is the necessary and sufficient for the logarithm…

math.CA2018

New Characterizations of Musielak-Orlicz-Sobolev Spaces via Sharp Ball Averaging Functions

Sibei Yang, Dachun Yang, Wen Yuan

In this article, the authors establish a new characterization of the Musielak--Orlicz--Sobolev space on , which includes the classical Orlicz--Sobolev space, the weig…

math.AP2024

Global BMO-Sobolev Estimates for Second-Order Linear Elliptic Equations on Lipschitz Domains

Hongjie Dong, Dachun Yang, Sibei Yang

Let and be a bounded Lipschitz domain. In this article, we establish first-order global regularity estimates in the scale of BMO spaces on f…

cs.CV2023

LoGoPrompt: Synthetic Text Images Can Be Good Visual Prompts for Vision-Language Models

Cheng Shi, Sibei Yang

Prompt engineering is a powerful tool used to enhance the performance of pre-trained models on downstream tasks. For example, providing the prompt "Let's think step by step" improv…

cs.CV2023

Grounded Image Text Matching with Mismatched Relation Reasoning

Yu Wu, Yana Wei, Haozhe Wang +3

This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities o…

math.AP2020

Weighted Global Regularity Estimates for Elliptic Problems with Robin Boundary Conditions in Lipschitz Domains

Sibei Yang, Dachun Yang, Wen Yuan

Let and be a bounded Lipschitz domain in . In this article, the authors investigate global (weighted) estimates for the gradient of solutions to Robin bo…

cs.CV2019

Non-Local Context Encoder: Robust Biomedical Image Segmentation against Adversarial Attacks

Xiang He, Sibei Yang, Guanbin Li? +3

Recent progress in biomedical image segmentation based on deep convolutional neural networks (CNNs) has drawn much attention. However, its vulnerability towards adversarial samples…

cs.CV2025

Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement

Qiyuan Dai, Hanzhuo Huang, Yu Wu +1

Generalized Category Discovery (GCD) aims to recognize unlabeled images from known and novel classes by distinguishing novel classes from known ones, while also transferring knowle…

math.FA2023

Predual Spaces of Hardy Spaces Related to Fractional Schrödinger Operators

Qiumeng Li, Haibo Lin, Sibei Yang

Let and . For any , the fractional Schrödinger operator is defined by \begin{equation*} L_α:=(-Δ)^{α/2}+a{|x|…

cs.CV2025

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

Jiajin Tang, Zhengxuan Wei, Yuchen Zhu +4

Temporal sentence grounding aims to identify exact moments in a video that correspond to a given textual query, typically addressed with detection transformer (DETR) solutions. How…

math.CA2016

Maximal Function Characterizations of Musielak-Orlicz-Hardy Spaces Associated to Non-negative Self-adjoint Operators Satisfying Gaussian Estimates

Dachun Yang, Sibei Yang

Let be a non-negative self-adjoint operator on whose heat kernels have the Gaussian upper bound estimates. Assume that the growth function $φ:\,\mathbb{R}^…

cs.CV2025

Discovering Influential Neuron Path in Vision Transformers

Yifan Wang, Yifei Liu, Yingdong Shi +5

Vision Transformer models exhibit immense power yet remain opaque to human understanding, posing challenges and risks for practical applications. While prior research has attempted…

math.AP2022

Heat Kernels and Hardy Spaces on Non-Tangentially Accessible Domains with Applications to Global Regularity of Inhomogeneous Dirichlet Problems

Sibei Yang, Dachun Yang

Let and be a bounded non-tangentially accessible domain (for short, NTA domain) of . Assume that is a second-order divergence form elliptic operato…

cs.CV2021

ConvNets vs. Transformers: Whose Visual Representations are More Transferable?

Hong-Yu Zhou, Chixiang Lu, Sibei Yang +1

Vision transformers have attracted much attention from computer vision researchers as they are not restricted to the spatial inductive bias of ConvNets. However, although Transform…

cs.CV2020

Relationship-Embedded Representation Learning for Grounding Referring Expressions

Sibei Yang, Guanbin Li, Yizhou Yu

Grounding referring expressions in images aims to locate the object instance in an image described by a referring expression. It involves a joint understanding of natural language…

math.CA2011

Real-variable Characterizations of Orlicz-Hardy Spaces on Strongly Lipschitz Domains of

Dachun Yang, Sibei Yang

Let be a strongly Lipschitz domain of , whose complement in is unbounded. Let be a second order divergence form elliptic operator on $L^2 (Ω)…

math.CA2012

Musielak-Orlicz Hardy Spaces Associated with Operators and Their Applications

Dachun Yang, Sibei Yang

Let be a metric space with doubling measure and a nonnegative self-adjoint operator in satisfying the Davies-Gaffney estimates. Let $φ:\,\math…

cs.CV2023

Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embodied Reference Understanding

Cheng Shi, Sibei Yang

Embodied Reference Understanding studies the reference understanding in an embodied fashion, where a receiver is required to locate a target object referred to by both language and…

math.AP2019

estimate for the gradient in

Duchao Liu, Beibei Wang, Sibei Yang

Under appropriate assumptions on the -fucntion, the estimate for the gradient of the minimizers of a class of energy functional in Musielak-Orlicz-Sobol…

cs.CV2025

Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video

Yulin Zhang, Cheng Shi, Yang Wang +1

Envision an AI capable of functioning in human-like settings, moving beyond mere observation to actively understand, anticipate, and proactively respond to unfolding events. Toward…

cs.CL2026

Bridging the Agent-World Gap: Text World Models for LLM-based Agents

Yixia Li, Hongru Wang, Peng Lai +13

Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet m…

cs.CV2026

WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs

Yulin Zhang, Cheng Shi, Sibei Yang

Recent advances in Multimodal Large Language Models have greatly improved visual understanding and reasoning, yet their quadratic attention and offline training protocols make them…

cs.CV2023

CoTDet: Affordance Knowledge Prompting for Task Driven Object Detection

Jiajin Tang, Ge Zheng, Jingyi Yu +1

Task driven object detection aims to detect object instances suitable for affording a task in an image. Its challenge lies in object categories available for the task being too div…

math.CA2012

Weighted Local Orlicz-Hardy Spaces on Domains and Their Applications in Inhomogeneous Dirichlet and Neumann Problems

Jun Cao, Der-Chen Chang, Dachun Yang +1

Let be either or a strongly Lipschitz domain of , and (the class of Muckenhoupt weights). Let be a second ord…

math.AP2016

Global Boundedness of the Gradient for a Class of Schrödinger Equations

Sibei Yang

In this paper, via applying the method developed by A. Cianchi and V. Maz'ya, the author obtains the global boundedness of the gradient for solutions to Dirichlet and Neumann probl…

cs.CV2023

Contrastive Grouping with Transformer for Referring Image Segmentation

Jiajin Tang, Ge Zheng, Cheng Shi +1

Referring image segmentation aims to segment the target referent in an image conditioning on a natural language expression. Existing one-stage methods employ per-pixel classificati…

cs.CV2025

Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context

Ge Zheng, Jiaye Qian, Jiajin Tang +1

Large Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-for…

cs.GR2023

DreamFace: Progressive Generation of Animatable 3D Faces under Text Guidance

Longwen Zhang, Qiwei Qiu, Hongyang Lin +7

Emerging Metaverse applications demand accessible, accurate, and easy-to-use tools for 3D digital human creations in order to depict different cultures and societies as if in the p…

cs.CV2026

MAC: A Benchmark for Multiple Attributes Compositional Zero-Shot Learning

Shuo Xu, Sai Wang, Xinyue Hu +3

Compositional Zero-Shot Learning (CZSL) aims to learn semantic primitives (attributes and objects) from seen compositions and recognize unseen attribute-object compositions. Existi…

cs.CV2025

Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM

Qiyuan Dai, Sibei Yang

Vision-Language Models (VLMs) have become prominent in open-world image recognition for their strong generalization abilities. Yet, their effectiveness in practical applications is…

math.CA2014

Lusin Area Function and Molecular Characterizations of Musielak-Orlicz Hardy Spaces and Their Applications

Shaoxiong Hou, Dachun Yang, Sibei Yang

Lusin Area Function and Molecular Characterizations of Musielak-Orlicz Hardy Spaces and Their ApplicationsLet be a growth function s…

cs.CR2026

EVA: Editing for Versatile Alignment against Jailbreaks

Yi Wang, Hongye Qiu, Yue Xu +4

Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks, where adversaries exploit te…

cs.CV2024

Part2Object: Hierarchical Unsupervised 3D Instance Segmentation

Cheng Shi, Yulin Zhang, Bin Yang +3

Unsupervised 3D instance segmentation aims to segment objects from a 3D point cloud without any annotations. Existing methods face the challenge of either too loose or too tight cl…

cs.CV2024

Plain-Det: A Plain Multi-Dataset Object Detector

Cheng Shi, Yuchen Zhu, Sibei Yang

Recent advancements in large-scale foundational models have sparked widespread interest in training highly proficient large vision models. A common consensus revolves around the ne…

math.CA2011

Weighted Local Orlicz-Hardy Spaces with Applications to Pseudo-differential Operators

Dachun Yang, Sibei Yang

Let be a concave function on of strictly lower type and . We introduce the weighted local Orl…

math.FA2022

Maximal Function and Riesz Transform Characterizations of Hardy Spaces Associated with Homogeneous Higher Order Elliptic Operators and Ball Quasi-Banach Function Spaces

Xiaosheng Lin, Dachun Yang, Sibei Yang +1

Let be a homogeneous divergence form higher order elliptic operator with complex bounded measurable coefficients on and a ball quasi-Banach function space on…

math.CA2012

Orlicz-Hardy Spaces Associated with Divergence Operators on Unbounded Strongly Lipschitz Domains of

Dachun Yang, Sibei Yang

Let be either or an unbounded strongly Lipschitz domain of , and be a continuous, strictly increasing, subadditive and positive function on $…

math.CA2012

Local Hardy Spaces of Musielak-Orlicz Type and Their Applications

Dachun Yang, Sibei Yang

Let $ϕ: \mathbb{R}^n\times[0,\fz)\rightarrow[0,\fz)$ be a function such that is an Orlicz function and $ϕ(\cdot,t)\in A^{\mathop\mathrm{loc}}_{\infty}(\mathbb{R}^n)…

cs.RO2024

RealDex: Towards Human-like Grasping for Robotic Dexterous Hand

Yumeng Liu, Yaxun Yang, Youzhuo Wang +9

In this paper, we introduce RealDex, a pioneering dataset capturing authentic dexterous hand grasping motions infused with human behavioral patterns, enriched by multi-view and mul…