Publications (142)
Borel Local Lemma: arbitrary random variables and limited exponential growth
Anton Bernshteyn, Jing Yu
The Lovász Local Lemma (the LLL for short) is a powerful tool in probabilistic combinatorics that is used to verify the existence of combinatorial objects with desirable propertie…
Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering
Zhuoqian Yang, Zengchang Qin, Jing Yu +1
One of the key issues of Visual Question Answering (VQA) is to reason with semantic clues in the visual content under the guidance of the question, how to model relational semantic…
Enhancing atomic-resolution in electron microscopy: A frequency-domain deep learning denoiser
Ivan Pinto-Huguet, Marc Botifoll, Xuli Chen +6
Atomic resolution electron microscopy, particularly high-angle annular dark-field scanning transmission electron microscopy, has become an essential tool for many scientific fields…
Syntax-BERT: Improving Pre-trained Transformers with Syntax Trees
Jiangang Bai, Yujing Wang, Yiren Chen +4
Pre-trained language models like BERT achieve superior performances in various NLP tasks without explicit consideration of syntactic information. Meanwhile, syntactic information h…
Derived discrete Hopf algebras with the Chevalley property
Jing Yu, Gongxiang Liu
We try to classify Hopf algebras with the Chevalley property according to their derived representation type. We show that a finite-dimensional indecomposable non-semisimple Hopf al…
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval
Yuanmin Tang, Xiaoting Qin, Jue Zhang +7
Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user…
A Novel Symmetry Constraint Of The Super cKdV System
Jing Yu, Jingsong He, Yi Cheng +1
A new (1+1)-dimensional integrable system, i. e. the super coupled Korteweg-de Vries (cKdV) system, has been constructed by a super extension of the well-known (1+1)-dimensional cK…
SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
Jing Yu, Yuqi Tang, Kehua Feng +8
Large Language Models (LLMs) have shown impressive capabilities in contextual understanding and reasoning. However, evaluating their performance across diverse scientific domains r…
DARWIN: A Highly Flexible Platform for Imaging Research in Radiology
Lufan Chang, Wenjing Zhuang, Richeng Wu +6
To conduct a radiomics or deep learning research experiment, the radiologists or physicians need to grasp the needed programming skills, which, however, could be frustrating and co…
Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning
Xiangyan Qu, Jing Yu, Keke Gai +5
Recent work shows that documents from encyclopedias serve as helpful auxiliary information for zero-shot learning. Existing methods align the entire semantics of a document with co…
LLAMA: Multi-Feedback Smart Contract Fuzzing Framework with LLM-Guided Seed Generation
Keke Gai, Haochen Liang, Jing Yu +2
Smart contracts play a pivotal role in blockchain ecosystems, and fuzzing remains an important approach to securing smart contracts. Even though mutation scheduling is a key factor…
Submodular flows and extreme flows on measurable spaces
Jing Yu, Junchi Zhang, Mingyang Zhou
The theory of submodular flows, introduced by Edmonds and Giles, is a cornerstone of combinatorial optimization, unifying network flows, matroid intersections and directed cut cove…
Edge-Based Blur Kernel Estimation Using Sparse Representation and Self-Similarity
Jing Yu, Zhenchun Chang, Chuangbai Xiao
Blind image deconvolution is the problem of recovering the latent image from the only observed blurry image when the blur kernel is unknown. In this paper, we propose an edge-based…
PGN: The RNN's New Successor is Effective for Long-Range Time Series Forecasting
Yuxin Jia, Youfang Lin, Jing Yu +3
Due to the recurrent structure of RNN, the long information propagation path poses limitations in capturing long-term dependencies, gradient explosion/vanishing issues, and ineffic…
Beyond Chunk-Local Extraction: Cross-Chunk Graph Augmentation for GraphRAG
Jiaming Zhang, Yibo Zhao, Jing Yu +2
GraphRAG extends retrieval-augmented generation by organizing corpora as explicit knowledge graphs, enabling graph-based retrieval for complex question answering. However, existing…
SLVC-DIDA: Signature-less Verifiable Credential-based Issuer-hiding and Multi-party Authentication for Decentralized Identity
Tianxiu Xie, Keke Gai, Jing Yu +2
As an emerging paradigm in digital identity, Decentralized Identity (DID) appears advantages over traditional identity management methods in a variety of aspects, e.g., enhancing u…
Coquasitriangular structures on Hopf algebras constructed via abelian extensions
Jing Yu, Xiangjun Zhen
The aim of this paper is to study coquasitriangular structures on a class of cosemisimple Hopf algebras of the form , constructed as abelian extensions…
A Comprehensive Survey for Evaluation Methodologies of AI-Generated Music
Zeyu Xiong, Weitao Wang, Jing Yu +2
In recent years, AI-generated music has made significant progress, with several models performing well in multimodal and complex musical genres and scenes. While objective metrics…
Algebraic independence of arithmetic gamma values and Carlitz zeta values
Chieh-Yu Chang, Matthew A. Papanikolas, Dinesh S. Thakur +1
We consider the values at proper fractions of the arithmetic gamma function and the values at positive integers of the zeta function for F_q[theta] and provide complete algebraic i…
RecGPT Technical Report
Chao Yi, Dian Chen, Gaoyang Guo +51
Recommender systems are among the most impactful applications of artificial intelligence, serving as critical infrastructure connecting users, merchants, and platforms. However, mo…
A2-DIDM: Privacy-preserving Accumulator-enabled Auditing for Distributed Identity of DNN Model
Tianxiu Xie, Keke Gai, Jing Yu +1
Recent booming development of Generative Artificial Intelligence (GenAI) has facilitated model commercialization to reinforce the model performance, including licensing or trading…
CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs
Zhen Zeng, Leijiang Gu, Feng Li +2
Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned responses in cross-cultural settings…
DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction Identification
Junzhou Chen, Zirui Zhang, Jing Yu +5
Driver distraction remains a leading cause of traffic accidents, posing a critical threat to road safety globally. As intelligent transportation systems evolve, accurate and real-t…
An effective criterion for Eulerian multizeta values in positive characteristic
Chieh-Yu Chang, Matthew A. Papanikolas, Jing Yu
Characteristic p multizeta values were initially studied by Thakur, who defined them as analogues of classical multiple zeta values of Euler. In the present paper we establish an e…
High-Performance Fine Defect Detection in Artificial Leather Using Dual Feature Pool Object Detection
Lin Huang, Weisheng Li, Yujuan Tan +2
In this study, the structural problems of the YOLOv5 model were analyzed emphatically. Based on the characteristics of fine defects in artificial leather, four innovative structure…
ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic Classification
Xinjie Lin, Gang Xiong, Gaopeng Gou +3
Encrypted traffic classification requires discriminative and robust traffic representation captured from content-invisible and imbalanced traffic data for accurate classification,…
Counting degree-constrained orientations
Jing Yu, Jie-Xiang Zhu
We study the enumeration of graph orientations under local degree constraints. Given a finite graph and a family of admissible sets $\{\mathsf P_v \subseteq \mathbb{Z}…
ALIEN: Analytic Latent Watermarking for Controllable Generation
Liangqi Lei, Keke Gai, Jing Yu +1
Watermarking is a technical alternative to safeguarding intellectual property and reducing misuse. Existing methods focus on optimizing watermarked latent variables to balance wate…
Trajectory-Aware Information Matching for Multi-Step Gradient Inversion in Federated Learning
Li Xia, Jing Yu, Zheng Liu +3
Federated learning enables distributed information sharing and collaborative model training without exposing raw client data. However, shared gradients or model updates may still c…
PCDiff: Proactive Control for Ownership Protection in Diffusion Models with Watermark Compatibility
Keke Gai, Ziyue Shen, Jing Yu +2
With the growing demand for protecting the intellectual property (IP) of text-to-image diffusion models, we propose PCDiff -- a proactive access control framework that redefines mo…
Hypergraph independence bounds: from maximum degree to average degree
Jing Yu, Junchi Zhang
We prove a transfer theorem for hereditary classes of -uniform hypergraphs. Let be such a class, and for write and for the maxim…
One-photon Solutions to Multiqubit Multimode quantum Rabi model
Jie Peng, Juncong Zheng, Jing Yu +6
General solutions to the quantum Rabi model involve subspaces with unbounded number of photons. However, for the multiqubit multimode case, we find special solutions with at most o…
System-Level Performance and Communication Tradeoff in Networked Control with Predictions
Yifei Wu, Jing Yu, Tongxin Li
Distributed control of large-scale systems is challenging due to the need for scalable and localized communication and computation. In this work, we introduce a Predictive System-L…
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
Yongrui Chen, Zhiqiang Liu, Jing Yu +21
Large Language Models (LLMs) have demonstrated substantial progress on reasoning tasks involving unstructured text, yet their capabilities significantly deteriorate when reasoning…
Geometric Gamma values and zeta values in positive characteristic
Chieh-Yu Chang, Matthew A. Papanikolas, Jing Yu
In analogy with values of the classical Euler Gamma-function at rational numbers and the Riemann zeta-function at positive integers, we consider Thakur's geometric Gamma-function e…
Macroscopic Quantum Tunneling Effect of Z2 Topological Order
Jing Yu, Su-Peng Kou
In this paper, macroscopic quantum tunneling (MQT) effect of Z2 topological order in the Wen-Plaquette model is studied. This kind of MQT is characterized by quantum tunneling proc…
The Bargmann symmetry constraint and binary nonlinearization of the super Dirac systems
Jing Yu, Jingsong He, Wen-Xiu Ma +1
An explicit Bargmann symmetry constraint is computed and its associated binary nonlinearization of Lax pairs is carried out for the super Dirac systems. Under the obtained symmetry…
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
Yulin Song, Guorui Sang, Jing Yu +1
Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have perfo…
YOLOCS: Object Detection based on Dense Channel Compression for Feature Spatial Solidification
Lin Huang, Weisheng Li, Yujuan Tan +3
In this study, we examine the associations between channel features and convolutional kernels during the processes of feature purification and gradient backpropagation, with a focu…
Learning the Uncertainty Sets for Control Dynamics via Set Membership: A Non-Asymptotic Analysis
Yingying Li, Jing Yu, Lauren Conger +2
This paper studies uncertainty set estimation for unknown linear systems. Uncertainty sets are crucial for the quality of robust control since they directly influence the conservat…
A Vision-Language Pre-training Model-Guided Approach for Mitigating Backdoor Attacks in Federated Learning
Keke Gai, Dongjue Wang, Jing Yu +2
Defending backdoor attacks in Federated Learning (FL) under heterogeneous client data distributions encounters limitations balancing effectiveness and privacy-preserving, while mos…
Outerplanar graphs with positive Lin-Lu-Yau curvature
George Brooks, Fadekemi Osaye, Anna Schenfisch +2
In this paper, we show that all simple outerplanar graphs with minimum degree at least and positive Lin-Lu-Yau Ricci curvature on every edge have maximum degree at most …
Watermarking Vision-Language Pre-trained Models for Multi-modal Embedding as a Service
Yuanmin Tang, Jing Yu, Keke Gai +4
Recent advances in vision-language pre-trained models (VLPs) have significantly increased visual understanding and cross-modal analysis capabilities. Companies have emerged to prov…
MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification
Xiangyan Qu, Jing Yu, Jiamin Zhuang +3
Zero-shot learning (ZSL) aims to train a model on seen classes and recognize unseen classes by knowledge transfer through shared auxiliary information. Recent studies reveal that d…
T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
Yili Li, Jing Yu, Keke Gai +3
Current text-video retrieval methods mainly rely on cross-modal matching between queries and videos to calculate their similarity scores, which are then sorted to obtain retrieval…
Robust Online Voltage Control with an Unknown Grid Topology
Christopher Yeh, Jing Yu, Yuanyuan Shi +1
Voltage control generally requires accurate information about the grid's topology in order to guarantee network stability. However, accurate topology identification is a challengin…
WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages
Jia Yu, Fei Yuan, Rui Min +20
This paper introduces the open-source dataset WanJuanSiLu, designed to provide high-quality training corpora for low-resource languages, thereby advancing the research and developm…
MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering
Yang Ding, Jing Yu, Bang Liu +3
Knowledge-based visual question answering requires the ability of associating external knowledge for open-ended cross-modal scene understanding. One limitation of existing solution…
Modeling Text with Graph Convolutional Network for Cross-Modal Information Retrieval
Jing Yu, Yuhang Lu, Zengchang Qin +4
Cross-modal information retrieval aims to find heterogeneous data of various modalities from a given query of one modality. The main challenge is to map different modalities into a…
MRI-based and metabolomics-based age scores act synergetically for mortality prediction shown by multi-cohort federated learning
Pedro Mateus, Swier Garst, Jing Yu +19
Biological age scores are an emerging tool to characterize aging by estimating chronological age based on physiological biomarkers. Various scores have shown associations with agin…
A class of (infinite-dimensional) cosemisimple Hopf algebras constructed via abelian extensions
Jing Yu, Gongxiang Liu, Kun Zhou +1
In this paper, we aim to study abelian extensions for some infinite group. We show that the Hopf algebra constructed through abelian extensions of $\Bbbk…
Preparation of NOON State Induced by Macroscopic Quantum Tunneling in an Ising Chain
Chun-Li Zang, Jing Yu, Wan-Li Yang +2
In this brief report, we propose a possible way, theoretically and experimentally, to generate a NOON state of the two degenerate ferromagnetic ground states of the Transverse Isin…
Align before Search: Aligning Ads Image to Text for Accurate Cross-Modal Sponsored Search
Yuanmin Tang, Jing Yu, Keke Gai +4
Cross-Modal sponsored search displays multi-modal advertisements (ads) when consumers look for desired products by natural language queries in search engines. Since multi-modal ads…
Online Adversarial Stabilization of Unknown Networked Systems
Jing Yu, Dimitar Ho, Adam Wierman
We investigate the problem of stabilizing an unknown networked linear system under communication constraints and adversarial disturbances. We propose the first provably stabilizing…
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
Jing Yu, Yibo Zhao, Jiapeng Zhu +4
The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxificati…
Towards Fast and Accurate Image-Text Retrieval with Self-Supervised Fine-Grained Alignment
Jiamin Zhuang, Jing Yu, Yang Ding +2
Image-text retrieval requires the system to bridge the heterogenous gap between vision and language for accurate retrieval while keeping the network lightweight-enough for efficien…
Intern-S1: A Scientific Multimodal Foundation Model
Lei Bai, Zhongrui Cai, Yuhang Cao +173
In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that…
Maximum in-general-position set in a random subset of
Yaobin Chen, Jiaxi Nie, Jing Yu +1
Let be the maximum possible size of a point set in general position in a -random subset of . We determine the order of magnitude of $α(…
End-to-End Learning and Intervention in Games
Jiayang Li, Jing Yu, Yu Marco Nie +1
In a social system, the self-interest of agents can be detrimental to the collective good, sometimes leading to social dilemmas. To resolve such a conflict, a central designer may…
AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection
Jianbo Gao, Keke Gai, Jing Yu +2
Recent advancement in large-scale Artificial Intelligence (AI) models offering multimodal services have become foundational in AI systems, making them prime targets for model theft…
DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
Xiaoze Jiang, Jing Yu, Zengchang Qin +4
Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of vis…
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
Jiahui Peng, Xinlin Zhuang, Jiantao Qiu +4
The performance of large language models (LLMs) is significantly affected by the quality and composition of their pre-training data, which is inherently diverse, spanning various l…
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
Bin Wang, Tianyao He, Linke Ouyang +40
Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art…
APEX: Academic Poster Editing Agentic Expert
Chengxin Shi, Qinnan Cai, Zeyuan Chen +5
Designing academic posters is a labor-intensive process requiring the precise balance of high-density content and sophisticated layout. While existing paper-to-poster generation me…
KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue
Xiaoze Jiang, Siyi Du, Zengchang Qin +2
Visual dialogue is a challenging task that needs to extract implicit information from both visual (image) and textual (dialogue history) contexts. Classical approaches pay more att…
The Simulation of Non-Abelian Statistics of Majorana Fermions in Ising Chain with Z2 Symmetry
Xiao-Ming Zhao, Jing Yu, Jing He +3
In this paper, we numerically study the non-Abelian statistics of the zero-energy Majorana fermions on the end of Majorana chain and show its application to quantum computing by ma…
A quiver approach to quasi-quantum groups with the Chevalley property
Jing Yu
In this paper, we develop a quiver approach to coquasi-Hopf algebras with the dual Chevalley property. We introduce a modified generalized path coalgebra $\Bbbk(\mathrm{Q},\mathcal…
Collaborative Belief Reasoning with LLMs for Efficient Multi-Agent Collaboration
Zhimin Wang, Duo Wu, Shaokang He +6
Effective real-world multi-agent collaboration requires not only accurate planning but also the ability to reason about collaborators' intents--a crucial capability for avoiding mi…
Evolving Attention with Residual Convolutions
Yujing Wang, Yaming Yang, Jiangang Bai +6
Transformer is a ubiquitous model for natural language processing and has attracted wide attentions in computer vision. The attention maps are indispensable for a transformer model…
Simple Yetter-Drinfeld modules over Generalized Liu algebras
Xiangjun Zhen, Gongxiang Liu, Jing Yu
Let be a generalized Liu algebra over an algebraically closed field of characteristic zero. We prove that all simple Yetter-Drinfeld modules over are finite-dimensional…
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
Keyan Ding, Jing Yu, Junjie Huang +3
Scientific research increasingly relies on specialized computational tools, yet effectively utilizing these tools demands substantial domain expertise. While Large Language Models…
Large-scale geometry of Borel graphs of polynomial growth
Anton Bernshteyn, Jing Yu
We study graphs of polynomial growth from the perspective of asymptotic geometry and descriptive set theory. The starting point of our investigation is a theorem of Krauthgamer and…
EASTER: Embedding Aggregation-based Heterogeneous Models Training in Vertical Federated Learning
Shuo Wang, Keke Gai, Jing Yu +3
Vertical federated learning has garnered significant attention as it allows clients to train machine learning models collaboratively without sharing local data, which protects the…
Exact results of the quantum phase transition for the topological order
Jing Yu, Su-Peng Kou, Xiao-Gang Wen
In this paper a duality between the d=2 Wen-plaquette model in a transverse field and the d=1 Ising model in a transverse field is used to learn the nature of the quantum phase tra…
Online Stabilization of Unknown Linear Time-Varying Systems
Jing Yu, Varun Gupta, Adam Wierman
This paper studies the problem of online stabilization of an unknown discrete-time linear time-varying (LTV) system under bounded non-stochastic (potentially adversarial) disturban…
System Identification Under Bounded Noise: Optimal Rates Beyond Least Squares
Xiong Zeng, Jing Yu, Necmiye Ozay
System identification is a fundamental problem in control and learning, particularly in high-stakes applications where data efficiency is critical. Classical approaches, such as th…
Solving Optimal Experimental Design with Sequential Quadratic Programming and Chebyshev Interpolation
Jing Yu, Mihai Anitescu
We propose an optimization algorithm to compute the optimal sensor locations in experimental design in the formulation of Bayesian inverse problems, where the parameter-to-observab…
Bi-directional Cognitive Thinking Network for Machine Reading Comprehension
Wei Peng, Yue Hu, Luxi Xing +4
We propose a novel Bi-directional Cognitive Knowledge Framework (BCKF) for reading comprehension from the perspective of complementary learning systems theory. It aims to simulate…
LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
Jing Yu, Yibo Zhao, Jiaming Zhang +1
Long-term memory enables LLM agents to leverage past interactions, but dialogue histories quickly exceed the context window, forcing agents to retrieve relevant subsets at query ti…
A pharmacokinetic -- viral kinetic model describes the effect of alisporivir monotherapy or in combination with peg-IFN on 2 hepatitis C virologic response
Thi Huyen Tram Nguyen, France Mentré, Micha Levi +2
Alisporivir is a cyclophilin inhibitor with demonstrated in vitro and in vivo activity against hepatitis C 11 virus (HCV). We estimated antiviral effectiveness of alisporivir alone…
On derived categories of module categories over multiring categories
Jing Yu
Let and be subcategories of tensor categories and , respectively, both of which are abelian categories with finitely many iso…
Embedding Borel graphs into grids of asymptotically optimal dimension
Anton Bernshteyn, Jing Yu
Let be a Borel graph all of whose finite subgraphs embed into the -dimensional grid with diagonals. We show that then itself admits a Borel embedding into the Schreier g…
Convolution-enhanced Evolving Attention Networks
Yujing Wang, Yaming Yang, Zhuo Li +7
Attention-based neural networks, such as Transformers, have become ubiquitous in numerous applications, including computer vision, natural language processing, and time-series anal…
Coarse-to-Careful: Seeking Semantic-related Knowledge for Open-domain Commonsense Question Answering
Luxi Xing, Yue Hu, Jing Yu +2
It is prevalent to utilize external knowledge to help machine answer questions that need background commonsense, which faces a problem that unlimited knowledge will transmit noisy…
CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation
Jing Yu, Yuan Chai, Yujing Wang +2
Scene graphs are semantic abstraction of images that encourage visual understanding and reasoning. However, the performance of Scene Graph Generation (SGG) is unsatisfactory when f…
-adic periods of Carlitz motives and Chowla-Selberg formula revisited
Chieh-Yu Chang, Fu-Tsun Wei, Jing Yu
Let be a finite place of . In this paper, we interpret -adic arithmetic gamma values in terms of the -adic crystalline-de Rham periods of Carlitz motive…
Achieving Sample Complexity for Bilinear Systems Identification under Bounded Noises
Hongyu Yi, Chenbei Lu, Jing Yu
This paper studies finite-sample set-membership identification for discrete-time bilinear systems under bounded symmetric log-concave disturbances. Our analysis considers trajector…
Scientific Large Language Models: A Survey on Biological & Chemical Domains
Qiang Zhang, Keyang Ding, Tianwen Lyv +22
Large Language Models (LLMs) have emerged as a transformative power in enhancing natural language comprehension, representing a significant stride toward artificial general intelli…
Frobenius difference equations and algebraic independence of zeta values in positive equal characteristic
Chieh-Yu Chang, Matthew A. Papanikolas, Jing Yu
In analogy with the Riemann zeta function at positive integers, for each finite field F_p^r with fixed characteristic p we consider Carlitz zeta values zeta_r(n) at positive intege…
Denoise-I2W: Mapping Images to Denoising Words for Accurate Zero-Shot Composed Image Retrieval
Yuanmin Tang, Jing Yu, Keke Gai +4
Zero-Shot Composed Image Retrieval (ZS-CIR) supports diverse tasks with a broad range of visual content manipulation intentions that can be related to domain, scene, object, and at…
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
Jialing Wang, Yue Zhao, Yuhao Zhang +5
Recent advances in Speech Large Language Models (Speech-LLMs) have made significant progress, greatly enhancing multimodal interaction capabilities.However, their application in lo…
YOLO-PRO: Enhancing Instance-Specific Object Detection with Full-Channel Global Self-Attention
Lin Huang, Yujuan Tan, Weisheng Li +5
This paper addresses the inherent limitations of conventional bottleneck structures (diminished instance discriminability due to overemphasis on batch statistics) and decoupled hea…
INet: Inter-Intra-slice Interpolation Network for Medical Slice Synthesis
Haofei Song, Xintian Mao, Jing Yu +2
Medical imaging is limited by acquisition time and scanning equipment. CT and MR volumes, reconstructed with thicker slices, are anisotropic with high in-plane resolution and low t…
Emergent Supersymmetric Many-Body Systems in Doped Z2 Topological Spin Liquid of the Toric-Code Model
Jing He, Jing Yu, Xing-Hai Zhang +1
In this paper, we studied the doped Z2 topological spin liquid of the toric-code model. We found that the doped holes become supersymmetric particles. The ground state of the doped…
InternLM2 Technical Report
Zheng Cai, Maosong Cao, Haojiong Chen +97
The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advan…
Coradically graded Hopf algebras of tame corepresentation type
Jing Yu, Gongxiang Liu
Let be an algebraically closed field of characteristic and let be a finite-dimensional Hopf algebra over with the dual Chevalley property. In this paper, we…
On Infinite-horizon System Level Synthesis Problems
Olle Kjellqvist, Jing Yu
System level synthesis is a promising approach that formulates structured optimal controller synthesis problems as convex problems. This work solves the distributed linear-quadrati…
Adaptive Prototype Knowledge Transfer for Federated Learning with Mixed Modalities and Heterogeneous Tasks
Keke Gai, Mohan Wang, Jing Yu +2
Multimodal Federated Learning (MFL) with mixed modalities enables unimodal and multimodal clients to collaboratively train models while ensuring clients' privacy. As a representati…
Semantic Modeling of Textual Relationships in Cross-Modal Retrieval
Jing Yu, Chenghao Yang, Zengchang Qin +3
Feature modeling of different modalities is a basic problem in current research of cross-modal information retrieval. Existing models typically project texts and images into one em…
YOLO-DS: Fine-Grained Feature Decoupling via Dual-Statistic Synergy Operator for Object Detection
Lin Huang, Yujuan Tan, Weisheng Li +6
One-stage object detection, particularly the YOLO series, strikes a favorable balance between accuracy and efficiency. However, existing YOLO detectors lack explicit modeling of he…