Publications (23)
On Reasoning with RDF Statements about Statements using Singleton Property Triples
Vinh Nguyen, Olivier Bodenreider, Krishnaprasad Thirunarayan +6
The Singleton Property (SP) approach has been proposed for representing and querying metadata about RDF triples such as provenance, time, location, and evidence. In this approach,…
QwenLong-CPRS: Towards -LLMs with Dynamic Context Optimization
Weizhou Shen, Chenliang Li, Fanqi Wan +12
This technical report presents QwenLong-CPRS, a context compression framework designed for explicit long-context optimization, addressing prohibitive computation overhead during th…
Greedy PIG: Adaptive Integrated Gradients
Kyriakos Axiotis, Sami Abu-al-haija, Lin Chen +2
Deep learning has become the standard approach for most machine learning tasks. While its impact is undeniable, interpreting the predictions of deep learning models from a human pe…
Exposing Provenance Metadata Using Different RDF Models
Gang Fu, Evan Bolton, Núria Queralt Rosinach +5
A standard model for exposing structured provenance metadata of scientific assertions on the Semantic Web would increase interoperability, discoverability, reliability, as well as…
edge2vec: Representation learning using edge semantics for biomedical knowledge discovery
Zheng Gao, Gang Fu, Chunping Ouyang +8
Representation learning provides new and powerful graph analytical approaches and tools for the highly valued data science challenge of mining knowledge graphs. Since previous grap…
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
Adel Javanmard, Lin Chen, Vahab Mirrokni +2
Due to the rise of privacy concerns, in many practical applications the training data is aggregated before being shared with the learner, in order to protect privacy of users' sens…
Feature Cross Search via Submodular Optimization
Lin Chen, Hossein Esfandiari, Gang Fu +2
In this paper, we study feature cross search as a fundamental primitive in feature engineering. The importance of feature cross search especially for the linear model has been know…
Downsizing Diffusion Models for Cardinality Estimation
Xinhe Mu, Zhaoqi Zhou, Zaijiu Shang +5
Learned cardinality estimation requires accurate model designs to capture the local characteristics of probability distributions. However, existing models may fail to accurately ca…
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
Michal Lukasik, Lin Chen, Harikrishna Narasimhan +7
Bipartite ranking is a fundamental supervised learning problem, with the goal of learning a ranking over instances with maximal Area Under the ROC Curve (AUC) against a single bina…
Correlation Matching Transformation Transformers for UHD Image Restoration
Cong Wang, Jinshan Pan, Wei Wang +5
This paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution spac…
Greedy Column Subset Selection: New Bounds and Distributed Algorithms
Jason Altschuler, Aditya Bhaskara, Gang Fu +3
The problem of column subset selection has recently attracted a large body of research, with feature selection serving as one obvious and important application. Among the technique…
Approximately Optimal Core Shapes for Tensor Decompositions
Mehrdad Ghadiri, Matthew Fahrbach, Gang Fu +1
This work studies the combinatorial optimization problem of finding an optimal core tensor shape, also called multilinear rank, for a size-constrained Tucker decomposition. We give…
Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic Data
Gang Fu, Qing Zhang, Lei Zhu +2
This paper aims to remove specular highlights from a single object-level image. Although previous methods have made some progresses, their performance remains somewhat limited, par…
How Powerful Potential of Attention on Image Restoration?
Cong Wang, Jinshan Pan, Yeying Jin +5
Transformers have demonstrated their effectiveness in image restoration tasks. Existing Transformer architectures typically comprise two essential components: multi-head self-atten…
SequentialAttention++ for Block Sparsification: Differentiable Pruning Meets Combinatorial Optimization
Taisuke Yasuda, Kyriakos Axiotis, Gang Fu +2
Neural network pruning is a key technique towards engineering large yet scalable, interpretable, and generalizable models. Prior work on the subject has developed largely along two…
Learning from Aggregated Data: Curated Bags versus Random Bags
Lin Chen, Gang Fu, Amin Karbasi +1
Protecting user privacy is a major concern for many machine learning systems that are deployed at scale and collect from a diverse set of population. One way to address this concer…
Deep Image-based Illumination Harmonization
Zhongyun Bao, Chengjiang Long, Gang Fu +4
Integrating a foreground object into a background scene with illumination harmonization is an important but challenging task in computer vision and augmented reality community. Exi…
DeepCrossAttention: Supercharging Transformer Residual Connections
Mike Heddes, Adel Javanmard, Kyriakos Axiotis +3
Transformer networks have achieved remarkable success across diverse domains, leveraging a variety of architectural innovations, including residual connections. However, traditiona…
Sequential Attention for Feature Selection
Taisuke Yasuda, MohammadHossein Bateni, Lin Chen +3
Feature selection is the problem of selecting a subset of features for a machine learning model that maximizes model quality subject to a budget constraint. For neural networks, pr…
WebDancer: Towards Autonomous Information Seeking Agency
Jialong Wu, Baixuan Li, Runnan Fang +10
Addressing intricate real-world problems necessitates in-depth information seeking and multi-step reasoning. Recent progress in agentic systems, exemplified by Deep Research, under…
Deep & Cross Network for Ad Click Predictions
Ruoxi Wang, Bin Fu, Gang Fu +1
Feature engineering has been the key to the success of many prediction models. However, the process is non-trivial and often requires manual feature engineering or exhaustive searc…
Accelerating Scientific Research with Gemini: Case Studies and Common Techniques
David P. Woodruff, Vincent Cohen-Addad, Lalit Jain +33
Recent advances in large language models (LLMs) have opened new avenues for accelerating scientific research. While models are increasingly capable of assisting with routine tasks,…
Tongyi DeepResearch Technical Report
Tongyi DeepResearch Team, Baixuan Li, Bo Zhang +54
We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous…