papers

Publications (296)

cs.CL2026

Large Language Models Could Be Rote Learners

Yuyang Xu, Renjun Hu, Haochao Ying +3

Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliabilit…

cs.CV2025

GBlobs: Local LiDAR Geometry for Improved Sensor Placement Generalization

Dušan Malić, Christian Fruhwirth-Reisinger, Alexander Prutsch +3

This technical report outlines the top-ranking solution for RoboSense 2025: Track 3, achieving state-of-the-art performance on 3D object detection under various sensor placements.…

cs.LG2026

A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources

Qingyu Song, Rui Liu, Wei Lin +11

Deploying Large Language Models (LLMs) on edge devices enhances privacy but faces performance hurdles due to limited resources. We introduce a systematic methodology to evaluate on…

hep-th2024

Meso-inflationary Peccei-Quinn symmetry breaking with non-minimal coupling

Yermek Aldabergenov, Ding Ding, Wei Lin +1

We study a realization of the inflationary scenario where the Peccei-Quinn (PQ) symmetry is spontaneously broken during inflation, facilitated by its non-minimal coupling to gravit…

cs.CV2023

TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification

M. Jehanzeb Mirza, Leonid Karlinsky, Wei Lin +3

Vision and Language Models (VLMs), such as CLIP, have enabled visual recognition of a potentially unlimited set of categories described by text prompts. However, for the best visua…

cs.LG2026

The Multi-Query Paradox in Zeroth-Order Optimization

Wei Lin, Qingyu Song, Hong Xu

Zeroth-order (ZO) optimization provides a powerful framework for problems where explicit gradients are unavailable and have to be approximated using only queries to function value.…

cs.LG2026

Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards

Yingyu Shan, Yuhang Guo, Zihao Cheng +7

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in training LLMs for reasoning tasks, but representative methods such as GRPO assign uniform c…

gr-qc2026

Bouncing cosmologies from Born-Infeld-type gravity

Yermek Aldabergenov, Wei Lin, Rongjian Li +2

We construct a Born-Infeld-type modification of gravity, where is the Gauss-Bonnet term, by embedding Born-Infeld electrodynamics in a five-dimensional p…

eess.SY2025

Distributed Platoon Control Under Quantization: Stability Analysis and Privacy Preservation

Kaixiang Zhang, Zhaojian Li, Wei Lin

Distributed control of connected and automated vehicles has attracted considerable interest for its potential to improve traffic efficiency and safety. However, such control scheme…

cs.LG2024

QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Ying Zhang, Peng Zhang, Mincong Huang +7

Quantization is a proven effective method for compressing large language models. Although popular techniques like W8A8 and W4A16 effectively maintain model performance, they often…

eess.SY2021

Cost Functions over Feasible Power Transfer Regions of Virtual Power Plants

Wei Lin, Changhong Zhao

A virtual power plant (VPP) facilitates the integration of distributed energy resources (DERs) for the transmission-level operation. A challenge in operating a VPP is to characteri…

cs.CL2026

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

Li Wang, Xiaohan Wang, Xiaodong Lu +5

Large language models (LLMs) have increasingly leveraged tool invocation to enhance their reasoning capabilities. However, existing approaches typically tightly couple tool invocat…

cs.LG2025

Bi-Level Decision-Focused Causal Learning for Large-Scale Marketing Optimization: Bridging Observational and Experimental Data

Shuli Zhang, Hao Zhou, Jiaqi Zheng +4

Online Internet platforms require sophisticated marketing strategies to optimize user retention and platform revenue -- a classical resource allocation problem. Traditional solutio…

eess.IV2025

ViG3D-UNet: Volumetric Vascular Connectivity-Aware Segmentation via 3D Vision Graph Representation

Bowen Liu, Chunlei Meng, Wei Lin +4

Accurate vascular segmentation is essential for coronary visualization and the diagnosis of coronary heart disease. This task involves the extraction of sparse tree-like vascular b…

cs.CV2018

Graph-Adaptive Pruning for Efficient Inference of Convolutional Neural Networks

Mengdi Wang, Qing Zhang, Jun Yang +2

In this work, we propose a graph-adaptive pruning (GAP) method for efficient inference of convolutional neural networks (CNNs). In this method, the network is viewed as a computati…

cs.GR2025

SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation

Shenggan Cheng, Yuanxin Wei, Lansong Diao +8

Leveraging the diffusion transformer (DiT) architecture, models like Sora, CogVideoX and Wan have achieved remarkable progress in text-to-video, image-to-video, and video editing t…

physics.optics2024

Bistability in spatiotemporal mode-locking with dynamic multimode gain

Zhijin Xiong, Yuankai Guo, Wei Lin +8

Three-dimensional (3D) dissipative soliton existed in spatiotemporal mode-locked (STML) multimode fiber laser has been demonstrated to be a promising formalism for generating high-…

cs.CL2025

Privacy-Preserving Reasoning with Knowledge-Distilled Parametric Retrieval Augmented Generation

Jinwen Chen, Hainan Zhang, Liang Pang +5

The current RAG system requires uploading plaintext documents to the cloud, risking private data leakage. Parametric RAG (PRAG) encodes documents as LoRA parameters within LLMs, of…

cs.LG2023

Neural Delay Differential Equations: System Reconstruction and Image Classification

Qunxi Zhu, Yao Guo, Wei Lin

Neural Ordinary Differential Equations (NODEs), a framework of continuous-depth neural networks, have been widely applied, showing exceptional efficacy in coping with representativ…

cs.CL2025

From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory

Siyu Xia, Zekun Xu, Jiajun Chai +7

Large Language Models (LLMs) based agents have demonstrated remarkable potential in autonomous task-solving across complex, open-ended environments. A promising approach for improv…

cs.CV2023

AIR-DA: Adversarial Image Reconstruction for Unsupervised Domain Adaptive Object Detection

Kunyang Sun, Wei Lin, Haoqin Shi +3

Unsupervised domain adaptive object detection is a challenging vision task where object detectors are adapted from a label-rich source domain to an unlabeled target domain. Recent…

cs.LG2026

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

Zihan Lin, Xiaohan Wang, Jie Cao +6

Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivi…

cs.AI2026

Emergence Transformer: Dynamical Temporal Attention Matters

Zihan Zhou, Bo-Wei Qin, Kai Du +1

The Transformer, a breakthrough architecture in artificial intelligence, owes its success to the attention mechanism, which utilizes long-range interactions in sequential data, ena…

cs.DC2019

AliGraph: A Comprehensive Graph Neural Network Platform

Rong Zhu, Kun Zhao, Hongxia Yang +5

An increasing number of machine learning tasks require dealing with large graph datasets, which capture rich and complex relationship among potentially billions of elements. Graph…

cs.CV2025

ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model

Juntian Zhang, Song Jin, Chuanqi Cheng +8

The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging…

cs.CV2024

Vision-Language Guidance for LiDAR-based Unsupervised 3D Object Detection

Christian Fruhwirth-Reisinger, Wei Lin, Dušan Malić +2

Accurate 3D object detection in LiDAR point clouds is crucial for autonomous driving systems. To achieve state-of-the-art performance, the supervised training of detectors requires…

cs.CV2025

Point-to-Region Loss for Semi-Supervised Point-Based Crowd Counting

Wei Lin, Chenyang Zhao, Antoni B. Chan

Point detection has been developed to locate pedestrians in crowded scenes by training a counter through a point-to-point (P2P) supervision scheme. Despite its excellent localizati…

cs.CL2016

Neural Networks Models for Entity Discovery and Linking

Dan Liu, Wei Lin, Shiliang Zhang +2

This paper describes the USTC_NELSLIP systems submitted to the Trilingual Entity Detection and Linking (EDL) track in 2016 TAC Knowledge Base Population (KBP) contests. We have bui…

cs.CV2024

Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs

M. Jehanzeb Mirza, Leonid Karlinsky, Wei Lin +5

Prompt ensembling of Large Language Model (LLM) generated category-specific prompts has emerged as an effective method to enhance zero-shot recognition ability of Vision-Language M…

cs.CV2024

A Fixed-Point Approach to Unified Prompt-Based Counting

Wei Lin, Antoni B. Chan

Existing class-agnostic counting models typically rely on a single type of prompt, e.g., box annotations. This paper aims to establish a comprehensive prompt-based counting framewo…

cs.CV2026

GenLie: A Global-Enhanced Lie Detection Network under Sparsity and Semantic Interference

Zongshun Zhang, Yao Liu, Qiao Liu +5

Video-based lie detection aims to identify deceptive behaviors from visual cues. Despite recent progress, its core challenge lies in learning sparse yet discriminative representati…

cs.CL2026

LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services

Jinwen Chen, Shiwen Zhang, Shuai Gong +6

In local-life service platforms, query suggestion reduces user effort by generating candidate queries from input prefixes. Traditional multi-stage systems rely heavily on historica…

cs.LG2023

Embedding Theory of Reservoir Computing and Reducing Reservoir Network Using Time Delays

Xing-Yue Duan, Xiong Ying, Si-Yang Leng +3

Reservoir computing (RC), a particular form of recurrent neural network, is under explosive development due to its exceptional efficacy and high performance in reconstruction or/an…

cs.CL2025

Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems

Song Jin, Juntian Zhang, Yuhan Liu +6

Evaluating and iterating upon recommender systems is crucial, yet traditional A/B testing is resource-intensive, and offline methods struggle with dynamic user-platform interaction…

cs.DC2019

FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads

Guoping Long, Jun Yang, Wei Lin

Performance optimization is the art of continuous seeking a harmonious mapping between the application domain and hardware. Recent years have witnessed a surge of deep learning (DL…

stat.ML2022

Heterogeneous Federated Learning on a Graph

Huiyuan Wang, Xuyang Zhao, Wei Lin

Federated learning, where algorithms are trained across multiple decentralized devices without sharing local data, is increasingly popular in distributed machine learning practice.…

cs.AI2025

Promoting Efficient Reasoning with Verifiable Stepwise Reward

Chuhuai Yue, Chengqi Dong, Yinan Gao +4

Large reasoning models (LRMs) have recently achieved significant progress in complex reasoning tasks, aided by reinforcement learning with verifiable rewards. However, LRMs often s…

cs.LG2026

Your Group-Relative Advantage Is Biased

Fengkai Yang, Zherui Chen, Xiaohan Wang +10

Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such…

cs.CV2025

Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation

Johannes Spoecklberger, Wei Lin, Pedro Hermosilla +3

Vision Foundation Models (VFMs) have become a de facto choice for many downstream vision tasks, like image classification, image segmentation, and object localization. However, the…

physics.optics2019

Supercontinuum generation without residual pump peak through multiple coherent pump seeds

Dong Qiu, Wei Lin, Hong-Jie Chen +8

Residual pump peak in fiber-based supercontinuum, as a general phenomenon, limits its practical application. We report a novel supercontinuum generation (SCG) in a conventional hig…

cs.AR2026

Towards transferable lightweight neuromorphic computing through a model-free temporal-switch framework

Zefeng Zhang, Chao Li, Siyao Chen +5

Lightweight neuromorphic computing offers a promising route to efficient AI, with particular benefits for resource-constrained edge deployments. However, its scalable deployment th…

cs.DC2024

Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling

Xinyi Zhang, Hanyu Zhao, Wencong Xiao +5

The era of large deep learning models has given rise to advanced training strategies such as 3D parallelism and the ZeRO series. These strategies enable various (re-)configurable e…

cs.IR2026

CORE: A Unified Cascaded Ordinal Relevance Estimation Framework for E-commerce Search

Zhi Jin, Xi Wang, Yunfei Li +3

Ranking relevance is a fundamental task in e-commerce search, directly affecting ranking quality and consumer experience. Although inherently an ordinal classification problem, it…

math.GT2022

A standard form of incompressible surfaces in 3-dimensional handlebodies

Wei Lin, Fengchun Lei

Applying Morse theory, we give a standard form for a class of surfaces which includes all the properly embedded incompressible surfaces in 3-dimensional handlebodies. We also give…

cs.IR2026

Unpaired Modality-Agnostic Generative Recommendation

Weihao Shen, Wei Chen, Fuwei Zhang +6

Generative Recommendation (GR) formulates recommendation as autoregressive generation over discrete semantic identifiers (IDs). Although recent multimodal GR methods improve semant…

cs.AI2026

AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning

Jingbo Sun, Wenyue Chong, Songjun Tu +7

Agentic retrieval-augmented generation (RAG) systems enable large language models (LLMs) to solve complex tasks through multi-step interaction with external retrieval tools. Howeve…

cs.CV2025

Pose-Free 3D Quantitative Phase Imaging of Flowing Cellular Populations

Enze Ye, Wei Lin, Shaochi Ren +5

High-throughput 3D quantitative phase imaging (QPI) in flow cytometry enables label-free, volumetric characterization of individual cells by reconstructing their refractive index (…

cs.CV2025

PerLA: Perceptive 3D Language Assistant

Guofeng Mei, Wei Lin, Luigi Riz +3

Enabling Large Language Models (LLMs) to understand the 3D physical world is an emerging yet challenging research direction. Current strategies for processing point clouds typicall…

astro-ph.EP2010

Measurements of Transit Timing Variations for WASP-5b

Akihiko Fukui, Norio Narita, Paul J. Tristram +34

We have observed 7 new transits of the `hot Jupiter' WASP-5b using a 61 cm telescope located in New Zealand, in order to search for transit timing variations (TTVs) which can be in…

cs.CL2018

Transfer Learning for Context-Aware Question Matching in Information-seeking Conversations in E-commerce

Minghui Qiu, Liu Yang, Feng Ji +6

Building multi-turn information-seeking conversation systems is an important and challenging research topic. Although several advanced neural text matching models have been propose…

cs.DC2018

FusionStitching: Deep Fusion and Code Generation for Tensorflow Computations on GPUs

Guoping Long, Jun Yang, Kai Zhu +1

In recent years, there is a surge on machine learning applications in industry. Many of them are based on popular AI frameworks like Tensorflow, Torch, Caffe, or MxNet, etc, and ar…

cond-mat.stat-mech2026

Exact Resonances Are Not Sufficient for Phonon Energy Diffusion

Wei Lin, Yong Zhang, Hong Zhao

Multi-phonon resonance conditions underpin kinetic theories of phonon transport and lattice thermalization. We show that exact resonance matching, nonzero interaction coefficients,…

stat.ME2024

DEEPEAST technique to enhance power in two-sample tests via the same-attraction function

Yiting Chen, Min Gao, Wei Lin +3

Data depth has emerged as an invaluable nonparametric measure for the ranking of multivariate samples. The main contribution of depth-based two-sample comparisons is the introducti…

math.OC2025

Neural Event-Triggered Control with Optimal Scheduling

Luan Yang, Jingdong Zhang, Qunxi Zhu +1

Learning-enabled controllers with stability certificate functions have demonstrated impressive empirical performance in addressing control problems in recent years. Nevertheless, d…

cs.IR2026

MTFM: A Scalable and Alignment-free Foundation Model for Industrial Recommendation in Meituan

Xin Song, Zhilin Guan, Ruidong Han +12

Industrial recommendation systems typically involve multiple scenarios, yet existing cross-domain (CDR) and multi-scenario (MSR) methods often require prohibitive resources and str…

q-fin.MF2017

Pricing VIX Derivatives With Free Stochastic Volatility Model

Wei Lin, Shenghong Li, Shane Chern

In this paper, we relax the power parameter of instantaneous variance and develop a new stochastic volatility plus jumps model that generalize the Heston model and 3/2 model as spe…

stat.ME2016

Large Covariance Estimation for Compositional Data via Composition-Adjusted Thresholding

Yuanpei Cao, Wei Lin, Hongzhe Li

High-dimensional compositional data arise naturally in many applications such as metagenomic data analysis. The observed data lie in a high-dimensional simplex, and conventional st…

cs.LG2025

Learning Provably Improves the Convergence of Gradient Descent

Qingyu Song, Wei Lin, Hong Xu

Learn to Optimize (L2O) trains deep neural network-based solvers for optimization, achieving success in accelerating convex problems and improving non-convex solutions. However, L2…

physics.soc-ph2025

Quantifying Transient Dynamics in Heterogeneous Networks under Various Inputs

Xiaoge Bao, Wei P. Dai, Jan Nagler +1

Understanding how transient dynamics unfold in response to localized inputs is central to predicting and controlling signal propagation in network systems, including neural process…

cs.CL2024

AlphaFin: Benchmarking Financial Analysis with Retrieval-Augmented Stock-Chain Framework

Xiang Li, Zhenyu Li, Chen Shi +5

The task of financial analysis primarily encompasses two key areas: stock trend prediction and the corresponding financial question answering. Currently, machine learning and deep…

stat.ME2017

Nonsparse learning with latent variables

Zemin Zheng, Jinchi Lv, Wei Lin

As a popular tool for producing meaningful and interpretable models, large-scale sparse learning works efficiently when the underlying structures are indeed or close to sparse. How…

physics.optics2017

Fast two-snapshot structured illumination for temporal focusing microscopy with enhanced axial resolution

Yunlong Meng, Wei Lin, Chenglin Li +1

We present a new two-snapshot structured light illumination (SLI) reconstruction algorithm for fast image acquisition. The new algorithm, which only requires two mutually π phase-…

q-bio.QM2011

Bifurcations of Emergent Bursting in a Neuronal Network

Yu Wu, Wenlian Lu, Wei Lin +2

Currently we routinely develop a complex neuronal network to explain observed but often paradoxical phenomena based upon biological recordings. Here we present a general approach t…

cs.CV2024

Comparison Visual Instruction Tuning

Wei Lin, Muhammad Jehanzeb Mirza, Sivan Doveh +4

Comparing two images in terms of Commonalities and Differences (CaD) is a fundamental human capability that forms the basis of advanced visual reasoning and interpretation. It is e…

cs.LG2026

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration

Songhao Wu, Ang Lv, Xiao Feng +5

The KV cache in large language models is a dominant factor in memory usage, limiting their broader applicability. Quantizing the cache to lower bit widths is an effective way to re…

cs.CV2022

Multi-Frequency-Aware Patch Adversarial Learning for Neural Point Cloud Rendering

Jay Karhade, Haiyue Zhu, Ka-Shing Chung +3

We present a neural point cloud rendering pipeline through a novel multi-frequency-aware patch adversarial learning framework. The proposed approach aims to improve the rendering r…

math.GT2021

On Crossing Ball Structure in Knot and Link Complements

Wei Lin

We develop a word mechanism applied in knot and link diagrams for the illustration of a diagrammatic property. We also give a necessary condition for determining incompressible and…

math.ST2023

Multivariate two-sample test statistics based on data depth

Yiting Chen, Wei Lin, Xiaoping Shi

Data depth has been applied as a nonparametric measurement for ranking multivariate samples. In this paper, we focus on homogeneity tests to assess whether two multivariate samples…

cs.CV2026

PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies

Lukas Selch, Yufang Hou, M. Jehanzeb Mirza +4

Large Multimodal Models (LMMs) are increasingly applied to scientific research, yet it remains unclear whether they can reliably understand and reason over the multimodal complexit…

cs.LG2025

Towards Robust Learning to Optimize with Theoretical Guarantees

Qingyu Song, Wei Lin, Juncheng Wang +1

Learning to optimize (L2O) is an emerging technique to solve mathematical optimization problems with learning-based methods. Although with great success in many real-world scenario…

q-fin.CP2015

Consistent Pricing of VIX and Equity Derivatives with the 4/2 Stochastic Volatility Plus Jumps Model

Wei Lin, Shenghong Li, Xingguo Luo +1

In this paper, we develop a 4/2 stochastic volatility plus jumps model, namely, a new stochastic volatility model including the Heston model and 3/2 model as special cases. Our mod…

cs.CV2024

Towards Multimodal In-Context Learning for Vision & Language Models

Sivan Doveh, Shaked Perek, M. Jehanzeb Mirza +5

State-of-the-art Vision-Language Models (VLMs) ground the vision and the language modality primarily via projecting the vision tokens from the encoder to language-like tokens, whic…

cs.IR2018

EENMF: An End-to-End Neural Matching Framework for E-Commerce Sponsored Search

Wenjin Wu, Guojun Liu, Hui Ye +5

E-commerce sponsored search contributes an important part of revenue for the e-commerce company. In consideration of effectiveness and efficiency, a large-scale sponsored search sy…

cs.AI2026

MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation Agents

Chuhan Qiao, Jianghua Huang, Daxing Zhao +5

Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential fo…

q-bio.MN2024

Exact first passage time distribution for second-order reactions in chemical networks

Changqian Rao, David Waxman, Wei Lin +1

The first passage time (FPT) is a generic measure that quantifies when a random quantity reaches a specific state. We consider the FTP distribution in nonlinear stochastic biochemi…

cs.CV2026

Fine-grained CLIP fine-tuning with self-annotated region alignment

Chenyang Zhao, Wei Lin, Antoni B. Chan +1

The paper proposes SFF-CLIP, a fine-tuning approach that uses only image-text pairs to align region features with phrase concepts via text-specific heat maps, improving CLIP's fine…

#vision-language#fine-grained representation#self-supervised fine-tuning#region alignment
math.GT2022

A meridian lemma for fully alternating links in thickened surfaces

Wei Lin

Menasco showed that a closed surface in the complement of a non-split prime alternating link in contains a circle isotopic in the link complement to a meridian of the links.…

cs.CV2020

Grasping Detection Network with Uncertainty Estimation for Confidence-Driven Semi-Supervised Domain Adaptation

Haiyue Zhu, Yiting Li, Fengjun Bai +6

Data-efficient domain adaptation with only a few labelled data is desired for many robotic applications, e.g., in grasping detection, the inference skill learned from a grasping da…

cs.CV2026

Ensemble Deep Learning Approaches for AI-Altered Video Detection

Laiba Khan, Hung-Mao Wu, Wei Lin +3

The increasing accessibility of artificial intelligence has led to a rapid rise in AI-generated videos, making it more difficult to distinguish between real and manipulated content…

cs.LG2025

pLSTM: parallelizable Linear Source Transition Mark networks

Korbinian Pöppel, Richard Freinschlag, Thomas Schmied +2

Modern recurrent architectures, such as xLSTM and Mamba, have recently challenged the Transformer in language modeling. However, their structure constrains their applicability to s…

cs.CL2021

AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search

Daoyuan Chen, Yaliang Li, Minghui Qiu +7

Large pre-trained language models such as BERT have shown their effectiveness in various natural language processing tasks. However, the huge parameter size makes them difficult to…

cs.CV2023

MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge

Wei Lin, Leonid Karlinsky, Nina Shvetsova +6

Large scale Vision-Language (VL) models have shown tremendous success in aligning representations between visual and text modalities. This enables remarkable progress in zero-shot…

math.GT2025

Amalgamations along surfaces with boundary in a handlebody

Siqi Ding, Fengchun Lei, Wei Lin +1

Let M be a connected orientable 3-manifold, and F a compact connected orientable surface properly embedded in M. If F cuts M into two connected 3-manifolds X and Y, that is, M=X \c…

cs.CV2025

Wan: Open and Advanced Large-Scale Video Generative Models

Team Wan, Ang Wang, Baole Ai +58

This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transfo…

cs.LG2025

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Jiajun Chai, Guojun Yin, Zekun Xu +9

Large language models excel at basic reasoning but struggle with tasks that require interaction with external tools. We present RLFactory, a plug-and-play reinforcement learning po…

cs.CV2020

One-shot Text Field Labeling using Attention and Belief Propagation for Structure Information Extraction

Mengli Cheng, Minghui Qiu, Xing Shi +2

Structured information extraction from document images usually consists of three steps: text detection, text recognition, and text field labeling. While text detection and text rec…

cs.IR2025

CAT-ID: Category-Tree Integrated Document Identifier Learning for Generative Retrieval In E-commerce

Xiaoyu Liu, Fuwei Zhang, Yiqing Wu +6

Generative retrieval (GR) has gained significant attention as an effective paradigm that integrates the capabilities of large language models (LLMs). It generally consists of two s…

cs.IR2025

IterQR: An Iterative Framework for LLM-based Query Rewrite in e-Commercial Search System

Shangyu Chen, Xinyu Jia, Yingfei Zhang +3

The essence of modern e-Commercial search system lies in matching user's intent and available candidates depending on user's query, providing personalized and precise service. Howe…

cs.CL2025

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

Renjun Hu, Yi Cheng, Libin Meng +4

The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge tha…

math.DS2024

A geometric approach for stability analysis of delay systems: Applications to network dynamics

Shijie Zhou, Yang Luan, Xuzhe Qian +1

Investigating the network stability or synchronization dynamics of multi-agent systems with time delays is of significant importance in numerous real-world applications. Such inves…

cs.IR2024

Context-based Fast Recommendation Strategy for Long User Behavior Sequence in Meituan Waimai

Zhichao Feng, Junjiie Xie, Kaiyuan Li +7

In the recommender system of Meituan Waimai, we are dealing with ever-lengthening user behavior sequences, which pose an increasing challenge to modeling user preference effectivel…

cs.AI2025

MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL

Zekun Xu, Siyu Xia, Chuhuai Yue +6

As large language models (LLMs) are increasingly used in Text-to-SQL tasks, Reinforcement Learning (RL) has become a common method for improving performance. Existing methods prima…

cs.DC2025

Efficient Long Context Fine-tuning with Chunk Flow

Xiulong Yuan, Hongtao Xu, Wenting Shen +10

Long context fine-tuning of large language models(LLMs) involves training on datasets that are predominantly composed of short sequences and a small proportion of longer sequences.…

cs.CV2025

GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

M. Jehanzeb Mirza, Mengjie Zhao, Zhuoyuan Mao +12

In this work, we propose GLOV, which enables Large Language Models (LLMs) to act as implicit optimizers for Vision-Language Models (VLMs) to enhance downstream vision tasks. GLOV p…

cs.LG2023

Nonasymptotic theory for two-layer neural networks: Beyond the bias-variance trade-off

Huiyuan Wang, Wei Lin

Large neural networks have proved remarkably effective in modern deep learning practice, even in the overparametrized regime where the number of active parameters is large relative…

cs.AR2024

Llumnix: Dynamic Scheduling for Large Language Model Serving

Biao Sun, Ziming Huang, Hanyu Zhao +4

Inference serving for large language models (LLMs) is the key to unleashing their potential in people's daily lives. However, efficient LLM serving remains challenging today becaus…

cs.DC2023

Ada-Grouper: Accelerating Pipeline Parallelism in Preempted Network by Adaptive Group-Scheduling for Micro-Batches

Siyu Wang, Zongyan Cao, Chang Si +3

Pipeline parallelism has been demonstrated to be a remarkable approach to improve throughput for training deep neural networks with billions of parameters over heterogeneous cluste…

cs.CL2025

AR-Med: Automated Relevance Enhancement in Medical Search via LLM-Driven Information Augmentation

Chuyue Wang, Jie Feng, Yuxi Wu +4

Accurate and reliable search on online healthcare platforms is critical for user safety and service efficacy. Traditional methods, however, often fail to comprehend complex and nua…

cs.CL2025

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Yuqian Fu, Tinghong Chen, Jiajun Chai +7

Large language models (LLMs) have achieved remarkable progress in reasoning tasks, yet the optimal integration of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) remai…

cs.DC2022

Optimizing DNN Compilation for Distributed Training with Joint OP and Tensor Fusion

Xiaodong Yi, Shiwei Zhang, Lansong Diao +6

This paper proposes DisCo, an automatic deep learning compilation module for data-parallel distributed training. Unlike most deep learning compilers that focus on training or infer…