papers

Publications (123)

cs.GR2024

Real-time High-resolution View Synthesis of Complex Scenes with Explicit 3D Visibility Reasoning

Tiansong Zhou, Yebin Liu, Xuangeng Chu +4

Rendering photo-realistic novel-view images of complex scenes has been a long-standing challenge in computer graphics. In recent years, great research progress has been made on enh…

cs.LG2026

ADaPT: Token-Level Decoupling for Efficient Large Reasoning Models

Tingyun Li, Zishang Jiang, Jinyi Han +8

Large reasoning models rely on long chain-of-thought to achieve strong performance, but applying such reasoning uniformly incurs high computational cost. Existing efficiency-orient…

cs.RO2025

JAM: Keypoint-Guided Joint Prediction after Classification-Aware Marginal Proposal for Multi-Agent Interaction

Fangze Lin, Ying He, Fei Yu +1

Predicting the future motion of road participants is a critical task in autonomous driving. In this work, we address the challenge of low-quality generation of low-probability mode…

cs.CV2025

Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning

Fei Yu, Quan Deng, Shengeng Tang +2

Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static…

cs.CV2024

Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering

Junxiao Xue, Quan Deng, Fei Yu +3

Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tas…

q-fin.ST2024

MDGNN: Multi-Relational Dynamic Graph Neural Network for Comprehensive and Dynamic Stock Investment Prediction

Hao Qian, Hongting Zhou, Qian Zhao +7

The stock market is a crucial component of the financial system, but predicting the movement of stock prices is challenging due to the dynamic and intricate relations arising from…

math.AG2026

Gorenstein singularities with -action and moduli spaces of holomorphic differentials

Dawei Chen, Fei Yu

Given a holomorphic differential on a smooth complex algebraic curve, we associate to it a Gorenstein curve singularity with -action via a test configuration. This con…

cs.CL2026

Toward Automated Robustness Evaluation of Mathematical Reasoning

Yutao Hou, Zeguan Xiao, Fei Yu +6

Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpected brittleness, often failing on…

cs.CV2025

OnlineHOI: Towards Online Human-Object Interaction Generation and Perception

Yihong Ji, Yunze Liu, Yiyao Zhuo +4

The perception and generation of Human-Object Interaction (HOI) are crucial for fields such as robotics, AR/VR, and human behavior understanding. However, current approaches model…

eess.SP2024

Multimodal Trustworthy Semantic Communication for Audio-Visual Event Localization

Yuandi Li, Zhe Xiang, Fei Yu +4

The exponential growth in wireless data traffic, driven by the proliferation of mobile devices and smart applications, poses significant challenges for modern communication systems…

cs.CV2022

Region-Aware Metric Learning for Open World Semantic Segmentation via Meta-Channel Aggregation

Hexin Dong, Zifan Chen, Mingze Yuan +5

As one of the most challenging and practical segmentation tasks, open-world semantic segmentation requires the model to segment the anomaly regions in the images and incrementally…

cs.CL2023

HuatuoGPT, towards Taming Language Model to Be a Doctor

Hongbo Zhang, Junying Chen, Feng Jiang +10

In this paper, we present HuatuoGPT, a large language model (LLM) for medical consultation. The core recipe of HuatuoGPT is to leverage both \textit{distilled data from ChatGPT} an…

physics.med-ph2018

High fidelity fibre-based physiological sensing deep in tissue

Tushar R. Choudhary, Michael G. Tanner, Alicia Megia-Fernandez +13

Physiological sensing deep in tissue, remains a clinical challenge. Here a flexible miniaturised sensing optrode providing a platform to perform minimally invasive in vivo in situ…

cs.CL2025

Order Matters: Investigate the Position Bias in Multi-constraint Instruction Following

Jie Zeng, Qianyu He, Qingyu Ren +5

Real-world instructions with multiple constraints pose a significant challenge to existing large language models (LLMs). An observation is that the LLMs exhibit dramatic performanc…

physics.optics2026

Electron-beam Writing of Spectrally Uniform Green Single-photon Emitters in Hexagonal Boron Nitride

Qingsong Tao, Fuyi Zhou, Zhijie Li +16

Scalable quantum photonic technologies require single-photon emitters whose positions and emission energies can be engineered simultaneously. Hexagonal boron nitride (hBN) is an at…

astro-ph.CO2010

The influence of quintessence on the motion of a binary system in cosmology

Fei Yu, Molin Liu, Yuanxing Gui

We employ the metric of Schwarzschild space surrounded by quintessential matter to study the trajectories of test masses on the motion of a binary system. The results, which are ob…

cs.CL2023

Data-Centric Financial Large Language Models

Zhixuan Chu, Huaiyu Guo, Xinyuan Zhou +9

Large language models (LLMs) show promise for natural language tasks but struggle when applied directly to complex domains like finance. LLMs have difficulty reasoning about and in…

cs.LG2023

A Better Match for Drivers and Riders: Reinforcement Learning at Lyft

Xabi Azagirre, Akshay Balwally, Guillaume Candeli +16

To better match drivers to riders in our ridesharing application, we revised Lyft's core matching algorithm. We use a novel online reinforcement learning approach that estimates th…

cs.AI2025

Robust Search with Uncertainty-Aware Value Models for Language Model Reasoning

Fei Yu, Yingru Li, Benyou Wang

Value model guided search is effective in steering LLM generation but suffers from a lack of robustness. This is due to verifier failure: imperfect VMs mistakenly prune valid reaso…

cs.CV2026

M3SR: Multi-Scale Multi-Perceptual Mamba for Efficient Spectral Reconstruction

Yuze Zhang, Lingjie Li, Qiuzhen Lin +3

The Mamba architecture has been widely applied to various low-level vision tasks due to its exceptional adaptability and strong performance. Although the Mamba architecture has bee…

cs.LG2026

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Jiacheng Chen, Xinyu Zhang, Shunkai Zhang +20

We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabili…

cs.CV2026

From Orbit to Ground: Generative City Photogrammetry from Extreme Off-Nadir Satellite Images

Fei Yu, Yu Liu, Luyang Tang +10

City-scale 3D reconstruction from satellite imagery presents the challenge of extreme viewpoint extrapolation, where our goal is to synthesize ground-level novel views from sparse…

cs.MM2024

DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis

Fa-Ting Hong, Yunfei Liu, Yu Li +3

Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generaliza…

cs.CL2025

A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models

Jinyi Han, Xinyi Wang, Haiquan Zhao +9

Recent advances in self-refinement have demonstrated significant potential for improving the outputs of large language models (LLMs) through iterative refinement. However, most exi…

eess.SP2025

Scene Understanding Enabled Semantic Communication with Open Channel Coding

Zhe Xiang, Fei Yu, Quan Deng +2

As communication systems transition from symbol transmission to conveying meaningful information, sixth-generation (6G) networks emphasize semantic communication. This approach pri…

cs.CV2023

MODA: Mapping-Once Audio-driven Portrait Animation with Dual Attentions

Yunfei Liu, Lijian Lin, Fei Yu +2

Audio-driven portrait animation aims to synthesize portrait videos that are conditioned by given audio. Animating high-fidelity and multimodal video portraits has a variety of appl…

cs.CL2026

Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs

Chenxing Wei, Hong Wang, Ying He +2

Large Language Models (LLMs) employ multi-turn interaction as a fundamental paradigm for completing complex tasks. However, their performance often degrades in extended interaction…

cs.CL2025

Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning

Fei Yu, Yingru Li, Benyou Wang

Large language models (LLMs) struggle with multi-step reasoning, where inference-time scaling has emerged as a promising strategy for performance improvement. Verifier-guided searc…

cs.CV2025

Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning

Yifan Xie, Fei Ma, Yi Bin +2

Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the signific…

math.AG2023

Filtration and splitting of the Hodge bundle on the non-varying strata of quadratic differentials

Dawei Chen, Fei Yu

We describe the Harder--Narasimhan filtration of the Hodge bundle for Teichmüller curves in the non-varying strata of quadratic differentials appearing in [CM2]. Moreover, we show…

cs.CV2021

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

Fei Yu, Jiji Tang, Weichong Yin +4

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL…

math.AG2016

Eigenvalues of Curvature, Lyapunov exponents and Harder-Narasimhan filtrations

Fei Yu

Inspired by Katz-Mazur theorem on crystalline cohomology and by Eskin-Kontsevich-Zorich's numerical experiments, we conjecture that the polygon of Lyapunov spectrum lies above (or…

physics.optics2021

Photoionization-induced broadband dispersive wave generated in an Ar-filled hollow-core photonic crystal fiber

Jianhua Fu, Yifei Chen, Zhiyuan Huang +7

The resonance band in hollow-core photonic crystal fiber (HC-PCF), while leading to high-loss region in the fiber transmission spectrum, has been successfully used for generating p…

cs.CV2026

ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space

Mingchao Sun, Luyang Tang, Yu Liu +34

The paper introduces ABot-3DWorld 0, a multimodal system that converts text, images, or video into high‑fidelity, explorable 3D worlds using a compact spatial representation and pa…

#3d reconstruction#multimodal generation#panoramic video#spatial world modeling
cs.CV2025

InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition

Junxiao Xue, Xiaozhen Liu, Xuecheng Wu +2

Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences…

cs.AI2024

OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Fei Yu, Anningzhe Gao, Benyou Wang

Large language models (LLMs) often struggle with maintaining accuracy throughout multiple multiple reasoning steps, especially in mathematical reasoning where an error in earlier s…

cs.LG2024

A Huber Loss Minimization Approach to Byzantine Robust Federated Learning

Puning Zhao, Fei Yu, Zhiguo Wan

Federated learning systems are susceptible to adversarial attacks. To combat this, we introduce a novel aggregator based on Huber loss minimization, and provide a comprehensive the…

gr-qc2008

Real Scalar Field Scattering with Polynomial Approximation around Schwarzschild-de Sitter Black-hole

Molin Liu, Hongya Liu, Jingfei Zhang +1

As one of the fitting methods, the polynomial approximation is effective to process sophisticated problem. In this paper, we employ this approach to handle the scattering of scalar…

cs.CV2025

Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM

Junxiao Xue, Quan Deng, Xuecheng Wu +7

Remote sensing change understanding (RSCU) is essential for analyzing remote sensing images and understanding how human activities affect the environment. However, existing dataset…

cs.CV2026

ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance

Dongfu Yin, Rourou Su, Cong Zhao +1

The paper introduces a method that removes asymmetric background clutter and enforces rotation-equivariant feature matching to improve detection of reflection symmetry in images, a…

#reflection symmetry detection#rotation equivariance#asymmetric denoising#deep learning
cs.AI2026

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

Zishang Jiang, Jinyi Han, Tingyun Li +7

Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large Language Models (LLMs). However, the effect…

physics.optics2023

Delivery of nanosecond laser pulses by multimode anti-resonant hollow core fiber at 1 um wavelength

Meng Zhao, Fei Yu, Dakun Wu +8

In this paper we explore the application of low-loss multimode anti-resonant hollow-core fiber (MM-AR-HCF) in the delivery of nanosecond laser pulses at 1 um wavelength. MM-AR-HCF…

cs.AI2026

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

Xianrui Zeng, Pengfei Liu, Yirui Zang +5

The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and generate hypotheses, most remain procedura…

eess.IV2021

BEFD: Boundary Enhancement and Feature Denoising for Vessel Segmentation

Mo Zhang, Fei Yu, Jie Zhao +2

Blood vessel segmentation is crucial for many diagnostic and research applications. In recent years, CNN-based models have leaded to breakthroughs in the task of segmentation, howe…

cs.CV2021

Unsupervised Domain Adaptation in Semantic Segmentation Based on Pixel Alignment and Self-Training

Hexin Dong, Fei Yu, Jie Zhao +2

This paper proposes an unsupervised cross-modality domain adaptation approach based on pixel alignment and self-training. Pixel alignment transfers ceT1 scans to hrT2 modality, hel…

astro-ph.CO2010

A more general interacting model of holographic dark energy

Fei Yu, Jingfei Zhang, Jianbo Lu +2

So far, there have been no theories or observational data that deny the presence of interaction between dark energy and dark matter. We extend naturally the holographic dark energy…

eess.IV2019

PGU-net+: Progressive Growing of U-net+ for Automated Cervical Nuclei Segmentation

Jie Zhao, Lei Dai, Mo Zhang +5

Automated cervical nucleus segmentation based on deep learning can effectively improve the quantitative analysis of cervical cancer. However, accurate nuclei segmentation is still…

cs.LG2026

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts

Xinyi Wang, Jinyi Han, Zishang Jiang +7

Reinforcement Learning (RL) has become a key driver for enhancing the long chain-of-thought (CoT) reasoning capabilities of Large Language Models (LLMs). However, prevalent methods…

math.AG2014

Weierstrass filtration on Teichmuller curves and Lyapunov exponents

Fei Yu, Kang Zuo

We define the Weierstrass filtration for Teichmuller curves and construct the Harder-Narasimhan filtration of the Hodge bundle of a Teichmuller curve in hyperelliptic loci and low-…

cs.SI2023

Stochastic Step-wise Feature Selection for Exponential Random Graph Models (ERGMs)

Helal El-Zaatari, Fei Yu, Michael R Kosorok

Statistical analysis of social networks provides valuable insights into complex network interactions across various scientific disciplines. However, accurate modeling of networks r…

astro-ph.CO2015

Statefinder hierarchy exploration of the extended Ricci dark energy

Fei Yu, Jing-Lei Cui, Jing-Fei Zhang +1

We apply the statefinder hierarchy plus the fractional growth parameter to explore the extended Ricci dark energy (ERDE) model, in which there are two independent coefficients

math.AG2026

An algebro-geometric perspective on the topology of moduli spaces of differentials

Dawei Chen, Fei Yu

Differentials on Riemann surfaces correspond to translation surfaces with conical singularities, and affine transformations acting on them preserve the orders of these singularitie…

cs.AI2026

Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking

Jinyi Han, Ying Huang, Ying Liao +11

Large Reasoning Models (LRMs) have achieved impressive performance on challenging tasks, yet their deep reasoning often incurs substantial computational costs. To achieve efficient…

stat.AP2014

Scalable Privacy-Preserving Data Sharing Methodology for Genome-Wide Association Studies

Fei Yu, Stephen E. Fienberg, Aleksandra Slavković +1

The protection of privacy of individual-level information in genome-wide association study (GWAS) databases has been a major concern of researchers following the publication of "an…

cs.AI2026

LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models

Qingyu Ren, Qianyu He, Jingwen Chang +9

Instruction following is critical for large language models, yet real-world instructions often involve multiple constraints with logical structures, such as parallel composition, s…

astro-ph.CO2018

Exploring interacting holographic dark energy in a perturbed universe with parameterized post-Friedmann approach

Lu Feng, Yun-He Li, Fei Yu +2

The model of holographic dark energy in which dark energy interacts with dark matter is investigated in this paper. In particular, we consider the interacting holographic dark ener…

physics.optics2022

Measurements of microjoule-level, few-femtosecond ultraviolet dispersive-wave pulses generated in gas-filled hollow capillary fibers

Cheng Zhang, Tiandao Chen, Jinyu Pan +11

High-energy ultraviolet pulse generation in gas-filled hollow capillary fibers (HCFs) through dispersive-wave-emission process, has attracted considerable attentions in recent seve…

cs.CV2026

TextSculptor: Training and Benchmarking Scene Text Editing

Yiheng Lin, Siyu Jiao, Xiaohan Lan +12

Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editi…

physics.optics2026

Physics-guided foundation model for universal speckle removal in ultrathin multimode fiber imaging

Xianrui Zeng, Yirui Zang, Pengfei Liu +4

Ultrathin multimode fibers (MMFs) promise endoscopes with hair-scale diameters for accessing sub-millimeter anatomy, but in MMF far-field imaging the required small collection aper…

cs.MM2025

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition

Junxiao Xue, Xiaozhen Liu, Xuecheng Wu +3

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the u…

cs.CV2025

MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach

Xin Zhang, Siting Huang, Xiangyang Luo +5

Face editing modifies the appearance of face, which plays a key role in customization and enhancement of personal images. Although much work have achieved remarkable success in tex…

cs.LG2025

ReDit: Reward Dithering for Improved LLM Policy Optimization

Chenxing Wei, Jiarui Yu, Ying Tiffany He +3

DeepSeek-R1 has successfully enhanced Large Language Model (LLM) reasoning capabilities through its rule-based reward system. While it's a ''perfect'' reward system that effectivel…

cs.CV2026

Recovering 3D Shapes from Ultra-Fast Motion-Blurred Images

Fei Yu, Shudan Guo, Shiqing Xin +3

We consider the problem of 3D shape recovery from ultra-fast motion-blurred images. While 3D reconstruction from static images has been extensively studied, recovering geometry fro…

cs.CL2025

Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation

Sirui Xia, Xintao Wang, Jiaqing Liang +5

Retrieval-Augmented Generation (RAG) has been widely adopted to enhance Large Language Models (LLMs) in knowledge-intensive tasks. To enhance credibility and verifiability in RAG s…

math.AG2013

A new inequality on the Hodge number of algebraic surfaces

Jun Lu, Sheng-Li Tan, Fei Yu +1

We get a new inequality on the Hodge number of fibred algebraic complex surfaces , which is a generalization of an inequality of Beauville. Our inequality implies t…

cs.CV2026

ABot-Earth 0.5: Generative 3D Earth Model

Ming Qian, Tianjian Ouyang, Mingchao Sun +25

We present ABot-Earth 0.5, a generative 3D framework designed to synthesize vast, seamless 3D environments from ubiquitous, geospatially referenced satellite imagery. To achieve th…

math.NA2019

Higher-order accurate diffuse-domain methods for partial differential equations with Dirichlet boundary conditions in complex, evolving geometries

Fei Yu, Zhenlin Guo, John Lowengrub

The diffuse-domain, or smoothed boundary, method is an attractive approach for solving partial differential equations in complex geometries because of its simplicity and flexibilit…

math.AG2016

Weierstrass filtration on Teichmüller curves and Lyapunov exponents: Upper bounds

Fei Yu, Kang Zuo

We get an upper bound of the slope of each graded quotient for the Harder-Narasimhan filtration of the Hodge bundle of a Teichmüller curve. As an application, we show that the sum…

cs.CL2026

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following

Qingyu Ren, Qianyu He, Powei Chang +5

Language models often struggle to follow multi-constraint instructions that are crucial for real-world applications. Existing reinforcement learning (RL) approaches suffer from dep…

gr-qc2007

Dirac quasinormal modes of a Schwarzschild black hole surrounded by free static spherically symmetric quintessence

Yu Zhang, Yuan-Xing Gui, Fei Yu

We evaluate the quasinormal modes of massless Dirac perturbation in a Schwarzschild black hole surrounded by the free static spherically symmetric quintessence by using the third-o…

cs.CL2024

Enhancing and Accelerating Large Language Models via Instruction-Aware Contextual Compression

Haowen Hou, Fei Ma, Binwen Bai +2

Large Language Models (LLMs) have garnered widespread attention due to their remarkable performance across various tasks. However, to mitigate the issue of hallucinations, LLMs oft…

physics.optics2023

Broadband Dispersive-Wave Emission Coupled with Two-Stage Soliton Self-Compression in Gas-Filled Anti-Resonant Hollow-Core Fibers

Jinyu Pan, Zhiyuan Huang, Yifei Chen +9

We studied the underlying mechanism of broadband dispersive-wave emission within a resonance band of gas-filled anti-resonant hollow-core fiber. Both theoretical and experimental r…

quant-ph2024

Method to deterministically generate large-amplitude optical cat states

Zheng-Hong Li, Fei Yu, Zhen-Ya Li +2

Cat states, as an important resource in the study of macroscopic quantum superposition and quantum information applications, have garnered widespread attention. To date, preparing…

cs.CL2025

Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation

Qianxi He, Qianyu He, Jiaqing Liang +4

Logical reasoning is essential for large language models (LLMs) to ensure accurate and coherent inference. However, LLMs struggle with reasoning order variations and fail to genera…

stat.ML2014

Differentially-Private Logistic Regression for Detecting Multiple-SNP Association in GWAS Databases

Fei Yu, Michal Rybar, Caroline Uhler +1

Following the publication of an attack on genome-wide association studies (GWAS) data proposed by Homer et al., considerable attention has been given to developing methods for rele…

cs.CV2025

Enhancing Long Video Question Answering with Scene-Localized Frame Grouping

Xuyi Yang, Wenhao Zhang, Hongbo Jin +5

Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video…

math.AG2015

On Okounkov's conjecture connecting Hilbert schemes of points and multiple q-zeta values

Zhenbo Qin, Fei Yu

We compute the generating series for the intersection pairings between the total Chern classes of the tangent bundles of the Hilbert schemes of points on a smooth projective surfac…

cs.AI2024

PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods

Yiying Wang, Xiaojing Li, Binzhu Wang +10

In domain-specific applications, GPT-4, augmented with precise prompts or Retrieval-Augmented Generation (RAG), shows notable potential but faces the critical tri-lemma of performa…

cs.CV2025

Uncertainty Quantification for Incomplete Multi-View Data Using Divergence Measures

Zhipeng Xue, Yan Zhang, Ming Li +3

Existing multi-view classification and clustering methods typically improve task accuracy by leveraging and fusing information from different views. However, ensuring the reliabili…

math.AG2012

Instability of Truncated Symmetric Powers of sheaves

Lingguang Li, Fei Yu

Let be a smooth projective variety of dimension over an algebraically closed field of characteristic . Let be the absolute Frobenius morphism,…

cs.CL2024

MileBench: Benchmarking MLLMs in Long Context

Dingjie Song, Shunian Chen, Guiming Hardy Chen +3

Despite the advancements and impressive performance of Multimodal Large Language Models (MLLMs) on benchmarks, their effectiveness in real-world, long-context, and multi-image task…

eess.IV2019

Annotation-Free Cardiac Vessel Segmentation via Knowledge Transfer from Retinal Images

Fei Yu, Jie Zhao, Yanjun Gong +6

Segmenting coronary arteries is challenging, as classic unsupervised methods fail to produce satisfactory results and modern supervised learning (deep learning) requires manual ann…

cs.CL2025

HERA: Improving Long Document Summarization using Large Language Models with Context Packaging and Reordering

Taiji Li, Hao Chen, Fei Yu +1

Despite the rapid growth of context length of large language models (LLMs) , LLMs still perform poorly in long document summarization. An important reason for this is that relevant…

cs.CV2026

What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs

Jiaping Lin, Fei Shen, Junzhe Li +4

Existing training-free approaches for GUI grounding often rely on multiple inference runs, such as iterative cropping or candidate aggregation, to identify target elements. Despite…

gr-qc2007

Quasinormal modes of a Schwarzschild black hole surrounded by free static spherically symmetric quintessence: Electromagnetic perturbations

Yu Zhang, Yuanxing Gui, Fei Yu +1

In this paper, we evaluated the quasinormal modes of electromagnetic perturbation in a Schwarzschild black hole surrounded by the static spherically symmetric quintessence by using…

cs.SI2019

Attentive Geo-Social Group Recommendation

Fei Yu, Feiyi Fan, Shouxu Jiang +1

Social activities play an important role in people's daily life since they interact. For recommendations based on social activities, it is vital to have not only the activity infor…

cond-mat.mtrl-sci2023

Chiral Topological superconductivity in the OAI/SC/FMI heterostructure avoiding the subband problem

Jingnan Hu, Fei Yu, Aiyun Luo +3

Implementing topological superconductivity (TSC) and Majorana states (MSs) is one of the most significant and challenging tasks in both fundamental physics and topological quantum…

physics.optics2023

Generation and characterization of frequency tuneable sub-15 fs pulses in a gas-filled hollow-core fiber pumped by a Yb:KGW laser

Mohammed Sabbah, Federico Belli, Christian Brahms +3

We investigate soliton self-compression and photoionization effects in an argon-filled antiresonant hollow-core photonic crystal fiber pumped with a commercial Yb:KGW laser. Before…

cs.CV2025

Object Isolated Attention for Consistent Story Visualization

Xiangyang Luo, Junhao Cheng, Yifan Xie +5

Open-ended story visualization is a challenging task that involves generating coherent image sequences from a given storyline. One of the main difficulties is maintaining character…

cs.AI2025

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following

Qingyu Ren, Qianyu He, Bowei Zhang +6

Reasoning models excel in complex problem solving but exhibit a concerning trade off between reasoning capabilities and instruction following abilities. Existing approaches for imp…

astro-ph.CO2013

Statefinder diagnosis for the extended holographic Ricci dark energy model without and with interaction

Fei Yu, Jing-Fei Zhang

We apply the statefinder diagnostic to the extended holographic Ricci dark energy (ERDE) model without and with interaction to study their behaviors. We plot the trajectories of va…

physics.optics2025

A power-in-bucket model enabled designs of nanostructure-enhanced waveguides for highly efficient wide-angle light couplings

Wenbo Luo, Yitong Gu, Jianwei Wang +5

Well-designed nanostructures on fiber facets can boost wide-angle light coupling and thus gain considerable attention because of the potential for intensive applications. However,…

cs.CL2025

Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models

Qingyu Ren, Jie Zeng, Qianyu He +5

It is crucial for large language models (LLMs) to follow instructions that involve multiple constraints. However, it is an unexplored area to enhance LLMs' ability to follow soft c…

cs.IR2015

Network-based recommendation algorithms: A review

Fei Yu, An Zeng, Sebastien Gillard +1

Recommender systems are a vital tool that helps us to overcome the information overload problem. They are being used by most e-commerce web sites and attract the interest of a broa…

cs.SD2024

Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization

Fei Yu, Zhe Xiang, Nan Che +4

Multimodal semantic communication, which integrates various data modalities such as text, images, and audio, significantly enhances communication efficiency and reliability. It has…

cs.AI2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

MiniMax, :, Aili Chen +219

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

cs.SI2025

Topic-aware Most Influential Community Search in Social Networks

Long Teng, Yanhao Wang, Zhe Lin +1

Influential community search (ICS) finds a set of densely connected and high-impact vertices from a social network. Although great effort has been devoted to ICS problems, most exi…

cs.CV2025

UniSync: A Unified Framework for Audio-Visual Synchronization

Tao Feng, Yifan Xie, Xun Guan +4

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this cha…

cs.AI2025

AgentBay: A Hybrid Interaction Sandbox for Seamless Human-AI Intervention in Agentic Systems

Yun Piao, Hongbo Min, Hang Su +28

The rapid advancement of Large Language Models (LLMs) is catalyzing a shift towards autonomous AI Agents capable of executing complex, multi-step tasks. However, these agents remai…

cs.CL2023

Natural Language Reasoning, A Survey

Fei Yu, Hongbo Zhang, Prayag Tiwari +1

This survey paper proposes a clearer view of natural language reasoning in the field of Natural Language Processing (NLP), both conceptually and practically. Conceptually, we provi…