papers

Publications (19)

cs.RO2026

Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory Prediction

Hao Zhou, Lu Qi, Jason Li +5

Trajectory prediction is critical for autonomous driving, enabling safe and efficient planning in dense, dynamic traffic. Most existing methods optimize prediction accuracy under f…

cs.CV2023

DLCA-Recon: Dynamic Loose Clothing Avatar Reconstruction from Monocular Videos

Chunjie Luo, Fei Luo, Yusen Wang +2

Reconstructing a dynamic human with loose clothing is an important but difficult task. To address this challenge, we propose a method named DLCA-Recon to create human avatars from…

cs.CV2026

MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

Jie Zhang, Qilang Ye, Hao Zhou +2

The dominant paradigm in video retrieval relies on embedding-based full-corpus scanning, which suffers from inherent computational inefficiency and the semantic asymmetry between i…

astro-ph.CO2025

A Systematic Literature Review of Machine Learning Techniques for Observational Constraints in Cosmology

Luis Rojas, Sebastián Espinoza, Esteban González +2

This paper presents a systematic literature review focusing on the application of machine learning techniques for deriving observational constraints in cosmology. The goal is to ev…

cs.CV2018

RedNet: Residual Encoder-Decoder Network for indoor RGB-D Semantic Segmentation

Jindong Jiang, Lunan Zheng, Fei Luo +1

Indoor semantic segmentation has always been a difficult task in computer vision. In this paper, we propose an RGB-D residual encoder-decoder architecture, named RedNet, for indoor…

physics.ins-det2014

Neutron Time-Of-Flight Spectrometer Based on HIRFL for Studies of Spallation Reactions Related to ADS Project

Suyalatu Zhang, Zhiqiang Chen, Rui Han +8

A Neutron Time-Of-Flight (NTOF) spectrometer based on Heavy Ion Research Facility in Lanzhou (HIRFL) is developed for studies of neutron production of proton induced spallation rea…

cs.LG2026

Reasoning emerges from constrained inference manifolds in large language models

Yanbiao Ma, Fei Luo, Linfeng Zhang +10

Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inference. Here we study reasonin…

cs.CV2026

Wi-Spike: A Low-power WiFi Human Multi-action Recognition Model with Spiking Neural Networks

Nengbo Zhang, Yao Ying, Lu Wang +3

WiFi-based human action recognition (HAR) has gained significant attention due to its non-intrusive and privacy-preserving nature. However, most existing WiFi sensing models predom…

cs.AI2026

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning

Zhe Qian, Nianbing Su, Zhonghua Wang +6

Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this limitation, we propose Self…

cs.CR2026

Argus: Reorchestrating Static Analysis via a Multi-Agent Ensemble for Full-Chain Security Vulnerability Detection

Zi Liang, Qipeng Xie, Jun He +7

Recent advancements in Large Language Models (LLMs) have sparked interest in their application to Static Application Security Testing (SAST), primarily due to their superior contex…

cs.LG2025

Geometric Prior-Guided Federated Prompt Calibration

Fei Luo, Ziwei Zhao, Mingxuan Wang +5

Federated Prompt Learning (FPL) offers a parameter-efficient solution for collaboratively training large models, but its performance is severely hindered by data heterogeneity, whi…

cs.AI2026

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models

Zhe Qian, Yanbiao Ma, Zhuohan Ouyang +7

Multimodal Large Reasoning Models (MLRMs) have achieved remarkable strides in visual reasoning through test time compute scaling, yet long chain reasoning remains prone to hallucin…

cs.CV2022

DGECN: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation

Tuo Cao, Fei Luo, Yanping Fu +3

Monocular 6D pose estimation is a fundamental task in computer vision. Existing works often adopt a two-stage pipeline by establishing correspondences and utilizing a RANSAC algori…

cs.LG2026

FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices

Changyu Li, Shuanghong Huang, Jiashen Liu +5

Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data. However, in mobile deployments, the traini…

cs.LG2026

Reliability-Calibrated Edge-IoT Early Fault Warning for Rotating Machinery with a Physics-Guided Tiny-Mamba Transformer

Changyu Li, Huabei Nie, Xiaoya Ni +4

Industrial Internet of Things (IIoT) systems increasingly rely on distributed vibration sensing to support predictive maintenance of rotating machinery. In practical deployments, h…

cs.AI2026

PI-TTA: Physics-Informed Source-Free Test-Time Adaptation for Robust Human Activity Recognition on Mobile Devices

Changyu Li, Lu Wang, Ming Lei +4

Source-free test-time adaptation (TTA) is appealing for mobile and wearable sensing because it enables on-device personalization from unlabeled test streams without centralizing pr…

cs.CV2025

CardiacMamba: A Multimodal RGB-RF Fusion Framework with State Space Models for Remote Physiological Measurement

Zheng Wu, Yiping Xie, Bo Zhao +4

Heart rate (HR) estimation via remote photoplethysmography (rPPG) offers a non-invasive solution for health monitoring. However, traditional single-modality approaches (RGB or Radi…

cs.CV2023

NeTO:Neural Reconstruction of Transparent Objects with Self-Occlusion Aware Refraction-Tracing

Zongcheng Li, Xiaoxiao Long, Yusen Wang +4

We present a novel method, called NeTO, for capturing 3D geometry of solid transparent objects from 2D images via volume rendering. Reconstructing transparent objects is a very cha…

cs.AI2025

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

Chenyue Zhou, Mingxuan Wang, Yanbiao Ma +19

Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incohere…