papers

Publications (204)

cs.GR2022

Artemis: Articulated Neural Pets with Appearance and Motion synthesis

Haimin Luo, Teng Xu, Yuheng Jiang +6

We, humans, are entering into a virtual era and indeed want to bring animals to the virtual world as well for companion. Yet, computer-generated (CGI) furry animals are limited by…

cs.CV2023

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

Hanzhuo Huang, Yufan Feng, Cheng Shi +3

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text pr…

cs.CV2021

MirrorNeRF: One-shot Neural Portrait Radiance Field from Multi-mirror Catadioptric Imaging

Ziyu Wang, Liao Wang, Fuqiang Zhao +3

Photo-realistic neural reconstruction and rendering of the human portrait are critical for numerous VR/AR applications. Still, existing solutions inherently rely on multi-view capt…

cs.CV2025

Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability

Yingdong Shi, Changming Li, Yifan Wang +5

Diffusion models have demonstrated impressive capabilities in synthesizing diverse content. However, despite their high-quality outputs, these models often perpetuate social biases…

cs.CV2023

HiFi4G: High-Fidelity Human Performance Rendering via Compact Gaussian Splatting

Yuheng Jiang, Zhehao Shen, Penghao Wang +5

We have recently seen tremendous progress in photo-real human modeling and rendering. Yet, efficiently rendering realistic human performance and integrating it into the rasterizati…

cs.CV2016

Occlusion-Model Guided Anti-Occlusion Depth Estimation in Light Field

Hao Zhu, Qing Wang, Jingyi Yu

Occlusion is one of the most challenging problems in depth estimation. Previous work has modeled the single-occluder occlusion in light field and get good results, however it is st…

cs.CV2025

Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis

Kaiyang Ji, Ye Shi, Zichen Jin +5

Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate pr…

cs.CV2024

OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers

Han Liang, Jiacheng Bao, Ruichi Zhang +6

We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods often fail or produce implausible motions with unseen text inputs, which…

cs.CV2015

Automatic Layer Separation using Light Field Imaging

Qiaosong Wang, Haiting Lin, Yi Ma +2

We propose a novel approach that jointly removes reflection or translucent layer from a scene and estimates scene depth. The input data are captured via light field imaging. The pr…

eess.IV2023

Disentangling Light Fields for Super-Resolution and Disparity Estimation

Yingqian Wang, Longguang Wang, Gaochang Wu +4

Light field (LF) cameras record both intensity and directions of light rays, and encode 3D scenes into 4D LF images. Recently, many convolutional neural networks (CNNs) have been p…

cs.CV2023

Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos

Liao Wang, Qiang Hu, Qihan He +5

The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has inspired numerous attempts on dynamic scenes. Current techniques that util…

physics.ins-det2014

A Precise Calculation of Delayed Coincidence Selection Efficiency and Accidental Coincidence Rate

Jingyi Yu, Zhe Wang, Shaomin Chen

A model is proposed to address issues on the precise background evaluation due to the complex data structure defined by the delayed coincidence method, which is widely used in reac…

cs.GR2026

SPREAD: Spatial-Physical REasoning via geometry Aware Diffusion

Minzhang Li, Kuixiang Shao, Xuebing Li +6

Automated 3D scene generation is pivotal for applications spanning virtual reality, digital content creation, and Embodied AI. While computer graphics prioritizes aesthetic layouts…

cs.RO2026

ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K

Kaixuan Wang, Tianxing Chen, Jiawei Liu +13

Learning in simulation provides a useful foundation for scaling robotic manipulation capabilities. However, this paradigm often suffers from a lack of data-generation-ready digital…

cs.GR2025

BANG: Dividing 3D Assets via Generative Exploded Dynamics

Longwen Zhang, Qixuan Zhang, Haoran Jiang +4

3D creation has always been a unique human strength, driven by our ability to deconstruct and reassemble objects using our eyes, mind and hand. However, current 3D design tools str…

cs.CV2023

Omni-Line-of-Sight Imaging for Holistic Shape Reconstruction

Binbin Huang, Xingyue Peng, Siyuan Shen +8

We introduce Omni-LOS, a neural computational imaging method for conducting holistic shape reconstruction (HSR) of complex objects utilizing a Single-Photon Avalanche Diode (SPAD)-…

eess.IV2021

IREM: High-Resolution Magnetic Resonance (MR) Image Reconstruction via Implicit Neural Representation

Qing Wu, Yuwei Li, Lan Xu +7

For collecting high-quality high-resolution (HR) MR image, we propose a novel image reconstruction network named IREM, which is trained on multiple low-resolution (LR) MR images an…

cs.GR2025

BG-Triangle: Bézier Gaussian Triangle for 3D Vectorization and Rendering

Minye Wu, Haizhao Dai, Kaixin Yao +2

Differentiable rendering enables efficient optimization by allowing gradients to be computed through the rendering process, facilitating 3D reconstruction, inverse rendering and ne…

cs.CV2022

LiDARCap: Long-range Marker-less 3D Human Motion Capture with LiDAR Point Clouds

Jialian Li, Jingyi Zhang, Zhiyong Wang +6

Existing motion capture datasets are largely short-range and cannot yet fit the need of long-range applications. We propose LiDARHuman26M, a new human motion capture dataset captur…

cs.RO2024

Implicit Swept Volume SDF: Enabling Continuous Collision-Free Trajectory Generation for Arbitrary Shapes

Jingping Wang, Tingrui Zhang, Qixuan Zhang +5

In the field of trajectory generation for objects, ensuring continuous collision-free motion remains a huge challenge, especially for non-convex geometries and complex environments…

physics.app-ph2023

Single-pixel p-graded-n junction spectrometers

Jingyi Wang, Beibei Pan, Zi Wang +8

Ultra-compact spectrometers are becoming increasingly popular for their promising applications in biomedical analysis, environmental monitoring, and food safety. In this work, we r…

cs.CV2021

Neural Video Portrait Relighting in Real-time via Consistency Modeling

Longwen Zhang, Qixuan Zhang, Minye Wu +2

Video portraits relighting is critical in user-facing human photography, especially for immersive VR/AR experience. Recent advances still fail to recover consistent relit result un…

cs.CV2021

PIANO: A Parametric Hand Bone Model from Magnetic Resonance Imaging

Yuwei Li, Minye Wu, Yuyao Zhang +2

Hand modeling is critical for immersive VR/AR, action understanding, or human healthcare. Existing parametric models account only for hand shape, pose, or texture, without modeling…

cs.CV2020

A Neural Rendering Framework for Free-Viewpoint Relighting

Zhang Chen, Anpei Chen, Guli Zhang +4

We present a novel Relightable Neural Renderer (RNR) for simultaneous view synthesis and relighting using multi-view image inputs. Existing neural rendering (NR) does not explicitl…

eess.IV2024

Sparse-view Signal-domain Photoacoustic Tomography Reconstruction Method Based on Neural Representation

Bowei Yao, Yi Zeng, Haizhao Dai +6

Photoacoustic tomography is a hybrid biomedical technology, which combines the advantages of acoustic and optical imaging. However, for the conventional image reconstruction method…

cs.CV2019

Generic Multiview Visual Tracking

Minye Wu, Haibin Ling, Ning Bi +3

Recent progresses in visual tracking have greatly improved the tracking performance. However, challenges such as occlusion and view change remain obstacles in real world deployment…

cs.CV2024

A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals

Jiangnan Tang, Jingya Wang, Kaiyang Ji +3

Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest c…

cs.CV2021

TightCap: 3D Human Shape Capture with Clothing Tightness Field

Xin Chen, Anqi Pang, Yang Wei +2

In this paper, we present TightCap, a data-driven scheme to capture both the human shape and dressed garments accurately with only a single 3D human scan, which enables numerous ap…

cs.CV2024

SMGDiff: Soccer Motion Generation using diffusion probabilistic models

Hongdi Yang, Chengyang Li, Zhenxuan Wu +5

Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate in…

cs.CV2021

Editable Free-viewpoint Video Using a Layered Neural Representation

Jiakai Zhang, Xinhang Liu, Xinyi Ye +6

Generating free-viewpoint videos is critical for immersive VR/AR experience but recent neural advances still lack the editing ability to manipulate the visual perception for large…

cs.CV2018

A Generic Multi-Projection-Center Model and Calibration Method for Light Field Cameras

Qi Zhang, Chunping Zhang, Jinbo Ling +2

Light field cameras can capture both spatial and angular information of light rays, enabling 3D reconstruction by a single exposure. The geometry of 3D reconstruction is affected b…

cs.CV2017

Robust Guided Image Filtering

Wei Liu, Xiaogang Chen, Chunhua Shen +3

The process of using one image to guide the filtering process of another one is called Guided Image Filtering (GIF). The main challenge of GIF is the structure inconsistency betwee…

cs.CV2026

PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery

Yijing Guo, Mengjun Chao, Luo Wang +5

Panoramic imagery offers a full 360° field of view and is increasingly common in consumer devices. However, it introduces non-pinhole distortions that challenge joint pose estimat…

cs.CV2025

UniDB: A Unified Diffusion Bridge Framework via Stochastic Optimal Control

Kaizhen Zhu, Mokai Pan, Yuexin Ma +4

Recent advances in diffusion bridge models leverage Doob's -transform to establish fixed endpoints between distributions, demonstrating promising results in image translation an…

cond-mat.soft2025

A scaling law for large-deformation contact in soft materials

Tong Mu, Shizhuo Weng, Changhong Linghu +10

Compression of soft bodies is central to biology, materials science, and robotics, yet existing contact theories break down at large deformations. Here, we develop a general framew…

cs.LG2026

FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation

Weichen Qin, Yufan Xie, Peihao Wang +8

Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. However, existing methods strugg…

cs.CV2023

NeMF: Inverse Volume Rendering with Neural Microflake Field

Youjia Zhang, Teng Xu, Junqing Yu +5

Recovering the physical attributes of an object's appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering. Rece…

cs.RO2025

ExFace: Expressive Facial Control for Humanoid Robots with Diffusion Transformers and Bootstrap Training

Dong Zhang, Jingwei Peng, Yuyang Jiao +3

This paper presents a novel Expressive Facial Control (ExFace) method based on Diffusion Transformers, which achieves precise mapping from human facial blendshapes to bionic robot…

cs.CV2023

IKOL: Inverse kinematics optimization layer for 3D human pose and shape estimation via Gauss-Newton differentiation

Juze Zhang, Ye Shi, Yuexin Ma +3

This paper presents an inverse kinematic optimization layer (IKOL) for 3D human pose and shape estimation that leverages the strength of both optimization- and regression-based met…

cs.CV2022

TensoRF: Tensorial Radiance Fields

Anpei Chen, Zexiang Xu, Andreas Geiger +2

We present TensoRF, a novel approach to model and reconstruct radiance fields. Unlike NeRF that purely uses MLPs, we model the radiance field of a scene as a 4D tensor, which repre…

cs.CV2026

CylinderSplat: 3D Gaussian Splatting with Cylindrical Triplanes for Panoramic Novel View Synthesis

Qiwei Wang, Xianghui Ze, Jingyi Yu +1

Feed-forward 3D Gaussian Splatting (3DGS) has shown great promise for real-time novel view synthesis, but its application to panoramic imagery remains challenging. Existing methods…

cs.CV2024

CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets

Longwen Zhang, Ziyu Wang, Qixuan Zhang +6

In the realm of digital creativity, our potential to craft intricate 3D worlds from imagination is often hampered by the limitations of existing digital tools, which demand extensi…

cs.CV2023

HybridCap: Inertia-aid Monocular Capture of Challenging Human Motions

Han Liang, Yannan He, Chengfeng Zhao +4

Monocular 3D motion capture (mocap) is beneficial to many applications. The use of a single camera, however, often fails to handle occlusions of different body parts and hence it i…

cs.CV2024

AerialGo: Walking-through City View Generation from Aerial Perspectives

Fuqiang Zhao, Yijing Guo, Siyuan Yang +6

High-quality 3D urban reconstruction is essential for applications in urban planning, navigation, and AR/VR. However, capturing detailed ground-level data across cities is both lab…

cs.CV2023

Relightable Neural Human Assets from Multi-view Gradient Illuminations

Taotao Zhou, Kai He, Di Wu +6

Human modeling and relighting are two fundamental problems in computer vision and graphics, where high-quality datasets can largely facilitate related research. However, most exist…

cs.CV2015

Resolving Scale Ambiguity Via XSlit Aspect Ratio Analysis

Wei Yang, Haiting Lin, Sing Bing Kang +1

In perspective cameras, images of a frontal-parallel 3D object preserve its aspect ratio invariant to its depth. Such an invariance is useful in photography but is unique to perspe…

cs.CV2025

DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness

Yiming Zhong, Qi Jiang, Jingyi Yu +1

A dexterous hand capable of grasping any object is essential for the development of general-purpose embodied intelligent robots. However, due to the high degree of freedom in dexte…

cs.CV2018

Personalized Saliency and its Prediction

Yanyu Xu, Shenghua Gao, Junru Wu +2

Nearly all existing visual saliency models by far have focused on predicting a universal saliency map across all observers. Yet psychology studies suggest that visual attention of…

cs.CV2024

V^3: Viewing Volumetric Videos on Mobiles via Streamable 2D Dynamic Gaussians

Penghao Wang, Zhirui Zhang, Liao Wang +5

Experiencing high-fidelity volumetric video as seamlessly as 2D videos is a long-held dream. However, current dynamic 3DGS methods, despite their high rendering quality, face chall…

cs.CV2023

NeuRBF: A Neural Fields Representation with Adaptive Radial Basis Functions

Zhang Chen, Zhong Li, Liangchen Song +4

We present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for s…

cs.RO2025

DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot Handover

Youzhuo Wang, Jiayi Ye, Chuyang Xiao +6

Handover between a human and a dexterous robotic hand is a fundamental yet challenging task in human-robot collaboration. It requires handling dynamic environments and a wide varie…

eess.IV2025

SPIDER: Structure-Preferential Implicit Deep Network for Biplanar X-ray Reconstruction

Tianqi Yu, Xuanyu Tian, Jiawen Yang +4

Biplanar X-ray imaging is widely used in health screening, postoperative rehabilitation evaluation of orthopedic diseases, and injury surgery due to its rapid acquisition, low radi…

cs.CV2025

CADSpotting: Robust Panoptic Symbol Spotting on Large-Scale CAD Drawings

Fuyi Yang, Jiazuo Mu, Yanshun Zhang +7

We introduce CADSpotting, an effective method for panoptic symbol spotting in large-scale architectural CAD drawings. Existing approaches often struggle with symbol diversity, scal…

cs.GR2025

DreamPrinting: Volumetric Printing Primitives for High-Fidelity 3D Printing

Youjia Wang, Ruixiang Cao, Teng Xu +4

Translating the rich visual fidelity of volumetric rendering techniques into physically realizable 3D prints remains an open challenge. We introduce DreamPrinting, a novel pipeline…

cs.CV2022

NIMBLE: A Non-rigid Hand Model with Bones and Muscles

Yuwei Li, Longwen Zhang, Zesong Qiu +6

Emerging Metaverse applications demand reliable, accurate, and photorealistic reproductions of human hands to perform sophisticated operations as if in the physical world. While re…

cs.CV2022

Light Field-Based Underwater 3D Reconstruction Via Angular Resampling

Yuqi Ding, Zhang Chen, Yu Ji +2

Recovering 3D geometry of underwater scenes is challenging because of non-linear refraction of light at the water-air interface caused by the camera housing. We present a light fie…

cs.CV2026

Improving 2D Diffusion Models for 3D Medical Imaging with Inter-Slice Consistent Stochasticity

Chenhe Du, Qing Wu, Xuanyu Tian +3

3D medical imaging is in high demand and essential for clinical diagnosis and scientific research. Currently, diffusion models (DMs) have become an effective tool for medical imagi…

cs.LG2024

Guidance with Spherical Gaussian Constraint for Conditional Diffusion

Lingxiao Yang, Shutong Ding, Yifan Cai +3

Recent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training…

cs.CV2024

LiveHPS: LiDAR-based Scene-level Human Pose and Shape Estimation in Free Environment

Yiming Ren, Xiao Han, Chengfeng Zhao +4

For human-centric large-scale scenes, fine-grained modeling for 3D human global pose and shape is significant for scene understanding and can benefit many real-world applications.…

cs.AI2026

CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

Minzhang Li, Mingrui Li, Weichen Qin +5

Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity. Current solvers are often li…

cs.CV2024

ScalableMap: Scalable Map Learning for Online Long-Range Vectorized HD Map Construction

Jingyi Yu, Zizhao Zhang, Shengfu Xia +1

We propose a novel end-to-end pipeline for online long-range vectorized high-definition (HD) map construction using on-board camera sensors. The vectorized representation of HD map…

cs.GR2025

CityGo: Lightweight Urban Modeling and Rendering with Proxy Buildings and Residual Gaussians

Weihang Liu, Yuhui Zhong, Yuke Li +8

Accurate and efficient modeling of large-scale urban scenes is critical for applications such as AR navigation, UAV based inspection, and smart city digital twins. While aerial ima…

cs.CV2024

Capturing the Unseen: Vision-Free Facial Motion Capture Using Inertial Measurement Units

Youjia Wang, Yiwen Wu, Hengan Zhou +9

We present Capturing the Unseen (CAPUS), a novel facial motion capture (MoCap) technique that operates without visual signals. CAPUS leverages miniaturized Inertial Measurement Uni…

eess.IV2021

Non-line-of-Sight Imaging via Neural Transient Fields

Siyuan Shen, Zi Wang, Ping Liu +5

We present a neural modeling framework for Non-Line-of-Sight (NLOS) imaging. Previous solutions have sought to explicitly recover the 3D geometry (e.g., as point clouds) or voxel d…

cs.CV2018

Semantic See-Through Rendering on Light Fields

Huangjie Yu, Guli Zhang, Yuanxi Ma +2

We present a novel semantic light field (LF) refocusing technique that can achieve unprecedented see-through quality. Different from prior art, our semantic see-through (SST) diffe…

cs.LG2025

Causal Mechanism Estimation in Multi-Sensor Systems Across Multiple Domains

Jingyi Yu, Tim Pychynski, Marco F. Huber

To gain deeper insights into a complex sensor system through the lens of causality, we present common and individual causal mechanism estimation (CICME), a novel three-step approac…

cs.CV2022

Generative Deformable Radiance Fields for Disentangled Image Synthesis of Topology-Varying Objects

Ziyu Wang, Yu Deng, Jiaolong Yang +2

3D-aware generative models have demonstrated their superb performance to generate 3D neural radiance fields (NeRF) from a collection of monocular 2D images even for topology-varyin…

cs.GR2024

Robust Dual Gaussian Splatting for Immersive Human-centric Volumetric Videos

Yuheng Jiang, Zhehao Shen, Yu Hong +5

Volumetric video represents a transformative advancement in visual media, enabling users to freely navigate immersive virtual experiences and narrowing the gap between digital and…

cs.CV2022

NeuralHOFusion: Neural Volumetric Rendering under Human-object Interactions

Yuheng Jiang, Suyi Jiang, Guoxing Sun +5

4D modeling of human-object interactions is critical for numerous applications. However, efficient volumetric capture and rendering of complex interaction scenarios, especially fro…

cs.CV2026

Autoregressive B-Rep Shape Generation with Parametric Surfaces

Dafei Qin, Rui Xu, Zeyu Shen +8

Generative CAD modeling has broad design and application potential. Despite significant advances in Boundary Representation (B-Rep) generation, the dominant representation in CAD,…

cs.GR2025

BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video

Yu Hong, Yize Wu, Zhehao Shen +5

Volumetric video enables immersive experiences by capturing dynamic 3D scenes, enabling diverse applications for virtual reality, education, and telepresence. However, traditional…

cs.RO2025

AffordDP: Generalizable Diffusion Policy with Transferable Affordance

Shijie Wu, Yihang Zhu, Yunao Huang +5

Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance gene…

cs.GR2024

GaussianHair: Hair Modeling and Rendering with Light-aware Gaussians

Haimin Luo, Min Ouyang, Zijun Zhao +6

Hairstyle reflects culture and ethnicity at first glance. In the digital era, various realistic human hairstyles are also critical to high-fidelity digital human assets for beauty…

cs.CV2017

Hyperspectral Light Field Stereo Matching

Kang Zhu, Yujia Xue, Qiang Fu +3

In this paper, we describe how scene depth can be extracted using a hyperspectral light field capture (H-LF) system. Our H-LF system consists of a 5 x 6 array of cameras, with each…

cs.CV2026

TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos

Yan Zeng, Haoran Jiang, Kaixin Yao +4

Automatically generating photorealistic and self-consistent appearances for untextured 3D models is a critical challenge in digital content creation. The advancement of large-scale…

cs.CV2024

CryoGEM: Physics-Informed Generative Cryo-Electron Microscopy

Jiakai Zhang, Qihe Chen, Yan Zeng +4

In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains.…

cs.CV2026

ForeSplat: Optimization-Aware Foresight for Feed-Forward 3D Gaussian Splatting

Yuke Li, Weihang Liu, Cheng Zhang +8

Feed-forward 3D Gaussian Splatting models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is fundamentally hindered by the scarcity o…

cs.CV2023

MotionGPT: Human Motion as a Foreign Language

Biao Jiang, Xin Chen, Wen Liu +3

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains ch…

cs.CV2023

LiDAR-aid Inertial Poser: Large-scale Human Motion Capture by Sparse Inertial and LiDAR Sensors

Yiming Ren, Chengfeng Zhao, Yannan He +5

We propose a multi-sensor fusion method for capturing challenging 3D human motions with accurate consecutive local poses and global trajectories in large-scale scenarios, only usin…

cs.CV2024

Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance

Qingcheng Zhao, Pengyu Long, Qixuan Zhang +6

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality…

cs.HC2025

The Ephemeral Shadow: Hyperreal Beings in Stimulative Performance

Dong Zhang, Yanjun Zhou, Jingyi Yu

The Ephemeral Shadow is an interactive art installation centered on the concept of "simulacrum," focusing on the reconstruction of subjectivity at the intersection of reality and v…

physics.optics2023

Optical ReLU-like Activation Function Based on a Semiconductor Laser with Optical Injection

Guanting Liu, Yiwei Shen, Ruiqian Li +3

Artificial neural networks usually consist of successive linear multiply-accumulate operations and nonlinear activation functions. However, most optical neural networks only achiev…

cs.CV2019

3D Face Reconstruction Using Color Photometric Stereo with Uncalibrated Near Point Lights

Zhang Chen, Yu Ji, Mingyuan Zhou +2

We present a new color photometric stereo (CPS) method that recovers high quality, detailed 3D face geometry in a single shot. Our system uses three uncalibrated near point lights…

cs.CV2026

PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding

Penghao Wang, Yiyang He, Xin Lv +4

Understanding objects at the level of their constituent parts is fundamental to advancing computer vision, graphics, and robotics. While datasets like PartNet have driven progress…

cs.CV2025

CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image

Kaixin Yao, Longwen Zhang, Xinhao Yan +6

Recovering high-quality 3D scenes from a single RGB image is a challenging task in computer graphics. Current methods often struggle with domain-specific limitations or low-quality…

cs.CV2025

MouseGPT: A Large-scale Vision-Language Model for Mouse Behavior Analysis

Teng Xu, Taotao Zhou, Youjia Wang +12

Analyzing animal behavior is crucial in advancing neuroscience, yet quantifying and deciphering its intricate dynamics remains a significant challenge. Traditional machine vision a…

cs.CV2025

Solving Energy-Independent Density for CT Metal Artifact Reduction via Neural Representation

Qing Wu, Xu Guo, Lixuan Chen +8

X-ray CT often suffers from shadowing and streaking artifacts in the presence of metallic materials, which severely degrade imaging quality. Physically, the linear attenuation coef…

eess.IV2024

DPER: Diffusion Prior Driven Neural Representation for Limited Angle and Sparse View CT Reconstruction

Chenhe Du, Xiyue Lin, Qing Wu +9

Limited-angle and sparse-view computed tomography (LACT and SVCT) are crucial for expanding the scope of X-ray CT applications. However, they face challenges due to incomplete data…

cs.CV2022

SCULPTOR: Skeleton-Consistent Face Creation Using a Learned Parametric Generator

Zesong Qiu, Yuwei Li, Dongming He +8

Recent years have seen growing interest in 3D human faces modelling due to its wide applications in digital human, character generation and animation. Existing approaches overwhelm…

cs.CV2022

Weakly Supervised 3D Multi-person Pose Estimation for Large-scale Scenes based on Monocular Camera and Single LiDAR

Peishan Cong, Yiteng Xu, Yiming Ren +5

Depth estimation is usually ill-posed and ambiguous for monocular camera-based 3D multi-person pose estimation. Since LiDAR can capture accurate depth information in long-range sce…

cs.CV2023

CoTDet: Affordance Knowledge Prompting for Task Driven Object Detection

Jiajin Tang, Ge Zheng, Jingyi Yu +1

Task driven object detection aims to detect object instances suitable for affording a task in an image. Its challenge lies in object categories available for the task being too div…

cs.LG2026

Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs

Zhangzhi Xiong, Minzhang Li, Haotian Yu +6

Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budg…

cs.CV2024

SCOPE: Sign Language Contextual Processing with Embedding from LLMs

Yuqi Liu, Wenqian Zhang, Sihan Ren +3

Sign languages, used by around 70 million Deaf individuals globally, are visual languages that convey visual and contextual information. Current methods in vision-based sign langua…

cs.CL2026

UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering

Yingdong Shi, Ruiming Zhang, Changming Li +4

Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an effective paradigm for control…

cs.RO2026

Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary

Zhirui Liu, Kaiyang Ji, Ke Yang +4

Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing…

cs.CV2024

MeshXL: Neural Coordinate Field for Generative 3D Foundation Models

Sijin Chen, Xin Chen, Anqi Pang +11

The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, giv…

cs.CV2023

Human-centric Scene Understanding for 3D Large-scale Scenarios

Yiteng Xu, Peishan Cong, Yichen Yao +6

Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human…

cs.CV2022

Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time

Liao Wang, Jiakai Zhang, Xinhang Liu +6

Implicit neural representations such as Neural Radiance Field (NeRF) have focused mainly on modeling static objects captured under multi-view settings where real-time rendering can…

cs.CV2025

EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment

Yufei Zhu, Yiming Zhong, Zemin Yang +4

Dexterous robotic hands often struggle to generalize effectively in complex environments due to the limitations of models trained on low-diversity data. However, the real world pre…

cs.CV2018

Robust 3D Human Motion Reconstruction Via Dynamic Template Construction

Zhong Li, Yu Ji, Wei Yang +2

In multi-view human body capture systems, the recovered 3D geometry or even the acquired imagery data can be heavily corrupted due to occlusions, noise, limited field of- view, etc…