papers

Publications (53)

cs.CV2022

Self-supervised 360 Room Layout Estimation

Hao-Wen Ting, Cheng Sun, Hwann-Tzong Chen

We present the first self-supervised method to train panoramic room layout estimation models without any labeled data. Unlike per-pixel dense depth that provides abundant correspon…

cs.CV2024

Neural-PBIR Reconstruction of Shape, Material, and Illumination

Cheng Sun, Guangyan Cai, Zhengqin Li +6

Reconstructing the shape and spatially varying surface appearances of a physical-world object as well as its surrounding illumination based on 2D images (e.g., photographs) of the…

cs.CV2025

FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors

Chin-Yang Lin, Chung-Ho Wu, Chang-Han Yeh +3

Neural Radiance Fields (NeRF) face significant challenges in extreme few-shot scenarios, primarily due to overfitting and long training times. Existing methods, such as FreeNeRF an…

physics.optics2012

Nonreciprocal resonant transmission/reflection based on a one-dimensional photonic crystal adjacent to the magneto-optical metal film

Cheng He, Chang-Sheng Yuan, Ming-Hui Lu +2

We report the design of nonreciprocal resonant transmission/reflection using a one-dimensional photonic crystal (1DPC) adjacent to the magneto-optical (MO) metal film. The nonrecip…

cs.CV2023

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

Tao Tu, Shun-Po Chuang, Yu-Lun Liu +5

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods…

cs.GR2022

Improved Direct Voxel Grid Optimization for Radiance Fields Reconstruction

Cheng Sun, Min Sun, Hwann-Tzong Chen

In this technical report, we improve the DVGO framework (called DVGOv2), which is based on Pytorch and uses the simplest dense grid representation. First, we re-implement part of t…

math.CO2022

A bijection between the sets of -Generalized Motzkin paths avoiding -patterns and -patterns

Yidong Sun, Cheng Sun, Xiuli Hao

A generalized Motzkin path, called G-Motzkin path for short, of length is a lattice path from to in the first quadrant of the XOY-plane that consists of up st…

math.CO2022

The -avoiding -Generalized Motzkin paths with vertical steps: bijections and statistic enumerations

Yidong Sun, Weichen Wang, Cheng Sun

A generalized Motzkin path, called G-Motzkin path for short, of length is a lattice path from to in the first quadrant of the XOY-plane that consists of up st…

cs.AI2024

SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation

Yi-Chia Chen, Wei-Hua Li, Cheng Sun +2

We introduce SAM4MLLM, an innovative approach which integrates the Segment Anything Model (SAM) with Multi-Modal Large Language Models (MLLMs) for pixel-aware tasks. Our method ena…

physics.optics2011

Hiding a Realistic Object Using a Broadband Terahertz Invisibility Cloak

Fan Zhou, Yongjun Bao, Wei Cao +4

The invisibility cloak has been a long-standing dream for many researchers over the decades. The introduction of transformational optics has revitalized this field by providing a g…

cs.CV2025

Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering

Cheng Sun, Jaesung Choe, Charles Loop +2

We propose an efficient radiance field rendering algorithm that incorporates a rasterization process on adaptive sparse voxels without neural networks or 3D Gaussians. There are tw…

cs.CV2022

Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction

Cheng Sun, Min Sun, Hwann-Tzong Chen

We present a super-fast convergence approach to reconstructing the per-scene radiance field from a set of images that capture the scene with known poses. This task, which is often…

cs.CV2021

Indoor Panorama Planar 3D Reconstruction via Divide and Conquer

Cheng Sun, Chi-Wei Hsiao, Ning-Hsu Wang +2

Indoor panorama typically consists of human-made structures parallel or perpendicular to gravity. We leverage this phenomenon to approximate the scene in a 360-degree image with (H…

cs.CV2022

SVMAC: Unsupervised 3D Human Pose Estimation from a Single Image with Single-view-multi-angle Consistency

Yicheng Deng, Cheng Sun, Jiahui Zhu +1

Recovering 3D human pose from 2D joints is still a challenging problem, especially without any 3D annotation, video information, or multi-view information. In this paper, we presen…

cs.GR2026

Real-time 3D Visualization of Radiance Fields on Light Field Displays

Jonghyun Kim, Cheng Sun, Michael Stengel +6

Radiance fields, including their recent efficient forms such as 3D Gaussian Splatting and Sparse Voxels, have revolutionized photorealistic 3D scene visualization by enabling high-…

cs.CV2025

ASSR-NeRF: Arbitrary-Scale Super-Resolution on Voxel Grid for High-Quality Radiance Fields Reconstruction

Ding-Jiun Huang, Zi-Ting Chou, Yu-Chiang Frank Wang +1

NeRF-based methods reconstruct 3D scenes by building a radiance field with implicit or explicit representations. While NeRF-based methods can perform novel view synthesis (NVS) at…

cs.CV2022

Multiview Regenerative Morphing with Dual Flows

Chih-Jung Tsai, Cheng Sun, Hwann-Tzong Chen

This paper aims to address a new task of image morphing under a multiview setting, which takes two sets of multiview images as the input and generates intermediate renderings that…

cs.CV2019

360-Indoor: Towards Learning Real-World Objects in 360° Indoor Equirectangular Images

Shih-Han Chou, Cheng Sun, Wen-Yen Chang +3

While there are several widely used object detection datasets, current computer vision algorithms are still limited in conventional images. Such images narrow our vision in a restr…

cs.CV2023

Hashing Neural Video Decomposition with Multiplicative Residuals in Space-Time

Cheng-Hung Chan, Cheng-Yang Yuan, Cheng Sun +1

We present a video decomposition method that facilitates layer-based editing of videos with spatiotemporally varying lighting and motion effects. Our neural model decomposes an inp…

cs.CV2026

OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding

Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang +1

We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization…

cs.CV2025

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

Chin-Yang Lin, Cheng Sun, Fu-En Yang +3

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansi…

cs.CV2024

3D Human Pose Estimation Based on 2D-3D Consistency with Synchronized Adversarial Training

Yicheng Deng, Cheng Sun, Yongqi Sun +1

3D human pose estimation from a single image is still a challenging problem despite the large amount of work that has been performed in this field. Generally, most methods directly…

cs.CL2025

Streamlining Biomedical Research with Specialized LLMs

Linqing Chen, Weilei Wang, Yubin Xia +30

In this paper, we propose a novel system that integrates state-of-the-art, domain-specific large language models with advanced information retrieval techniques to deliver comprehen…

cs.CV2026

DVSM: Decoder-only View Synthesis Model Done Right

Cheng Sun, Jaesung Choe, Min-Hung Chen +2

Recent Large View Synthesis Models (LVSMs) advocate an encoder-decoder architecture that separates reconstruction and rendering into distinct networks. We re-examine this design. T…

cs.CV2026

3AM: 3egment Anything with Geometric Consistency in Videos

Yang-Che Sun, Cheng Sun, Chin-Yang Lin +4

Video object segmentation methods like SAM2 achieve strong performance through memory-based architectures but struggle under large viewpoint changes due to reliance on appearance f…

cs.CV2023

Seg2Reg: Differentiable 2D Segmentation to 1D Regression Rendering for 360 Room Layout Reconstruction

Cheng Sun, Wei-En Tai, Yu-Lin Shih +5

State-of-the-art single-view 360-degree room layout reconstruction methods formulate the problem as a high-level 1D (per-column) regression task. On the other hand, traditional low…

cs.CV2022

Data Efficient 3D Learner via Knowledge Transferred from 2D Model

Ping-Chung Yu, Cheng Sun, Min Sun

Collecting and labeling the registered 3D point cloud is costly. As a result, 3D resources for training are typically limited in quantity compared to the 2D images counterpart. In…

physics.bio-ph2011

Direct measurement of the correlated dynamics of the protein-backbone and proximal waters of hydration in mechanically strained elastin

Cheng Sun, Odingo Mitchell, Jiaxin Huang +1

We report on the direct measurement of the correlation times of the protein backbone carbons and proximal waters of hydration in mechanically strained elastin by nuclear magnetic r…

physics.optics2008

Cloaking of Matter Waves

Shuang Zhang, Dentcho A. Genov, Cheng Sun +1

Invariant transformation for quantum mechanical systems is proposed. A cloaking of matter wave can be realized at given energy by designing the potential and effective mass of the…

cs.CV2026

Customized Visual Storytelling with Unified Multimodal LLMs

Wei-Hua Li, Cheng Sun, Chu-Song Chen

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story…

physics.optics2010

Three-Dimensional Cloaking Device Operates at Terahertz Frequencies

Fan Zhou, Yongjun Bao, Wei Cao +3

The invisibility cloak has been a long-standing dream for many researchers over the decades. By transforming space and light propagation, a three-dimensional (3D) object can be per…

cs.CV2026

MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance

Yoonwoo Jeong, Cheng Sun, Yu-Chiang Frank Wang +2

Promptable segmentation has emerged as a powerful paradigm in computer vision, enabling users to guide models in parsing complex scenes with prompts such as clicks, boxes, or textu…

cs.CV2023

PanoMixSwap Panorama Mixing via Structural Swapping for Indoor Scene Understanding

Yu-Cheng Hsieh, Cheng Sun, Suraj Dengale +1

The volume and diversity of training data are critical for modern deep learningbased methods. Compared to the massive amount of labeled perspective images, 360 panoramic images fal…

cs.CV2018

A Spatial and Temporal Features Mixture Model with Body Parts for Video-based Person Re-Identification

Jie Liu, Cheng Sun, Xiang Xu +2

The video-based person re-identification is to recognize a person under different cameras, which is a crucial task applied in visual surveillance system. Most previous methods main…

cond-mat.mtrl-sci2022

Machine learning in nuclear materials research

Dane Morgan, Ghanshyam Pilania, Adrien Couet +3

Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes and transmutation, high temperature and temperature grad…

cs.CV2025

Segment Anything, Even Occluded

Wei-En Tai, Yu-Lin Shih, Cheng Sun +2

Amodal instance segmentation, which aims to detect and segment both visible and invisible parts of objects in images, plays a crucial role in various applications including autonom…

physics.optics2011

Construction of Chiral Metamaterial with a Helix Array

Xiang Xiong, Xiao-Chun Chen, Zhao-Wu Wang +5

Here we report the designing of chiral metamaterial with metallic helix array. The effective electric and magnetic dipoles, which originate from the induced surface electric curren…

cs.LG2025

A Variance-Reduced Cubic-Regularized Newton for Policy Optimization

Cheng Sun, Zhen Zhang, Shaofu Yang

In this paper, we study a second-order approach to policy optimization in reinforcement learning. Existing second-order methods often suffer from suboptimal sample complexity or re…

cs.CL2024

PatentGPT: A Large Language Model for Intellectual Property

Zilong Bai, Ruiji Zhang, Linqing Chen +24

In recent years, large language models(LLMs) have attracted significant attention due to their exceptional performance across a multitude of natural language process tasks, and hav…

cs.CV2026

Advancing Structured Priors for Sparse-Voxel Surface Reconstruction

Ting-Hsun Chi, Chu-Rong Chen, Chi-Tun Hsu +4

Reconstructing accurate surfaces with radiance fields has progressed rapidly, yet two promising explicit representations, 3D Gaussian Splatting and sparse-voxel rasterization, exhi…

cs.CV2021

HoHoNet: 360 Indoor Holistic Understanding with Latent Horizontal Features

Cheng Sun, Min Sun, Hwann-Tzong Chen

We present HoHoNet, a versatile and efficient framework for holistic understanding of an indoor 360-degree panorama using a Latent Horizontal Feature (LHFeat). The compact LHFeat f…

cs.CV2021

Moving in a 360 World: Synthesizing Panoramic Parallaxes from a Single Panorama

Ching-Yu Hsu, Cheng Sun, Hwann-Tzong Chen

We present Omnidirectional Neural Radiance Fields (OmniNeRF), the first method to the application of parallax-enabled novel panoramic view synthesis. Recent works for novel view sy…

physics.optics2015

Ultrafast All-optical Modulation Exploiting the Vibrational Dynamic of Metallic Meta-atoms

Biqin Dong, Xiangfan Chen, Fan Zhou +3

Optical control over elementary molecular vibration establishes fundamental capabilities for exploiting the broad range of optical linear and nonlinear phenomena. However, experime…

cs.CV2019

Flat2Layout: Flat Representation for Estimating Layout of General Room Types

Chi-Wei Hsiao, Cheng Sun, Min Sun +1

This paper proposes a new approach, Flat2Layout, for estimating general indoor room layout from a single-view RGB image whereas existing methods can only produce layout topologies…

cs.LG2025

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL

Yu-Heng Hung, Kai-Jie Lin, Yu-Heng Lin +3

Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in…

physics.optics2009

Construction of Chiral Metamaterial with U-Shaped Resonator Assembly

Xiang Xiong, Wei-Hua Sun, Yong-Jun Bao +7

Chiral structure can be applied to construct metamaterial with negative refractive index (NRI). In an assembly of double-layered metallic U-shaped resonators with two resonant freq…

cs.CV2025

R2SM: Referring and Reasoning for Selective Masks

Yu-Lin Shih, Wei-En Tai, Cheng Sun +2

We introduce a new task, Referring and Reasoning for Selective Masks (R2SM), which extends text-guided segmentation by incorporating mask-type selection driven by user intent. This…

cs.CV2025

Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting

Yoonwoo Jeong, Cheng Sun, Frank Wang +2

Recent advancements in computer vision have successfully extended Open-vocabulary segmentation (OVS) to the 3D domain by leveraging 3D Gaussian Splatting (3D-GS). Despite this prog…

cs.CL2024

PharmaGPT: Domain-Specific Large Language Models for Bio-Pharmaceutical and Chemistry

Linqing Chen, Weilei Wang, Zilong Bai +33

Large language models (LLMs) have revolutionized Natural Language Processing (NLP) by minimizing the need for complex feature engineering. However, the application of LLMs in speci…

cs.CV2019

HorizonNet: Learning Room Layout with 1D Representation and Pano Stretch Data Augmentation

Cheng Sun, Chi-Wei Hsiao, Min Sun +1

We present a new approach to the problem of estimating the 3D room layout from a single panoramic image. We represent room layout as three 1D vectors that encode, at each image col…

cond-mat.mtrl-sci2026

Current-tunable room temperature ferromagnetism and current-driven phase transitions

Jianping Guo, Peng Rao, Xinhao Huang +9

It is generally assumed that the application of a charge-current in ferromagnetic metals suppresses their ferromagnetic order through trivial Joule heating. Here, we demonstrate th…

cs.CV2021

Specialize and Fuse: Pyramidal Output Representation for Semantic Segmentation

Chi-Wei Hsiao, Cheng Sun, Hwann-Tzong Chen +1

We present a novel pyramidal output representation to ensure parsimony with our "specialize and fuse" process for semantic segmentation. A pyramidal "output" representation consist…

cs.CL2025

LongCat-Flash Technical Report

Meituan LongCat Team, Bayan, Bei Li +179

We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming f…