Publications (53)
Self-supervised 360 Room Layout Estimation
Hao-Wen Ting, Cheng Sun, Hwann-Tzong Chen
We present the first self-supervised method to train panoramic room layout estimation models without any labeled data. Unlike per-pixel dense depth that provides abundant correspon…
Neural-PBIR Reconstruction of Shape, Material, and Illumination
Cheng Sun, Guangyan Cai, Zhengqin Li +6
Reconstructing the shape and spatially varying surface appearances of a physical-world object as well as its surrounding illumination based on 2D images (e.g., photographs) of the…
FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors
Chin-Yang Lin, Chung-Ho Wu, Chang-Han Yeh +3
Neural Radiance Fields (NeRF) face significant challenges in extreme few-shot scenarios, primarily due to overfitting and long training times. Existing methods, such as FreeNeRF an…
Nonreciprocal resonant transmission/reflection based on a one-dimensional photonic crystal adjacent to the magneto-optical metal film
Cheng He, Chang-Sheng Yuan, Ming-Hui Lu +2
We report the design of nonreciprocal resonant transmission/reflection using a one-dimensional photonic crystal (1DPC) adjacent to the magneto-optical (MO) metal film. The nonrecip…
ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection
Tao Tu, Shun-Po Chuang, Yu-Lun Liu +5
We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods…
Improved Direct Voxel Grid Optimization for Radiance Fields Reconstruction
Cheng Sun, Min Sun, Hwann-Tzong Chen
In this technical report, we improve the DVGO framework (called DVGOv2), which is based on Pytorch and uses the simplest dense grid representation. First, we re-implement part of t…
A bijection between the sets of -Generalized Motzkin paths avoiding -patterns and -patterns
Yidong Sun, Cheng Sun, Xiuli Hao
A generalized Motzkin path, called G-Motzkin path for short, of length is a lattice path from to in the first quadrant of the XOY-plane that consists of up st…
The -avoiding -Generalized Motzkin paths with vertical steps: bijections and statistic enumerations
Yidong Sun, Weichen Wang, Cheng Sun
A generalized Motzkin path, called G-Motzkin path for short, of length is a lattice path from to in the first quadrant of the XOY-plane that consists of up st…
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
Yi-Chia Chen, Wei-Hua Li, Cheng Sun +2
We introduce SAM4MLLM, an innovative approach which integrates the Segment Anything Model (SAM) with Multi-Modal Large Language Models (MLLMs) for pixel-aware tasks. Our method ena…
Hiding a Realistic Object Using a Broadband Terahertz Invisibility Cloak
Fan Zhou, Yongjun Bao, Wei Cao +4
The invisibility cloak has been a long-standing dream for many researchers over the decades. The introduction of transformational optics has revitalized this field by providing a g…
Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering
Cheng Sun, Jaesung Choe, Charles Loop +2
We propose an efficient radiance field rendering algorithm that incorporates a rasterization process on adaptive sparse voxels without neural networks or 3D Gaussians. There are tw…
Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction
Cheng Sun, Min Sun, Hwann-Tzong Chen
We present a super-fast convergence approach to reconstructing the per-scene radiance field from a set of images that capture the scene with known poses. This task, which is often…
Indoor Panorama Planar 3D Reconstruction via Divide and Conquer
Cheng Sun, Chi-Wei Hsiao, Ning-Hsu Wang +2
Indoor panorama typically consists of human-made structures parallel or perpendicular to gravity. We leverage this phenomenon to approximate the scene in a 360-degree image with (H…
SVMAC: Unsupervised 3D Human Pose Estimation from a Single Image with Single-view-multi-angle Consistency
Yicheng Deng, Cheng Sun, Jiahui Zhu +1
Recovering 3D human pose from 2D joints is still a challenging problem, especially without any 3D annotation, video information, or multi-view information. In this paper, we presen…
Real-time 3D Visualization of Radiance Fields on Light Field Displays
Jonghyun Kim, Cheng Sun, Michael Stengel +6
Radiance fields, including their recent efficient forms such as 3D Gaussian Splatting and Sparse Voxels, have revolutionized photorealistic 3D scene visualization by enabling high-…
ASSR-NeRF: Arbitrary-Scale Super-Resolution on Voxel Grid for High-Quality Radiance Fields Reconstruction
Ding-Jiun Huang, Zi-Ting Chou, Yu-Chiang Frank Wang +1
NeRF-based methods reconstruct 3D scenes by building a radiance field with implicit or explicit representations. While NeRF-based methods can perform novel view synthesis (NVS) at…
Multiview Regenerative Morphing with Dual Flows
Chih-Jung Tsai, Cheng Sun, Hwann-Tzong Chen
This paper aims to address a new task of image morphing under a multiview setting, which takes two sets of multiview images as the input and generates intermediate renderings that…
360-Indoor: Towards Learning Real-World Objects in 360° Indoor Equirectangular Images
Shih-Han Chou, Cheng Sun, Wen-Yen Chang +3
While there are several widely used object detection datasets, current computer vision algorithms are still limited in conventional images. Such images narrow our vision in a restr…
Hashing Neural Video Decomposition with Multiplicative Residuals in Space-Time
Cheng-Hung Chan, Cheng-Yang Yuan, Cheng Sun +1
We present a video decomposition method that facilitates layer-based editing of videos with spatiotemporally varying lighting and motion effects. Our neural model decomposes an inp…
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang +1
We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization…
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
Chin-Yang Lin, Cheng Sun, Fu-En Yang +3
LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansi…
3D Human Pose Estimation Based on 2D-3D Consistency with Synchronized Adversarial Training
Yicheng Deng, Cheng Sun, Yongqi Sun +1
3D human pose estimation from a single image is still a challenging problem despite the large amount of work that has been performed in this field. Generally, most methods directly…
Streamlining Biomedical Research with Specialized LLMs
Linqing Chen, Weilei Wang, Yubin Xia +30
In this paper, we propose a novel system that integrates state-of-the-art, domain-specific large language models with advanced information retrieval techniques to deliver comprehen…
DVSM: Decoder-only View Synthesis Model Done Right
Cheng Sun, Jaesung Choe, Min-Hung Chen +2
Recent Large View Synthesis Models (LVSMs) advocate an encoder-decoder architecture that separates reconstruction and rendering into distinct networks. We re-examine this design. T…
3AM: 3egment Anything with Geometric Consistency in Videos
Yang-Che Sun, Cheng Sun, Chin-Yang Lin +4
Video object segmentation methods like SAM2 achieve strong performance through memory-based architectures but struggle under large viewpoint changes due to reliance on appearance f…
Seg2Reg: Differentiable 2D Segmentation to 1D Regression Rendering for 360 Room Layout Reconstruction
Cheng Sun, Wei-En Tai, Yu-Lin Shih +5
State-of-the-art single-view 360-degree room layout reconstruction methods formulate the problem as a high-level 1D (per-column) regression task. On the other hand, traditional low…
Data Efficient 3D Learner via Knowledge Transferred from 2D Model
Ping-Chung Yu, Cheng Sun, Min Sun
Collecting and labeling the registered 3D point cloud is costly. As a result, 3D resources for training are typically limited in quantity compared to the 2D images counterpart. In…
Direct measurement of the correlated dynamics of the protein-backbone and proximal waters of hydration in mechanically strained elastin
Cheng Sun, Odingo Mitchell, Jiaxin Huang +1
We report on the direct measurement of the correlation times of the protein backbone carbons and proximal waters of hydration in mechanically strained elastin by nuclear magnetic r…
Cloaking of Matter Waves
Shuang Zhang, Dentcho A. Genov, Cheng Sun +1
Invariant transformation for quantum mechanical systems is proposed. A cloaking of matter wave can be realized at given energy by designing the potential and effective mass of the…
Customized Visual Storytelling with Unified Multimodal LLMs
Wei-Hua Li, Cheng Sun, Chu-Song Chen
Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story…
Three-Dimensional Cloaking Device Operates at Terahertz Frequencies
Fan Zhou, Yongjun Bao, Wei Cao +3
The invisibility cloak has been a long-standing dream for many researchers over the decades. By transforming space and light propagation, a three-dimensional (3D) object can be per…
MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance
Yoonwoo Jeong, Cheng Sun, Yu-Chiang Frank Wang +2
Promptable segmentation has emerged as a powerful paradigm in computer vision, enabling users to guide models in parsing complex scenes with prompts such as clicks, boxes, or textu…
PanoMixSwap Panorama Mixing via Structural Swapping for Indoor Scene Understanding
Yu-Cheng Hsieh, Cheng Sun, Suraj Dengale +1
The volume and diversity of training data are critical for modern deep learningbased methods. Compared to the massive amount of labeled perspective images, 360 panoramic images fal…
A Spatial and Temporal Features Mixture Model with Body Parts for Video-based Person Re-Identification
Jie Liu, Cheng Sun, Xiang Xu +2
The video-based person re-identification is to recognize a person under different cameras, which is a crucial task applied in visual surveillance system. Most previous methods main…
Machine learning in nuclear materials research
Dane Morgan, Ghanshyam Pilania, Adrien Couet +3
Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes and transmutation, high temperature and temperature grad…
Segment Anything, Even Occluded
Wei-En Tai, Yu-Lin Shih, Cheng Sun +2
Amodal instance segmentation, which aims to detect and segment both visible and invisible parts of objects in images, plays a crucial role in various applications including autonom…
Construction of Chiral Metamaterial with a Helix Array
Xiang Xiong, Xiao-Chun Chen, Zhao-Wu Wang +5
Here we report the designing of chiral metamaterial with metallic helix array. The effective electric and magnetic dipoles, which originate from the induced surface electric curren…
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
Cheng Sun, Zhen Zhang, Shaofu Yang
In this paper, we study a second-order approach to policy optimization in reinforcement learning. Existing second-order methods often suffer from suboptimal sample complexity or re…
PatentGPT: A Large Language Model for Intellectual Property
Zilong Bai, Ruiji Zhang, Linqing Chen +24
In recent years, large language models(LLMs) have attracted significant attention due to their exceptional performance across a multitude of natural language process tasks, and hav…
Advancing Structured Priors for Sparse-Voxel Surface Reconstruction
Ting-Hsun Chi, Chu-Rong Chen, Chi-Tun Hsu +4
Reconstructing accurate surfaces with radiance fields has progressed rapidly, yet two promising explicit representations, 3D Gaussian Splatting and sparse-voxel rasterization, exhi…
HoHoNet: 360 Indoor Holistic Understanding with Latent Horizontal Features
Cheng Sun, Min Sun, Hwann-Tzong Chen
We present HoHoNet, a versatile and efficient framework for holistic understanding of an indoor 360-degree panorama using a Latent Horizontal Feature (LHFeat). The compact LHFeat f…
Moving in a 360 World: Synthesizing Panoramic Parallaxes from a Single Panorama
Ching-Yu Hsu, Cheng Sun, Hwann-Tzong Chen
We present Omnidirectional Neural Radiance Fields (OmniNeRF), the first method to the application of parallax-enabled novel panoramic view synthesis. Recent works for novel view sy…
Ultrafast All-optical Modulation Exploiting the Vibrational Dynamic of Metallic Meta-atoms
Biqin Dong, Xiangfan Chen, Fan Zhou +3
Optical control over elementary molecular vibration establishes fundamental capabilities for exploiting the broad range of optical linear and nonlinear phenomena. However, experime…
Flat2Layout: Flat Representation for Estimating Layout of General Room Types
Chi-Wei Hsiao, Cheng Sun, Min Sun +1
This paper proposes a new approach, Flat2Layout, for estimating general indoor room layout from a single-view RGB image whereas existing methods can only produce layout topologies…
BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
Yu-Heng Hung, Kai-Jie Lin, Yu-Heng Lin +3
Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in…
Construction of Chiral Metamaterial with U-Shaped Resonator Assembly
Xiang Xiong, Wei-Hua Sun, Yong-Jun Bao +7
Chiral structure can be applied to construct metamaterial with negative refractive index (NRI). In an assembly of double-layered metallic U-shaped resonators with two resonant freq…
R2SM: Referring and Reasoning for Selective Masks
Yu-Lin Shih, Wei-En Tai, Cheng Sun +2
We introduce a new task, Referring and Reasoning for Selective Masks (R2SM), which extends text-guided segmentation by incorporating mask-type selection driven by user intent. This…
Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
Yoonwoo Jeong, Cheng Sun, Frank Wang +2
Recent advancements in computer vision have successfully extended Open-vocabulary segmentation (OVS) to the 3D domain by leveraging 3D Gaussian Splatting (3D-GS). Despite this prog…
PharmaGPT: Domain-Specific Large Language Models for Bio-Pharmaceutical and Chemistry
Linqing Chen, Weilei Wang, Zilong Bai +33
Large language models (LLMs) have revolutionized Natural Language Processing (NLP) by minimizing the need for complex feature engineering. However, the application of LLMs in speci…
HorizonNet: Learning Room Layout with 1D Representation and Pano Stretch Data Augmentation
Cheng Sun, Chi-Wei Hsiao, Min Sun +1
We present a new approach to the problem of estimating the 3D room layout from a single panoramic image. We represent room layout as three 1D vectors that encode, at each image col…
Current-tunable room temperature ferromagnetism and current-driven phase transitions
Jianping Guo, Peng Rao, Xinhao Huang +9
It is generally assumed that the application of a charge-current in ferromagnetic metals suppresses their ferromagnetic order through trivial Joule heating. Here, we demonstrate th…
Specialize and Fuse: Pyramidal Output Representation for Semantic Segmentation
Chi-Wei Hsiao, Cheng Sun, Hwann-Tzong Chen +1
We present a novel pyramidal output representation to ensure parsimony with our "specialize and fuse" process for semantic segmentation. A pyramidal "output" representation consist…
LongCat-Flash Technical Report
Meituan LongCat Team, Bayan, Bei Li +179
We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming f…