Publications (39)
Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
Lin Li, Chuhan Zhang, Dong Zhang +3
Open-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pr…
Finite-temperature simulations of strongly correlated systems
Chong Sun
This thesis describes several topics related to finite temperature studies of strongly correlated systems: finite temperature density matrix embedding theory (FT-DMET), finite temp…
Towards Efficient Multi-Scale Deformable Attention on NPU
Chenghuan Huang, Zhigeng Xu, Chong Sun +2
Multi-scale deformable attention (MSDA) is a flexible and powerful feature extraction mechanism for visual tasks, but its random-access grid sampling strategy poses significant opt…
SELFIES and the future of molecular string representations
Mario Krenn, Qianxiang Ai, Senja Barthel +28
Artificial intelligence (AI) and machine learning (ML) are expanding in popularity for broad applications to challenging tasks in chemistry and materials science. Examples include…
Towards Software Development For Social Robotics Systems
Chong Sun, Jiongyan Zhang, Cong Liu +4
In this paper we introduce the core results of the project on software development for social robotics systems. The usability of maintenance and control features is crucial for man…
Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation
Lin Li, Chuhan Zhang, Dong Zhang +3
Today's open vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge…
Video-GPT via Next Clip Diffusion
Shaobin Zhuang, Zhipeng Huang, Ying Zhang +6
GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alte…
VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features
Sifei Li, Binxin Yang, Chunji Yin +4
Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achievi…
Electron localization in disordered quantum systems at finite temperatures
Chong Sun
We study electron localization in disordered quantum systems, focusing on both individual eigenstates and thermal states. We employ complex polarization as a numerical indicator to…
ROI Pooled Correlation Filters for Visual Tracking
Yuxuan Sun, Chong Sun, Dong Wang +2
The ROI (region-of-interest) based pooling method performs pooling operations on the cropped ROI regions for various samples and has shown great success in the object detection met…
What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
Yingquan Wang, Pingping Zhang, Chong Sun +2
Object Re-IDentification (ReID) aims to recognize individuals across non-overlapping camera views. While recent advances have achieved remarkable progress, most existing models are…
Correlation Tracking via Joint Discrimination and Reliability Learning
Chong Sun, Dong Wang, Huchuan Lu +1
For visual tracking, an ideal filter learned by the correlation filter (CF) method should take both discrimination and reliability information. However, existing attempts usually f…
VACoT: Rethinking Visual Data Augmentation with VLMs
Zhengzhuo Xu, Chong Sun, SiNan Du +3
While visual data augmentation remains a cornerstone for training robust vision models, it has received limited attention in visual language models (VLMs), which predominantly rely…
Variational quantum iterative power algorithms for global optimization
Thi Ha Kyaw, Micheline B. Soley, Brandon Allen +4
We introduce a family of variational quantum algorithms called quantum iterative power algorithms (QIPA) that outperform existing hybrid near-term quantum algorithms of the same ki…
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
Xingsong Ye, Yongkun Du, Jiaxin Zhang +5
WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general…
Learning Spatial-Aware Regressions for Visual Tracking
Chong Sun, Dong Wang, Huchuan Lu +1
In this paper, we analyze the spatial information of deep features, and propose two complementary regressions for robust visual tracking. First, we propose a kernelized ridge regre…
Spatial-Scale Aligned Network for Fine-Grained Recognition
Lizhao Gao, Haihua Xu, Chong Sun +2
Existing approaches for fine-grained visual recognition focus on learning marginal region-based representations while neglecting the spatial and scale misalignments, leading to inf…
Ab initio quantum embedding at finite temperature with density matrix embedding theory
Laurence Giordano, Y. Stanley Tan, Zhi-Hao Cui +1
We present a finite-temperature extension of density matrix embedding theory (FT-DMET) for realistic crystalline systems. We describe a practical framework for constructing extende…
V-Thinker: Interactive Thinking with Images
Runqi Qiao, Qiuna Tan, Minghan Yang +11
Empowering Large Multimodal Models (LMMs) to deeply integrate image interaction with long-horizon reasoning capabilities remains a long-standing challenge in this field. Recent adv…
Block2: a comprehensive open source framework to develop and apply state-of-the-art DMRG algorithms in electronic structure and beyond
Huanchen Zhai, Henrik R. Larsson, Seunghoon Lee +10
Block2 is an open source framework to implement and perform density matrix renormalization group and matrix product state algorithms. Out-of-the-box it supports the eigenstate, tim…
Text-guided Visual Prompt DINO for Generic Segmentation
Yuchen Guan, Chong Sun, Canmiao Fu +3
Recent advancements in multimodal vision models have highlighted limitations in late-stage feature fusion and suboptimal query selection for hybrid prompts open-world segmentation,…
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
Runqi Qiao, Qiuna Tan, Peiqing Yang +11
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various tasks, but still struggle with complex mathematical reasoning. Existing research p…
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Yongyi Su, Haojie Zhang, Shijie Li +11
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…
Monte-Carlo Simulations of Spin-Crossover Phenomena Based on a Vibronic Ising-like Model with Realistic Parameters
Hong-zhou Ye, Chong Sun, Hong Jiang
Materials with spin-crossover (SCO) properties hold great potentials in information storage and therefore have received a lot of concerns in the recent decades. The hysteresis phen…
Ground-state phase diagram of the three-band Hubbard model from density matrix embedding theory
Zhi-Hao Cui, Chong Sun, Ushnish Ray +3
We determine the ground-state phase diagram of the three-band Hubbard model across a range of model parameters using density matrix embedding theory. We study the atomic-scale natu…
The Python Simulations of Chemistry Framework: 10 years of an open-source quantum chemistry project
Qiming Sun, Matthew R Hermes, Xiaojie Wu +100
Over the past decade, the Python-based Simulations of Chemistry Framework (PySCF) has developed into a widely used open-source platform for electronic structure theory and quantum…
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Runqi Qiao, Qiuna Tan, Guanting Dong +15
Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks,…
Determining eigenstates and thermal states on a quantum computer using quantum imaginary time evolution
Mario Motta, Chong Sun, Adrian Teck Keng Tan +5
The accurate computation of Hamiltonian ground, excited, and thermal states on quantum computers stands to impact many problems in the physical and computer sciences, from quantum…
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
Zhipeng Huang, Shaobin Zhuang, Canmiao Fu +7
Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack th…
When to Reject a Ground State Preparation Algorithm
Katerina Gratsea, Chong Sun, Peter D. Johnson
In recent years substantial research effort has been devoted to quantum algorithms for ground state energy estimation (GSEE) in chemistry and materials. Given the many heuristic an…
Towards Domain Generalization in Object Detection
Xingxuan Zhang, Zekai Xu, Renzhe Xu +5
Despite the striking performance achieved by modern detectors when training and test data are sampled from the same or similar distribution, the generalization ability of detectors…
Recent developments in the PySCF program package
Qiming Sun, Xing Zhang, Samragni Banerjee +46
PYSCF is a Python-based general-purpose electronic structure platform that both supports first-principles simulations of molecules and solids, as well as accelerates the developmen…
Scalable Autoregressive 3D Molecule Generation
Austin H. Cheng, Chong Sun, Alán Aspuru-Guzik
Generative models of 3D molecular structure play a rapidly growing role in the design and simulation of molecules. Diffusion models currently dominate the space of 3D molecule gene…
Selected non-orthogonal configuration interaction with compressed single and double excitations
Chong Sun, Fei Gao, Gustavo E. Scuseria
Addressing both dynamic and static correlation accurately is a primary goal in electronic structure theory. Non-orthogonal configuration interaction (NOCI) is a versatile tool for…
QDK/Chemistry: A Modular Toolkit for Quantum Chemistry Applications
Nathan A. Baker, Brian Bilodeau, Chi Chen +23
We present QDK/Chemistry, a software toolkit for quantum chemistry workflows targeting quantum computers. The toolkit addresses a key challenge in the field: while quantum algorith…
Get In Video: Add Anything You Want to the Video
Shaobin Zhuang, Zhipeng Huang, Binxin Yang +7
Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique v…
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
Hao Yu, Jiabo Zhan, Kang Liu +8
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regi…
Finite temperature density matrix embedding theory
Chong Sun, Ushnish Ray, Zhi-Hao Cui +3
We describe a formulation of the density matrix embedding theory at finite temperature. We present a generalization of the ground-state bath orbital construction that embeds a mean…
Waveflow: boundary-conditioned normalizing flows applied to fermionic wavefunctions
Luca Thiede, Chong Sun, Alán Aspuru-Guzik
An efficient and expressive wavefunction ansatz is key to scalable solutions for complex many-body electronic structures. While Slater determinants are predominantly used for const…