papers

Publications (39)

cs.CV2025

Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation

Lin Li, Chuhan Zhang, Dong Zhang +3

Open-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pr…

cond-mat.str-el2023

Finite-temperature simulations of strongly correlated systems

Chong Sun

This thesis describes several topics related to finite temperature studies of strongly correlated systems: finite temperature density matrix embedding theory (FT-DMET), finite temp…

cs.PF2025

Towards Efficient Multi-Scale Deformable Attention on NPU

Chenghuan Huang, Zhigeng Xu, Chong Sun +2

Multi-scale deformable attention (MSDA) is a flexible and powerful feature extraction mechanism for visual tasks, but its random-access grid sampling strategy poses significant opt…

physics.chem-ph2022

SELFIES and the future of molecular string representations

Mario Krenn, Qianxiang Ai, Senja Barthel +28

Artificial intelligence (AI) and machine learning (ML) are expanding in popularity for broad applications to challenging tasks in chemistry and materials science. Examples include…

cs.RO2017

Towards Software Development For Social Robotics Systems

Chong Sun, Jiongyan Zhang, Cong Liu +4

In this paper we introduce the core results of the project on software development for social robotics systems. The usability of maintenance and control features is crucial for man…

cs.CV2025

Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation

Lin Li, Chuhan Zhang, Dong Zhang +3

Today's open vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge…

cs.CV2025

Video-GPT via Next Clip Diffusion

Shaobin Zhuang, Zhipeng Huang, Ying Zhang +6

GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alte…

cs.SD2024

VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

Sifei Li, Binxin Yang, Chunji Yin +4

Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achievi…

cond-mat.dis-nn2024

Electron localization in disordered quantum systems at finite temperatures

Chong Sun

We study electron localization in disordered quantum systems, focusing on both individual eigenstates and thermal states. We employ complex polarization as a numerical indicator to…

cs.CV2019

ROI Pooled Correlation Filters for Visual Tracking

Yuxuan Sun, Chong Sun, Dong Wang +2

The ROI (region-of-interest) based pooling method performs pooling operations on the cropped ROI regions for various samples and has shown great success in the object detection met…

cs.CV2025

What Makes You Unique? Attribute Prompt Composition for Object Re-Identification

Yingquan Wang, Pingping Zhang, Chong Sun +2

Object Re-IDentification (ReID) aims to recognize individuals across non-overlapping camera views. While recent advances have achieved remarkable progress, most existing models are…

cs.CV2018

Correlation Tracking via Joint Discrimination and Reliability Learning

Chong Sun, Dong Wang, Huchuan Lu +1

For visual tracking, an ideal filter learned by the correlation filter (CF) method should take both discrimination and reliability information. However, existing attempts usually f…

cs.CV2025

VACoT: Rethinking Visual Data Augmentation with VLMs

Zhengzhuo Xu, Chong Sun, SiNan Du +3

While visual data augmentation remains a cornerstone for training robust vision models, it has received limited attention in visual language models (VLMs), which predominantly rely…

quant-ph2022

Variational quantum iterative power algorithms for global optimization

Thi Ha Kyaw, Micheline B. Soley, Brandon Allen +4

We introduce a family of variational quantum algorithms called quantum iterative power algorithms (QIPA) that outperform existing hybrid near-term quantum algorithms of the same ki…

cs.CV2026

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods

Xingsong Ye, Yongkun Du, Jiaxin Zhang +5

WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general…

cs.CV2018

Learning Spatial-Aware Regressions for Visual Tracking

Chong Sun, Dong Wang, Huchuan Lu +1

In this paper, we analyze the spatial information of deep features, and propose two complementary regressions for robust visual tracking. First, we propose a kernelized ridge regre…

cs.CV2020

Spatial-Scale Aligned Network for Fine-Grained Recognition

Lizhao Gao, Haihua Xu, Chong Sun +2

Existing approaches for fine-grained visual recognition focus on learning marginal region-based representations while neglecting the spatial and scale misalignments, leading to inf…

physics.comp-ph2026

Ab initio quantum embedding at finite temperature with density matrix embedding theory

Laurence Giordano, Y. Stanley Tan, Zhi-Hao Cui +1

We present a finite-temperature extension of density matrix embedding theory (FT-DMET) for realistic crystalline systems. We describe a practical framework for constructing extende…

cs.CV2025

V-Thinker: Interactive Thinking with Images

Runqi Qiao, Qiuna Tan, Minghan Yang +11

Empowering Large Multimodal Models (LMMs) to deeply integrate image interaction with long-horizon reasoning capabilities remains a long-standing challenge in this field. Recent adv…

physics.chem-ph2023

Block2: a comprehensive open source framework to develop and apply state-of-the-art DMRG algorithms in electronic structure and beyond

Huanchen Zhai, Henrik R. Larsson, Seunghoon Lee +10

Block2 is an open source framework to implement and perform density matrix renormalization group and matrix product state algorithms. Out-of-the-box it supports the eigenstate, tim…

cs.CV2025

Text-guided Visual Prompt DINO for Generic Segmentation

Yuchen Guan, Chong Sun, Canmiao Fu +3

Recent advancements in multimodal vision models have highlighted limitations in late-stage feature fusion and suboptimal query selection for hybrid prompts open-world segmentation,…

cs.AI2025

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Runqi Qiao, Qiuna Tan, Peiqing Yang +11

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various tasks, but still struggle with complex mathematical reasoning. Existing research p…

cs.CV2025

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Yongyi Su, Haojie Zhang, Shijie Li +11

Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…

physics.chem-ph2014

Monte-Carlo Simulations of Spin-Crossover Phenomena Based on a Vibronic Ising-like Model with Realistic Parameters

Hong-zhou Ye, Chong Sun, Hong Jiang

Materials with spin-crossover (SCO) properties hold great potentials in information storage and therefore have received a lot of concerns in the recent decades. The hysteresis phen…

cond-mat.str-el2020

Ground-state phase diagram of the three-band Hubbard model from density matrix embedding theory

Zhi-Hao Cui, Chong Sun, Ushnish Ray +3

We determine the ground-state phase diagram of the three-band Hubbard model across a range of model parameters using density matrix embedding theory. We study the atomic-scale natu…

physics.chem-ph2026

The Python Simulations of Chemistry Framework: 10 years of an open-source quantum chemistry project

Qiming Sun, Matthew R Hermes, Xiaojie Wu +100

Over the past decade, the Python-based Simulations of Chemistry Framework (PySCF) has developed into a widely used open-source platform for electronic structure theory and quantum…

cs.AI2024

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Runqi Qiao, Qiuna Tan, Guanting Dong +15

Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks,…

quant-ph2020

Determining eigenstates and thermal states on a quantum computer using quantum imaginary time evolution

Mario Motta, Chong Sun, Adrian Teck Keng Tan +5

The accurate computation of Hamiltonian ground, excited, and thermal states on quantum computers stands to impact many problems in the physical and computer sciences, from quantum…

cs.CV2025

WeGen: A Unified Model for Interactive Multimodal Generation as We Chat

Zhipeng Huang, Shaobin Zhuang, Canmiao Fu +7

Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack th…

quant-ph2022

When to Reject a Ground State Preparation Algorithm

Katerina Gratsea, Chong Sun, Peter D. Johnson

In recent years substantial research effort has been devoted to quantum algorithms for ground state energy estimation (GSEE) in chemistry and materials. Given the many heuristic an…

cs.CV2022

Towards Domain Generalization in Object Detection

Xingxuan Zhang, Zekai Xu, Renzhe Xu +5

Despite the striking performance achieved by modern detectors when training and test data are sampled from the same or similar distribution, the generalization ability of detectors…

physics.chem-ph2020

Recent developments in the PySCF program package

Qiming Sun, Xing Zhang, Samragni Banerjee +46

PYSCF is a Python-based general-purpose electronic structure platform that both supports first-principles simulations of molecules and solids, as well as accelerates the developmen…

cs.LG2025

Scalable Autoregressive 3D Molecule Generation

Austin H. Cheng, Chong Sun, Alán Aspuru-Guzik

Generative models of 3D molecular structure play a rapidly growing role in the design and simulation of molecules. Diffusion models currently dominate the space of 3D molecule gene…

physics.chem-ph2024

Selected non-orthogonal configuration interaction with compressed single and double excitations

Chong Sun, Fei Gao, Gustavo E. Scuseria

Addressing both dynamic and static correlation accurately is a primary goal in electronic structure theory. Non-orthogonal configuration interaction (NOCI) is a versatile tool for…

quant-ph2026

QDK/Chemistry: A Modular Toolkit for Quantum Chemistry Applications

Nathan A. Baker, Brian Bilodeau, Chi Chen +23

We present QDK/Chemistry, a software toolkit for quantum chemistry workflows targeting quantum computers. The toolkit addresses a key challenge in the field: while quantum algorith…

cs.CV2025

Get In Video: Add Anything You Want to the Video

Shaobin Zhuang, Zhipeng Huang, Binxin Yang +7

Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique v…

cs.AI2026

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hao Yu, Jiabo Zhan, Kang Liu +8

End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regi…

cond-mat.str-el2019

Finite temperature density matrix embedding theory

Chong Sun, Ushnish Ray, Zhi-Hao Cui +3

We describe a formulation of the density matrix embedding theory at finite temperature. We present a generalization of the ground-state bath orbital construction that embeds a mean…

cs.LG2024

Waveflow: boundary-conditioned normalizing flows applied to fermionic wavefunctions

Luca Thiede, Chong Sun, Alán Aspuru-Guzik

An efficient and expressive wavefunction ansatz is key to scalable solutions for complex many-body electronic structures. While Slater determinants are predominantly used for const…