papers

Publications (158)

cs.CV2025

DreamOmni2: Multimodal Instruction-based Editing and Generation

Bin Xia, Bohao Peng, Yuechen Zhang +10

Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical…

cs.CL2025

Behavioral Fingerprinting of Large Language Models

Zehua Pei, Hui-Ling Zhen, Ying Zhang +5

Current benchmarks for Large Language Models (LLMs) primarily focus on performance metrics, often failing to capture the nuanced behavioral characteristics that differentiate them.…

cs.SE2026

Agentic Electronic Design Automation: A Handoff Perspective

Jiawei Liu, Peiyi Han, Yuntao Lu +3

Electronic design automation (EDA) is inherently multi-stage and handoff-heavy. Design artifacts, flow scripts, and engineering decisions cross tool, session, and organizational bo…

cs.AR2024

Timing-driven Approximate Logic Synthesis Based on Double-chase Grey Wolf Optimizer

Xiangfei Hu, Yuyang Ye, Tinghuan Chen +2

With the shrinking technology nodes, timing optimization becomes increasingly challenging. Approximate logic synthesis (ALS) can perform local approximate changes (LACs) on circuit…

cs.CE2026

CSCO: A Backside-PDN-Aware Clock-Signal Co-Optimization Framework for Improved PPA

Zixiao Wang, Leilei Jin, Zhen Zhuang +2

The paper presents CSCO, a data‑driven framework that jointly allocates limited backside power‑delivery network resources between clock and signal nets to improve power, performanc…

#backside power delivery network#clock-signal co-optimization#ir-drop reduction#signal integrity
cs.LG2025

MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling

Yu Zhang, Hui-Ling Zhen, Mingxuan Yuan +1

Training large language models with FP8 formats offers significant efficiency gains. However, the reduced numerical precision of FP8 poses challenges for stable and accurate traini…

cs.AR2024

HDLdebugger: Streamlining HDL debugging with Large Language Models

Xufeng Yao, Haoyang Li, Tsz Ho Chan +5

In the domain of chip design, Hardware Description Languages (HDLs) play a pivotal role. However, due to the complex syntax of HDLs and the limited availability of online resources…

cs.CV2023

p-Laplacian Adaptation for Generative Pre-trained Vision-Language Models

Haoyuan Wu, Xinyun Zhang, Peng Xu +3

Vision-Language models (VLMs) pre-trained on large corpora have demonstrated notable success across a range of downstream tasks. In light of the rapidly increasing size of pre-trai…

cs.LG2026

Generalized Kullback-Leibler Divergence Loss

Jiequan Cui, Beier Zhu, Qingshan Xu +5

In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss…

cs.CV2025

Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?

Tianyuan Qu, Longxiang Tang, Bohao Peng +3

The rise of Large Vision-Language Models (LVLMs) has significantly advanced video understanding. However, efficiently processing long videos remains a challenge due to the ``Sampli…

cs.LG2024

MixPE: Quantization and Hardware Co-design for Efficient LLM Inference

Yu Zhang, Mingzi Wang, Lancheng Zou +4

Transformer-based large language models (LLMs) have achieved remarkable success as model sizes continue to grow, yet their deployment remains challenging due to significant computa…

physics.app-ph2024

Open-Source Differentiable Lithography Imaging Framework

Guojin Chen, Hao Geng, Bei Yu +1

The rapid evolution of the electronics industry, driven by Moore's law and the proliferation of integrated circuits, has led to significant advancements in modern society, includin…

cs.CL2024

Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA

Yuan Pu, Zhuolun He, Tairu Qiu +2

Retrieval augmented generation (RAG) enhances the accuracy and reliability of generative AI models by sourcing factual information from external databases, which is extensively emp…

cs.AR2014

Methodology for standard cell compliance and detailed placement for triple patterning lithography

Bei Yu, Xiaoqing Xu, Jhih-Rong Gao +1

As the feature size of semiconductor process further scales to sub-16nm technology node, triple patterning lithography (TPL) has been regarded one of the most promising lithography…

cs.CL2026

One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient

Rui Ming, Haoyuan Wu, Shoubo Hu +2

Supervised fine-tuning (SFT) is the predominant method for adapting large language models (LLMs), yet it often struggles with generalization compared to reinforcement learning (RL)…

cs.AI2024

Intelligent OPC Engineer Assistant for Semiconductor Manufacturing

Guojin Chen, Haoyu Yang, Bei Yu +1

Advancements in chip design and manufacturing have enabled the processing of complex tasks such as deep learning and natural language processing, paving the way for the development…

cs.SD2025

MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech

Chengyao Wang, Zhisheng Zhong, Bohao Peng +7

We present MGM-Omni, a unified Omni LLM for omni-modal understanding and expressive, long-horizon speech generation. Unlike cascaded pipelines that isolate speech synthesis, MGM-Om…

cs.LG2025

TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices

Jianlei Yang, Jiacheng Liao, Fanding Lei +6

Developing deep learning models on tiny devices (e.g. Microcontroller units, MCUs) has attracted much attention in various embedded IoT applications. However, it is challenging to…

cs.CL2026

FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

Zehua Pei, Hui-Ling Zhen, Xianzhi Yu +3

Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains limited. We trace this gap to h…

cs.CV2025

DreamOmni3: Scribble-based Editing and Generation

Bin Xia, Bohao Peng, Jiyang Liu +8

Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based ed…

cs.CL2025

Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts

Haoyuan Wu, Haoxing Chen, Xiaodong Chen +10

The Mixture of Experts (MoE) architecture is a cornerstone of modern state-of-the-art (SOTA) large language models (LLMs). MoE models facilitate scalability by enabling sparse para…

cs.CV2026

ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models

Yuqi Liu, Liangyu Chen, Jiazhen Liu +4

Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Learning with Verifiable Rewards (RL…

cs.LG2026

Class-frequency Guided Noise Schedule for Diffusion Models

Jiequan Cui, Beier Zhu, Qingshan Xu +3

In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, l…

cs.CL2025

On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language Understanding

Haoyuan Wu, Rui Ming, Jilong Gao +6

Large language models (LLMs) achieve remarkable performance in code generation tasks. However, a significant performance disparity persists between popular programming languages (e…

cs.AI2026

AutoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction

Shuo Ren, Yaohui Han, Yifan Shi +6

Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its difficulty is fixed, and ever…

cs.LG2026

KCLNet: Electrically Equivalence-Oriented Graph Representation Learning for Analog Circuits

Peng Xu, Yapeng Li, Tinghuan Chen +2

Digital circuits representation learning has made remarkable progress in the electronic design automation domain, effectively supporting critical tasks such as testability analysis…

cs.AI2026

AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits

Shuo Ren, Yaohui Han, Libo Shen +4

As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with no 2D analogue, and post-ro…

cs.CL2025

DiLA: Enhancing LLM Tool Learning with Differential Logic Layer

Yu Zhang, Hui-Ling Zhen, Zehua Pei +4

Considering the challenges faced by large language models (LLMs) in logical reasoning and planning, prior efforts have sought to augment LLMs with access to external solvers. While…

cs.CV2024

Generative Video Propagation

Shaoteng Liu, Tianyu Wang, Jui-Hsien Wang +8

Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative vid…

cs.CV2026

VisionZip: Longer is Better but Not Necessary in Vision Language Models

Senqiao Yang, Yukang Chen, Zhuotao Tian +4

Recent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them much longer than text tokens and significantly raisin…

cs.LG2025

Generative Distribution Distillation

Jiequan Cui, Beier Zhu, Qingshan Xu +6

In this paper, we formulate the knowledge distillation (KD) as a conditional generative problem and propose the \textit{Generative Distribution Distillation (GenDD)} framework. A n…

cs.LG2025

DiffPattern-Flex: Efficient Layout Pattern Generation via Discrete Diffusion

Zixiao Wang, Wenqian Zhao, Yunheng Shen +4

Recent advancements in layout pattern generation have been dominated by deep generative models. However, relying solely on neural networks for legality guarantees raises concerns i…

cs.AR2023

Floorplet: Performance-aware Floorplan Framework for Chiplet Integration

Shixin Chen, Shanyi Li, Zhen Zhuang +5

A chiplet is an integrated circuit that encompasses a well-defined subset of an overall system's functionality. In contrast to traditional monolithic system-on-chips (SoCs), chiple…

cs.AI2026

AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization

Shuo Ren, Zijin Cheng, Yaohui Han +7

Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation requires a costly EDA run thro…

cs.CL2026

Diversity or Precision? A Deep Dive into Next Token Prediction

Haoyuan Wu, Hai Wang, Jiajia Wu +5

Recent advancements have shown that reinforcement learning (RL) can substantially improve the reasoning abilities of large language models (LLMs). The effectiveness of such RL trai…

cs.CV2023

DevelSet: Deep Neural Level Set for Instant Mask Optimization

Guojin Chen, Ziyang Yu, Hongduo Liu +2

With the feature size continuously shrinking in advanced technology nodes, mask optimization is increasingly crucial in the conventional design flow, accompanied by an explosive gr…

cs.CV2026

Low-Light Video Enhancement with An Effective Spatial-Temporal Decomposition Paradigm

Xiaogang Xu, Kun Zhou, Tao Hu +4

Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition s…

cs.CR2020

Attacking Split Manufacturing from a Deep Learning Perspective

Haocheng Li, Satwik Patnaik, Abhrajit Sengupta +5

The notion of integrated circuit split manufacturing which delegates the front-end-of-line (FEOL) and back-end-of-line (BEOL) parts to different foundries, is to prevent overproduc…

cs.LG2025

AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs

Yao Lai, Souradip Poddar, Sungyoung Lee +5

Despite recent advances, analog front-end design still relies heavily on expert intuition and iterative simulations, which limits the potential for automation. We present AnalogCod…

cs.CV2026

UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings

Jiajun Qin, Yuan Pu, Zhuolun He +3

Current vision-language models have been explored for multi-modal embedding tasks like information retrieval. However, they face significant challenges in real-world queries and ta…

cs.CV2022

DSGN++: Exploiting Visual-Spatial Relation for Stereo-based 3D Detectors

Yilun Chen, Shijia Huang, Shu Liu +2

Camera-based 3D object detectors are welcome due to their wider deployment and lower price than LiDAR sensors. We first revisit the prior stereo detector DSGN for its stereo volume…

cs.AR2014

Layout decomposition for triple patterning lithography

Bei Yu, Kun Yuan, Boyang Zhang +2

As minimum feature size and pitch spacing further decrease, triple patterning lithography (TPL) is a possible 193nm extension along the paradigm of double patterning lithography (D…

cs.AR2024

The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models

Lei Chen, Yiqi Chen, Zhufei Chu +36

Within the Electronic Design Automation (EDA) domain, AI-driven solutions have emerged as formidable tools, yet they typically augment rather than redefine existing methodologies.…

cs.AR2024

The Survey of Chiplet-based Integrated Architecture: An EDA perspective

Shixin Chen, Hengyuan Zhang, Zichao Ling +2

Enhancing performance while reducing costs is the fundamental design philosophy of integrated circuits (ICs). With advancements in packaging technology, interposer-based chiplet ar…

cs.CL2026

Scaling Native Multimodal Pre-Training From Scratch

Haoyuan Wu, Aoqi Wu, Hai Wang +3

Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world.…

cs.CL2026

MemDLM: Memory-Enhanced DLM Training

Zehua Pei, Hui-Ling Zhen, Weizhe Lin +4

Diffusion Language Models (DLMs) offer attractive advantages over Auto-Regressive (AR) models, such as full-attention parallel decoding and flexible generation. However, standard D…

cs.CV2023

PFENet++: Boosting Few-shot Semantic Segmentation with the Noise-filtered Context-aware Prior Mask

Xiaoliu Luo, Zhuotao Tian, Taiping Zhang +3

In this work, we revisit the prior mask guidance proposed in ``Prior Guided Feature Enrichment Network for Few-Shot Segmentation''. The prior mask serves as an indicator that highl…

cs.NE2025

Evolution of Optimization Algorithms for Global Placement via Large Language Models

Xufeng Yao, Jiaxi Jiang, Yuxuan Zhao +3

Optimization algorithms are widely employed to tackle complex problems, but designing them manually is often labor-intensive and requires significant expertise. Global placement is…

cs.CV2021

Conditional Temporal Variational AutoEncoder for Action Video Prediction

Xiaogang Xu, Yi Wang, Liwei Wang +2

To synthesize a realistic action sequence based on a single human image, it is crucial to model both motion patterns and diversity in the action video. This paper proposes an Actio…

cs.CL2025

ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving

Haoyuan Wu, Xueyi Chen, Rui Ming +4

Large language models (LLMs) demonstrate significant reasoning capabilities, particularly through long chain-of-thought (CoT) processes, which can be elicited by reinforcement lear…

cs.CV2025

Modular Customization of Diffusion Models via Blockwise-Parameterized Low-Rank Adaptation

Mingkang Zhu, Xi Chen, Zhongdao Wang +3

Recent diffusion model customization has shown impressive results in incorporating subject or style concepts with a handful of images. However, the modular composition of multiple…

cs.CV2025

Low-Light Video Enhancement via Spatial-Temporal Consistent Decomposition

Xiaogang Xu, Kun Zhou, Tao Hu +4

Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition s…

cs.LG2026

PreMoE: Proactive Inference for Efficient Mixture-of-Experts

Zehua Pei, Ying Zhang, Hui-Ling Zhen +6

Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization.…

cs.LG2025

PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models

Lancheng Zou, Shuo Yin, Zehua Pei +3

Channel permutation is a powerful technique for enhancing the accuracy of N:M sparse models by reordering the channels of weight matrices to prioritize the retention of important w…

cs.LG2019

Are Adversarial Perturbations a Showstopper for ML-Based CAD? A Case Study on CNN-Based Lithographic Hotspot Detection

Kang Liu, Haoyu Yang, Yuzhe Ma +5

There is substantial interest in the use of machine learning (ML) based techniques throughout the electronic computer-aided design (CAD) flow, particularly those based on deep lear…

cs.OH2019

CAD Tool Design Space Exploration via Bayesian Optimization

Yuzhe Ma, Ziyang Yu, Bei Yu

The design complexity is increasing as the technology node keeps scaling down. As a result, the electronic design automation (EDA) tools also become more and more complex. There ar…

cs.DC2025

AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism

Wendong Xu, Chujie Chen, He Xiao +8

Large Language Model (LLM) inference services demand exceptionally high availability and low latency, yet multi-GPU Tensor Parallelism (TP) makes them vulnerable to single-GPU fail…

cs.AR2026

CLIP-3D: Closed-Loop Evaluation of Performance and Physical Constraints for 3D ICs

Shuo Ren, Libo Shen, Yaohui Han +6

The paper introduces CLIP-3D, a methodology that integrates power, cache, and thermal models with architectural simulation to evaluate the performance and physical constraints of 3…

#3d ic design#thermal-aware floorplanning#performance evaluation#architectural simulation
cs.LG2025

Circuit Representation Learning with Masked Gate Modeling and Verilog-AIG Alignment

Haoyuan Wu, Haisheng Zheng, Yuan Pu +1

Understanding the structure and function of circuits is crucial for electronic design automation (EDA). Circuits can be formulated as And-Inverter graphs (AIGs), enabling efficient…

cs.LG2024

MoreauPruner: Robust Pruning of Large Language Models against Weight Perturbations

Zixiao Wang, Jingwei Zhang, Wenqian Zhao +2

Few-shot gradient methods have been extensively utilized in existing model pruning methods, where the model weights are regarded as static values and the effects of potential weigh…

cs.AR2026

CPPL: A Circuit Prompt Programming Language

Shuo Yin, Yihe Wang, Lancheng Zou +6

Large language models (LLMs) have shown promise in register-transfer level (RTL) design automation, but direct RTL generation remains difficult to validate, optimize, and integrate…

cs.CV2023

AdaOPC: A Self-Adaptive Mask Optimization Framework For Real Design Patterns

Wenqian Zhao, Xufeng Yao, Ziyang Yu +4

Optical proximity correction (OPC) is a widely-used resolution enhancement technique (RET) for printability optimization. Recently, rigorous numerical optimization and fast machine…

cs.AR2022

Eventor: An Efficient Event-Based Monocular Multi-View Stereo Accelerator on FPGA Platform

Mingjun Li, Jianlei Yang, Yingjie Qi +6

Event cameras are bio-inspired vision sensors that asynchronously represent pixel-level brightness changes as event streams. Event-based monocular multi-view stereo (EMVS) is a tec…

cs.LG2025

Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents

Mingkang Zhu, Xi Chen, Bei Yu +2

Large language model (LLM) agents increasingly rely on external tools such as search engines to solve complex, multi-step problems, and reinforcement learning (RL) has become a key…

cs.AR2020

DAMO: Deep Agile Mask Optimization for Full Chip Scale

Guojin Chen, Wanli Chen, Yuzhe Ma +2

Continuous scaling of the VLSI system leaves a great challenge on manufacturing and optical proximity correction (OPC) is widely applied in conventional design flow for manufactura…

cs.CL2024

Towards Versatile and Efficient Visual Knowledge Integration into Pre-trained Language Models with Cross-Modal Adapters

Xinyun Zhang, Haochen Tan, Han Wu +1

Humans learn language via multi-modal knowledge. However, due to the text-only pre-training scheme, most existing pre-trained language models (PLMs) are hindered from the multi-mod…

cs.AI2024

Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks

Haoyuan Wu, Haisheng Zheng, Zhuolun He +1

Large language models (LLMs) have demonstrated considerable proficiency in general natural language processing (NLP) tasks. Instruction tuning, a successful paradigm, enhances the…

cs.AR2014

GLOW: A global router for low-power thermal-reliable interconnect synthesis using photonic wavelength multiplexing

Duo Ding, Bei Yu, David Z. Pan

In this paper, we examine the integration potential and explore the design space of low power thermal reliable on-chip interconnect synthesis featuring nanophotonics Wavelength Div…

cs.CV2024

Decoupled Kullback-Leibler Divergence Loss

Jiequan Cui, Zhuotao Tian, Zhisheng Zhong +3

In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss…

cs.CV2025

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models

Yuqi Liu, Qin Jin, Tianyuan Qu +4

Understanding accurate atomic temporal event is essential for video comprehension. However, current video-language benchmarks often fall short to evaluate Large Multi-modal Models'…

physics.soc-ph2025

Causal Language in Observational Studies: Sociocultural Backgrounds and Team Composition

Jun Wang, Bei Yu

The use of causal language in observational studies has raised concerns about overstatement in scientific communication. While some argue that such language should be reserved for…

cs.AR2014

Triple Patterning Lithography (TPL) Layout Decomposition using End-Cutting

Bei Yu, Jhih-Rong Gao, David Z. Pan

Triple patterning lithography (TPL) is one of the most promising techniques in the 14nm logic node and beyond. However, traditional LELELE type TPL technology suffers from native c…

cs.CV2024

FlexPose: Pose Distribution Adaptation with Limited Guidance

Zixiao Wang, Junwu Weng, Mengyuan Liu +1

Numerous well-annotated human key-point datasets are publicly available to date. However, annotating human poses for newly collected images is still a costly and time-consuming pro…

cs.AR2025

Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToT

Zhiteng Chao, Yonghao Wang, Xinyu Zhang +9

Large language models (LLMs) hold promise for automating integrated circuit (IC) engineering using register transfer level (RTL) hardware description languages (HDLs) like Verilog.…

cs.OH2015

E-BLOW: E-Beam Lithography Overlapping aware Stencil Planning for MCC System

Bei Yu, Kun Yuan, Jhih-Rong Gao +1

Electron beam lithography (EBL) is a promising maskless solution for the technology beyond 14nm logic node. To overcome its throughput limitation, industry has proposed character p…

cs.CV2023

Generalized Parametric Contrastive Learning

Jiequan Cui, Zhisheng Zhong, Zhuotao Tian +3

In this paper, we propose the Generalized Parametric Contrastive Learning (GPaCo/PaCo) which works well on both imbalanced and balanced data. Based on theoretical analysis, we obse…

cs.LG2025

From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion

Zehua Pei, Hui-Ling Zhen, Xianzhi Yu +3

Structured pruning of Generative Pre-trained Transformers (GPTs) offers a promising path to efficiency but often suffers from irreversible performance degradation due to the discar…

cs.AR2014

TRIAD: a triple patterning lithography aware detailed router

Yen-Hung Lin, Bei Yu, David Z. Pan +1

TPL-friendly detailed routers require a systematic approach to detect TPL conflicts. However, the complexity of conflict graph (CG) impedes directly detecting TPL conflicts in CG.…

cs.AI2026

SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

Zehua Pei, Hui-Ling Zhen, Shixiong Kai +4

Large Language Model (LLM) agents are increasingly deployed in environments that generate massive, dynamic contexts. However, a critical bottleneck remains: while agents have acces…

cs.CV2025

DreamOmni: Unified Image Generation and Editing

Bin Xia, Yuechen Zhang, Jingyao Li +5

Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster…

cs.LG2024

Classes Are Not Equal: An Empirical Study on Image Recognition Fairness

Jiequan Cui, Beier Zhu, Xin Wen +3

In this paper, we present an empirical study on image recognition fairness, i.e., extreme class accuracy disparity on balanced data like ImageNet. We experimentally demonstrate tha…

cs.AR2014

EPIC: Efficient prediction of IC manufacturing hotspots with a unified meta-classification formulation

Duo Ding, Bei Yu, Joydeep Ghosh +1

In this paper we present EPIC, an efficient and effective predictor for IC manufacturing hotspots in deep sub-wavelength lithography. EPIC proposes a unified framework to combine d…

cs.CV2021

Parametric Contrastive Learning

Jiequan Cui, Zhisheng Zhong, Shu Liu +2

In this paper, we propose Parametric Contrastive Learning (PaCo) to tackle long-tailed recognition. Based on theoretical analysis, we observe supervised contrastive loss tends to b…

cs.LG2019

A Unified Approximation Framework for Compressing and Accelerating Deep Neural Networks

Yuzhe Ma, Ran Chen, Wei Li +4

Deep neural networks (DNNs) have achieved significant success in a variety of real world applications, i.e., image classification. However, tons of parameters in the networks restr…

cs.LG2025

From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models

Mingkang Zhu, Xi Chen, Bei Yu +2

Large reasoning models (LRMs) generate intermediate reasoning traces before producing final answers, yielding strong gains on multi-step and mathematical tasks. Yet aligning LRMs w…

cs.CV2026

RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing

Tianyuan Qu, Lei Ke, Xiaohang Zhan +6

The paper presents RePlan, a framework that first reasons about natural‑language instructions to identify specific image regions and then edits those regions using a diffusion mode…

#instruction-based image editing#region planning#vision-language reasoning#diffusion models
cs.AR2014

A High-Performance Triple Patterning Layout Decomposer with Balanced Density

Bei Yu, Yen-Hung Lin, Gerard Luk-Pat +3

Triple patterning lithography (TPL) has received more and more attentions from industry as one of the leading candidate for 14nm/11nm nodes. In this paper, we propose a high perfor…

cs.AR2018

Adaptive 3D-IC TSV Fault Tolerance Structure Generation

Song Chen, Qi Xu, Bei Yu

In three dimensional integrated circuits (3D-ICs), through silicon via (TSV) is a critical technique in providing vertical connections. However, the yield and reliability is one of…

cs.AR2026

Chiplet3D: Pin- and Thermal-Aware 3D Chiplet Floorplanning via Convolution-Embedded MILP

Shuo Ren, Libo Shen, Yaohui Han +4

As traditional Moore's Law scaling slows down, 3D-ICs stack multiple active dies vertically to sustain performance scaling. However, this vertical stacking traps heat inside, makin…

cs.CV2025

DreamVE: Unified Instruction-based Image and Video Editing

Bin Xia, Jiyang Liu, Yuechen Zhang +6

Instruction-based editing holds vast potential due to its simple and efficient interactive editing format. However, instruction-based editing, particularly for video, has been cons…

cs.AI2026

Graph World Models: Concepts, Taxonomy, and Future Directions

Jiawei Liu, Senqiao Yang, Mingjun Wang +2

As one of the mainstream models of artificial intelligence, world models allow agents to learn the representation of the environment for efficient prediction and planning. However,…

cs.AR2023

Analytical Die-to-Die 3D Placement with Bistratal Wirelength Model and GPU Acceleration

Peiyu Liao, Yuxuan Zhao, Dawei Guo +2

In this paper, we present a new analytical 3D placement framework with a bistratal wirelength model for F2F-bonded 3D ICs with heterogeneous technology nodes based on the electrost…

cs.AR2014

L-Shape based Layout Fracturing for E-Beam Lithography

Bei Yu, Jhih-Rong Gao, David Z. Pan

Layout fracturing is a fundamental step in mask data preparation and e-beam lithography (EBL) writing. To increase EBL throughput, recently a new L-shape writing strategy is propos…

cs.AR2026

LongRTL: Graph-Similarity-Guided LLM-driven Long Context RTL Optimization

Yuyang Ye, Che-Kuan Shen, Xiangfei Hu +5

Large Language Models (LLMs) show great promise in RTL code generation and optimization. However, real-world RTL designs are typically long, entangled, and poorly modularized, posi…

cs.AR2014

Voltage and Level-Shifter Assignment Driven Floorplanning

Bei Yu, Sheqin Dong, Song Chen +1

Low Power Design has become a significant requirement when the CMOS technology entered the nanometer era. Multiple-Supply Voltage (MSV) is a popular and effective method for both d…

cs.LG2026

Consistent Distributed Ranking of Generative Models via Kernel Distances

Zixiao Wang, Farzan Farnia, Zhenghao Lin +2

Ranking generative models based on the fidelity and diversity of their outputs is required to identify the best generator in a group of candidate generative AI models. To rank a gr…

cs.RO2026

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models

Zixuan Wang, Yuxin Chen, Yuqi Liu +6

Vision-Language-Action (VLA) models typically map visual observations and linguistic instructions directly to control signals. This "black-box" mapping forces a single forward pass…

cs.AI2024

LLM-Enhanced Bayesian Optimization for Efficient Analog Layout Constraint Generation

Guojin Chen, Keren Zhu, Seunggeun Kim +4

Analog layout synthesis faces significant challenges due to its dependence on manual processes, considerable time requirements, and performance instability. Current Bayesian Optimi…

cs.LG2021

Routing Towards Discriminative Power of Class Capsules

Haoyu Yang, Shuhe Li, Bei Yu

Capsule networks are recently proposed as an alternative to modern neural network architectures. Neurons are replaced with capsule units that represent specific features or entitie…

cs.AR2014

Multi-Voltage and Level-Shifter Assignment Driven Floorplanning

Bei Yu, Sheqin Dong, Statoshi Goto

As technology scales, low power design has become a significant requirement for SOC designers. Among the existing techniques, Multiple-Supply Voltage (MSV) is a popular and effecti…