6 papers
Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Calling
Jeffrey T. H. Wong, Zixi Zhang, Junyi Liu +1
Existing Multi-Agent Systems (MAS) typically rely on homogeneous model configurations, failing to exploit the diverse expertise inherent in different post-trained architectures. We…
Deep Kernel Fusion for Transformers
Zixi Zhang, Zhiwen Mo, Yiren Zhao +1
Agentic LLM inference with long contexts is increasingly limited by memory bandwidth rather than compute. In this setting, SwiGLU MLP blocks, whose large weights exceed cache capac…
Group Policy Gradient
Junhua Chen, Zixi Zhang, Hantao Zhong +1
We introduce Group Policy Gradient (GPG), a family of critic-free policy-gradient estimators for general MDPs. Inspired by the success of GRPO's approach in Reinforcement Learning…
Sample-Efficient Online Control Policy Learning with Real-Time Recursive Model Updates
Zixin Zhang, James Avtges, Todd D. Murphey
Data-driven control methods need to be sample-efficient and lightweight, especially when data acquisition and computational resources are limited -- such as during learning on hard…
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
Zixi Zhang, Balint Szekely, Pedro Gimenes +5
Hardware design verification (DV) is a process that checks the functional equivalence of a hardware design against its specifications, improving hardware reliability and robustness…
Unlocking the Global Synergies in Low-Rank Adapters
Zixi Zhang, Cheng Zhang, Xitong Gao +3
Low-rank Adaption (LoRA) has been the de-facto parameter-efficient fine-tuning technique for large language models. We present HeteroLoRA, a light-weight search algorithm that leve…