6 papers
An AI4AI Framework for Visual Token Pruning
Zhen Liu, Wenli Huang, Wei Song +3
Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and…
GraphIR: Architecture-Level Search States for LLM-Guided Neural Architecture Evolution
Zhen Liu, Wanqi Zhou, Shuanghao Bai +3
Large language models (LLMs) enable neural architecture search (NAS) directly over executable neural network programs. However, code-level flexibility does not provide the architec…
Overcoming the Weakest-Link Effect in LLM-Driven Program Optimization via Heterogeneous Edit Recombination
Jingwen Fu, Zhen Liu, Yuhan Liu +2
Large language models (LLMs) are increasingly used to solve complex problems by searching over program space, offering a general paradigm for scientific problems that can be natura…
Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search
Zhen Liu, Yuhan Liu, Jinjun Wang +3
This paper focuses on a key challenge in Neural Architecture Search (NAS): integrating established architectural knowledge while exploring new designs under expensive evaluations.…
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
Kangyi Wu, Pengna Li, Jingwen Fu +4
Emotional talking face generation aims to animate a human face in given reference images and generate a talking video that matches the content and emotion of driving audio. However…
Semantic-aware Representation Learning for Homography Estimation
Yuhan Liu, Qianxin Huang, Siqi Hui +5
Homography estimation is the task of determining the transformation from an image pair. Our approach focuses on employing detector-free feature matching methods to address this iss…