9 papers
HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
Anyang Tong, Xiang Niu, ZhiPing Liu +4
Existing multimodal Retrieval-Augmented Generation (RAG) methods for visually rich documents (VRD) are often biased towards retrieving salient knowledge(e.g., prominent text and vi…
Understanding Network Behaviors through Natural Language Question-Answering
Mingzhe Xing, Chang Tian, Jianan Zhang +4
Modern large-scale networks introduce significant complexity in understanding network behaviors, increasing the risk of misconfiguration. Prior work proposed to understand network…
Controllable Stylistic Text Generation with Train-Time Attribute-Regularized Diffusion
Fan Zhou, Chang Tian, Tim Van de Cruys
Generating stylistic text with specific attributes is a key problem in controllable text generation. Recently, diffusion models have emerged as a powerful paradigm for both visual…
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
Aleksa Jelaca, Ying Jiao, Chang Tian +1
Text-to-image generation has advanced rapidly with large-scale multimodal training, yet fine-grained controllability remains a critical challenge. Counterfactual controllability, d…
Structured Information for Improving Spatial Relationships in Text-to-Image Generation
Sander Schildermans, Chang Tian, Ying Jiao +1
Text-to-image (T2I) generation has advanced rapidly, yet faithfully capturing spatial relationships described in natural language prompts remains a major challenge. Prior efforts h…
Using Causality for Enhanced Prediction of Web Traffic Time Series
Chang Tian, Mingzhe Xing, Zenglin Shi +3
Predicting web service traffic has significant social value, as it can be applied to various practical scenarios, including but not limited to dynamic resource scaling, load balanc…