1 citations · 1 across the 9 of their papers we have counts for
9 papers
WinDOM: Self-Family Distillation for Small-Model GUI Grounding
Chengheng Li-Chen, Zhiqian Zhou, Hao Chen +1
Small (2B) GUI-grounding agents are attractive for on-device deployment, accessibility tooling, and low-cost iteration, but at this scale they face two open recipe questions:…
FiLM-Coordinated Dual-Branch Transformer for Global-Local Dependency Modeling in Language Modeling
Zhiqiang Zhou, Xu Ling, Junliang Dai
Standard Transformers use a single self-attention pathway to model both global dependencies and local patterns, creating tension between long-range structural reasoning and fine-gr…
Gen-VCoT: Generative Visual Chain-of-Thought Reasoning via Diffusion-Based RGB Intermediate Representations
Zhiqiang Zhou, Junliang Dai, Xu ling
Multimodal large language models (MLLMs) excel at visual reasoning but rely on text-based chain-of-thought (CoT), lacking interpretable visual intermediates. Existing methods use o…
TriAdReview: Triangular Adversarial Review Architecture for Multi-Model Technical Document Generation
Zhiqiang Zhou, Junliang Dai, Xu Ling
Large language models (LLMs) are increasingly used for technical document generation, yet single-model outputs often suffer from over-engineering, security blind spots, and incompl…
CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference
Xuezhen Xie, Zhiqiang Zhou
Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass. Multi-token prediction (MTP) offers a promising accelerati…
Feature Alignment Determines Fusion Strategy: A Comparative Study of Cross-Attention and Concatenation in Multimodal Learning
Zhiqiang Zhou, Xuezhen Xie
The choice between cross-attention and concatenation for multimodal fusion remains governed by practitioner intuition rather than principled understanding. In this paper, we demons…