7 papers
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference
Junyi Luo, Xinting Jiang, Tai-Hao Wen +9
Microscaling (MX) is now the standard for low-bit large language model (LLM) inference. Its 4-bit form MXFP4 still loses substantial accuracy, because existing MX formats fix eithe…
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
Zhaoyang Chu, Jiarui Hu, Xingyu Jiang +8
We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 ter…
LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
Xinting Jiang, Junyi Luo, Ruichen Qi +4
Sub-billion-parameter Transformer language models are increasingly deployed on edge devices, where the privacy, latency, and operating-cost advantages of on-device inference are co…
Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice
Xiusheng Huang, Xin Jiang, Jun Zhao +2
Accurate and effective discrete image tokenization is crucial for long image sequence processing. However, current methods rigidly compress all content at a fixed rate, ignoring th…
EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
Yiqun Yao, Naitong Yu, Xiang Li +7
We introduce EgoMem, the first lifelong memory agent tailored for full-duplex models that process real-time omnimodal streams. EgoMem enables real-time models to recognize multiple…
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
Yiqun Yao, Xiang Li, Xin Jiang +5
Full-duplex dialog models aim to listen and speak simultaneously, delivering rapid responses to dynamic user input. Among different solutions to full-duplexity, a native solution m…