5 papers
DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation
Peyman Hosseini, Ondrej Bohdal, Ahmed Alajrami +6
Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this ability typically depends on large models,…
Decoding Text Spans for Efficient and Accurate Named-Entity Recognition
Andrea Maracani, Savas Ozkan, Junyi Zhu +2
Named Entity Recognition (NER) is a key component in industrial information extraction pipelines, where systems must satisfy strict latency and throughput constraints in addition t…
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
Ahmed Bourouis, Savas Ozkan, Andrea Maracani +2
We tackle a new problem: generating geometrically consistent multi-view scenes from a single freehand sketch. Freehand sketches are the most geometrically impoverished input one co…
Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
Junyi Zhu, Savas Ozkan, Andrea Maracani +3
Deploying natural language processing (NLP) models on mobile platforms requires models that can adapt across diverse applications while remaining efficient in memory and computatio…
Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-Distillation
Andrea Maracani, Savas Ozkan, Sijun Cho +6
Scaling architectures have been proven effective for improving Scene Text Recognition (STR), but the individual contribution of vision encoder and text decoder scaling remain under…