Publications (13)
Alleviating the Long-Tail Problem in Conversational Recommender Systems
Zhipeng Zhao, Kun Zhou, Xiaolei Wang +4
Conversational recommender systems (CRS) aim to provide the recommendation service via natural language conversations. To develop an effective CRS, high-quality CRS datasets are ve…
Temperature-dependent Gilbert damping of Co2FeAl thin films with different degree of atomic order
Ankit Kumar, Fan Pan, Sajid Husain +5
Half-metallicity and low magnetic damping are perpetually sought for in spintronics materials and full Heusler alloys in this respect provide outstanding properties. However, it is…
Extended spin model in atomistic simulations of alloys
Fan Pan, Jonathan Chico, Anna Delin +2
An extended atomistic spin model allowing for studies of the finite temperature magnetic properties of alloys is proposed. The model is obtained by extending the Heisenberg Hamilto…
A systematic study of magnetodynamic properties at finite temperatures in doped permalloy from first principles calculations
Fan Pan, Jonathan Chico, Johan Hellsvik +3
By means of first principles calculations, we have systematically investigated how the magnetodynamic properties Gilbert damping, magnetization and exchange stiffness are affected…
Luxical: High-Speed Lexical-Dense Text Embeddings
DatologyAI, :, Luke Merrick +31
Frontier language model quality increasingly hinges on our ability to organize web-scale text corpora for training. Today's dominant tools trade off speed and flexibility: lexical…
ÃberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset
DatologyAI, :, Aldo Gael Carranza +32
Multilinguality is a core capability for modern foundation models, yet training high-quality multilingual models remains challenging due to uneven data availability across language…
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
DatologyAI, :, Siddharth Joshi +30
Empirical evaluation serves as the primary compass guiding research progress in foundation models. Despite a large body of work focused on training frontier vision-language models…
Interpretable Tsetlin Machine-based Premature Ventricular Contraction Identification
Jinbao Zhang, Xuan Zhang, Lei Jiao +3
Neural network-based models have found wide use in automatic long-term electrocardiogram (ECG) analysis. However, such black box models are inadequate for analysing physiological s…
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
DatologyAI, :, Pratyush Maini +28
Recent advances in large language model (LLM) pretraining have shown that simply scaling data quantity eventually leads to diminishing returns, hitting a data wall. In response, th…
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone
DatologyAI, :, Siddharth Joshi +32
Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less establi…
Improving Conversational Recommendation Systems via Counterfactual Data Simulation
Xiaolei Wang, Kun Zhou, Xinyu Tang +4
Conversational recommender systems (CRSs) aim to provide recommendation services via natural language conversations. Although a number of approaches have been proposed for developi…
Magnon properties of random alloys
Fan Pan, Anna Delin, Anders Bergman +1
We study magnon properties in terms of spin stiffness, Curie temperatures and magnon spectrum of Fe-Ni, Co-Ni and Fe-Co random alloys using a combination of electronic structure ca…
The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
Christina Baek, Ricardo Pio Monti, David Schwab +31
Real-world model deployments demand strong performance on narrow domains where data is often scarce. Typically, practitioners finetune models to specialize them, but this risks ove…