1 citations · 1 across the 1 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025★ 1 cited
ZeroLM: Data-Free Transformer Architecture Search for Language Models
Zhen-Song Chen, Hong-Wei Ding, Xian-Jia Wang +1
Neural architecture search (NAS) provides a systematic framework for automating the design of neural network architectures, yet its widespread adoption is hindered by prohibitive c…
cs.CL2024★ 1 cited
RakutenAI-7B: Extending Large Language Models for Japanese
Rakuten Group, Aaron Levine, Connie Huang +27
We introduce RakutenAI-7B, a suite of Japanese-oriented large language models that achieve the best performance on the Japanese LM Harness benchmarks among the open 7B models. Alon…