Yandex has published the Alice AI Search Pretrain language model on Hugging Face, trained from scratch and forming the basis of AI responses in "Search".
The model is designed to process a large number of queries and documents with relatively low computational costs. It uses a hybrid Encoder-Decoder and Mixture of Experts (MoE) architecture. The former is responsible for selecting information and forming a concise answer, while the latter activates only the necessary parts of the neural network. With a total of 35 billion parameters, approximately 600 million parameters are used to process one token – less than 2%.
In blind tests, Alice AI Search Pretrain showed higher quality responses in Russian compared to Qwen 3.5 2B and 4B, as well as T5 Gemma 2 4B-4B Base. In terms of results, the model is comparable to Qwen 3.5 35B-A3B, while requiring fewer computational resources.
Since July, a fine-tuned version of the model has been used in Yandex Search. Responses appear directly below the search bar. According to the company, more than 49 million people use fast AI responses monthly.