The system combines large language models and diffusion synthesis
The Russian company Neuro.net has introduced a new generation of Neuro TTS speech synthesis technology for AI voice solutions. The development is designed for integration into business services via API and can be used in AI voice agents.

At the core of Neuro TTS is an architecture that combines large language models and diffusion synthesis. This approach allows for considering text context and acoustic features of the voice. Upon integration, companies can customize the sound character of the virtual assistant.
The technology was created considering Neuro.net's experience in automating contact centers, where speech synthesis is part of the full dialogue cycle: the system recognizes the client's request, determines their intent, forms a response, and vocalizes it.
One of the features of Neuro TTS is streaming generation. The first audio data begins to arrive approximately 120 ms after the request, so playback of a replica can be started even before its generation is complete.
Neuro TTS is already available for use in production services. You can familiarize yourself with the voices and test the technology on the speech.neuro.net platform.