Kakao has advanced the voice generation technology of its proprietary omni artificial intelligence (AI) model, Kanana-o. The enhancements allow the system to control not just natural reading but also tone, emotion, speed, and volume based on user natural language commands.
On August 4, Kakao unveiled the voice generation technology of Kanana-o through its tech blog. The new model generates voice based on the desired speaking style input by users in natural language.
For instance, it can understand commands like "read very quickly," "read in a low voice," "read in a sad voice," or "read in a Gyeongsang dialect," adjusting speed, volume, pitch, emotion, and intonation accordingly.
It can also perform role-based commands such as "like a sports commentator," "like an announcer," or "like reading a storybook." The system can handle complex commands that include multiple conditions, such as "read quickly in a sad voice with a lower tone." While primarily trained in Korean, it can also execute the same speaking commands in English.
The performance has also improved. Kanana-o scored 94.50 on the Korean benchmark for evaluating command execution in voice generation, surpassing the GPT-4 Mini TTS, which scored 91.10.
Additionally, the speed and efficiency of voice generation have been enhanced. Kakao applied its proprietary voice tokenizer, LM-SPT, which compresses voice into fewer tokens, reducing the amount of data the AI needs to process and increasing voice generation speed.
LM-SPT has demonstrated excellent performance in evaluations comparing the voice understanding and generation capabilities of various language models in both Korean and English.
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.

