Audio & Voice
VisionStory supports 88 languages, including English, Chinese, Spanish, Arabic, Portuguese, Russian, Japanese, Punjabi, German, French, Korean, Turkish, Tamil, Vietnamese, Hindi, Bengali, Urdu, Persian, Italian, Indonesian, Thai, Marathi, Telugu, Ukrainian, Malay, Romanian, Polish, Dutch, Gujarati, and Kannada.
VisionStory provides a library of 1,000+ voices, which you can filter by gender, age, and use case. If you don’t find a voice that fits your needs, you can create a custom AI voice clone by uploading or recording your own audio.
Some languages have fewer voice options because those voices are specially optimized for that language. However, many English voices can speak multiple languages, so you still have flexibility when choosing a voice for your project.
Voice cloning lets you create a custom AI voice that imitates a specific voice by uploading or recording audio. To clone a voice, make sure your audio is recorded clearly in a quiet environment for the best results.
To use voice cloning in video generation, you’ll need to subscribe to the Pro plan or above.
Voice cloning is available in over 32 languages. The list of supported languages may change, so please check the voice cloning feature for the latest options. Please note: while cloning is free, you’ll need a subscription to use the cloned voice in video creation.
Preview audio lets you generate the speech for your talking video before creating the final video. This feature allows you to check the voice, pronunciation, and pauses to make sure they meet your expectations. You can adjust the voice as needed before generating the video, which uses credits. To use preview audio, you need to be subscribed to the Pro Plan or higher, and each plan comes with a different preview quota.
The stopwatch icon and +0.5s feature let you add a 0.5-second pause in the generated voice. You can use multiple stopwatch icons in a row to create longer pauses as needed in your video.
URL import lets you bring in audio by downloading and extracting it from a supported link for use in video creation. At this time, VisionStory supports links from YouTube and TikTok. If you’d like us to add support for other sites, please let us know. You can also use the voice changer feature to modify the imported audio while keeping the original content.
The remove noise feature helps eliminate background noise from your audio when you import or record it, ensuring clearer sound quality for your videos. To use this feature, you’ll need a Pro Plan or higher.
The voice changer feature lets you modify the voice in your audio, so you can create unique versions while keeping the original content. To use this feature, you’ll need a Pro Plan or higher.
The emotion in the voice is determined by the text you provide. As you use different wording, the text-to-speech (TTS) system will automatically apply the appropriate emotion, so there’s no need for extra controls.
When using the stopwatch feature, each stopwatch icon adds a 0.5-second pause. You can use them back-to-back to create longer pauses, up to a total of 3 seconds. However, it’s best not to use more than two consecutive pauses in a single text segment, as this could cause the AI to generate unexpected sounds or audio glitches.