Skip to main content
The tts() method converts text into spoken audio using AI voice models from OpenAI, ElevenLabs, and other providers.

Basic Text-to-Speech

Voice Selection

Choose from different voices:

Available Voices

Try different voices to find the best match for your content. Use alloy or echo for professional content, fable for narratives, and nova for upbeat messages.

Audio Format

Control the output audio format:

Format Options

Speech Speed

Adjust playback speed:

Speed Guide

Speeds below 0.5 or above 2.0 may sound unnatural. Stick to 0.75-1.5 for most use cases.

Parameters

Request Parameters

Response

Best Practices

1. Format Text for Speech

2. Use SSML for Control

Some models support SSML (Speech Synthesis Markup Language):

3. Choose Appropriate Format

4. Always Handle Errors

5. Monitor Costs

Examples

Podcast Intro

Product Announcement

Educational Content

Audiobook Narration

Voice Assistant Response

Multilingual Content

Advanced Usage

Batch Generation

Generate multiple audio files:

Download Audio File

Base64 Audio Embedding

Override Model

Use a specific TTS model:

Custom Metadata

Track TTS generation:

Use Cases

E-Learning Platforms

Convert course content to audio for accessibility:

IVR Systems

Generate prompts for phone systems:

Content Creation

Add voiceovers to videos:

Accessibility

Make written content accessible:

Next Steps

Video

Add audio to video content

Chat

Generate text for TTS

Gates & Routing

How Verlon routes TTS requests

Cost Tracking

Monitor TTS costs