description Amazon Polly (Text-to-Speech) Overview
While technically a Text-to-Speech (TTS) service, Polly is included as it is often confused with STT and is a key component of cloud voice solutions. It allows users to convert text into highly natural-sounding audio files. It is crucial for creating audiobooks or automated voice narration where the output must sound human-generated.
help Amazon Polly (Text-to-Speech) FAQ
Is Amazon Polly text-to-speech or speech-to-text?
Polly is strictly text-to-speech — it converts written text into synthesized audio. If you need the reverse, speech-to-text, Amazon's corresponding service is Amazon Transcribe, and the two are frequently confused in cloud voice architectures.
How is Amazon Polly priced?
Polly is usage-based, charging per million characters synthesized, with different rates for standard, neural, and long-form voices. New AWS accounts get a free tier covering millions of characters per month for the first 12 months, which is enough for small audiobook or app-voice pilots.
Does Amazon Polly support SSML for controlling pronunciation and pauses?
Yes, Polly has extensive SSML support, letting you insert pauses, phonetic pronunciations, breathing effects, and newscaster-style emphasis. Neural voices support a subset of SSML but deliver substantially more natural prosody than the standard voices.
Can I use Amazon Polly to produce an audiobook?
Yes — Polly outputs MP3 and OGG audio and gained long-form neural voices in the early 2020s specifically aimed at lengthy content like audiobooks. It integrates with other AWS services such as Lambda and S3 for automated batch conversion of manuscripts into audio files.
explore Explore More
Similar to Amazon Polly (Text-to-Speech)
compare_arrows Compare: Google Cloud Speech-to-Text... See all arrow_forwardReviews & Comments
Write a Review
Be the first to review
Share your thoughts with the community and help others make better decisions.