Amazon Transcribe FAQs
General
Open allAmazon Transcribe is an AWS Artificial Intelligence (AI) service that makes it easy for you to convert speech to text. Using Automatic Speech Recognition (ASR) technology, you can use Amazon Transcribe for a variety of business applications, including transcription of voice-based customer service calls, generation of subtitles on audio/video content, and conduct (text-based) content analysis on audio/video content.
Amazon Transcribe converts audio input into text, which opens the door for various text analytics applications on voice input. For instance, by using Amazon Comprehend on the converted text data from Amazon Transcribe, you can perform sentiment analysis or extract entities and key phrases. Similarly, by integrating with Amazon Translate and Amazon Polly, you can accept voice input in one language, translate it into another, and generate voice output, effectively enabling multilingual conversations. It is also possible to integrate Amazon Transcribe with Amazon Kendra or Amazon OpenSearch to index and perform text-based search across an audio/video library. To learn more, check out the Live Call Analytics and Agent Assist, Post Call Analytics, MediaSearch, or Content Analysis solution.
Amazon Transcribe is designed to handle a wide range of speech and acoustic characteristics, including variations in volume, pitch, and speaking rate. The quality and content of the audio signal (including but not limited to factors such as background noise, overlapping speakers, accented speech, or switches between languages within a single audio file) may affect the accuracy of service output. We are constantly updating the service to improve its ability to accommodate additional acoustic variation and content types.
Using Amazon Transcribe
Open allThe easiest way to get started is to submit a job using the console to transcribe an audio file. You can also call the service directly from the AWS Command Line Interface, or use one of the supported SDKs of your choice to integrate with your applications. Either way, you can start using Amazon Transcribe to generate automated transcripts for your audio files with just a few lines of code.
Yes. Amazon Transcribe allows you to open a bidirectional stream over HTTP2. You can send an audio stream to the service while receiving a text stream in return in real time. Please refer to the documentation page for more details.
Supported media types differ between batch transcriptions and streaming transcriptions, though lossless formats are recommended for both. Please refer to the documentation page for more details.
For information on language support, please refer to this documentation page.
Amazon Transcribe for the most part is device agnostic. In general, it works with any device that includes an on-device microphone such as phones, PCs, tablets, and IoT devices (such as car audio systems). Amazon Transcribe API will be able to detect the quality of the audio stream being input at the device (8kHz VS 16kHz) and will appropriately select the acoustic models for converting speech to text. Furthermore, developers can call Amazon Transcribe API through their applications to access speech-to-text conversion capability.
Amazon Transcribe service calls are limited to four hours (or 2 GB) per API call for our batch service. The streaming service can accommodate open connections up to four hours long.
Amazon Transcribe batch service supports .NET, Go, Java, JavaScript, PHP, Python, and Ruby. Amazon Transcribe real-time service supports Java SDK, Ruby SDK, and C++ SDK. Additional SDK support is coming. For more details, visit the Resources and documentation page.
The speech recognition output depends on a number of factors in addition to custom vocabulary entries, so there can be no assurance that if a term is included in the custom vocabulary, it will be correctly recognized. However, one of the most frequent reasons for this is that the words are pronounced significantly different than they are written.
In these cases, it is recommended to create multiple phrase entries in the custom vocabulary file for the same word to cover for the possible variations in pronounciation. You may use DisplayAs column for the desired output for these phrase entries. Refer to the custom vocabulary documentation for more information.
Custom vocabularies are optimized for a small list of targeted words; larger vocabularies may lead to over-generation of custom words, especially when they contain words that are pronounced in a similar way. If you have a large list, please try reducing it to rare words and words that are actually expected to occur in your audio files. If you have a large vocabulary covering multiple use cases, split it into separate lists for different use cases. The words that are short and sound similar to many other words may lead to over-generation (too many custom words appearing in the output). It is preferable to combine these words with surrounding words and list them as hyphen-separated phrases. For example, the custom word “A.D.” could be included as part of a phrase such as “A.D.-converter.”
Yes, Amazon Transcribe supports automatic content redaction or PII redaction for both batch and streaming APIs.
Please refer to the Amazon Transcribe documentation for information on the language availability of automatic content redaction / PII redaction.
No, this feature does not remove sensitive personal information from the source audio. However, Amazon Transcribe Call Analytics removes sensitive personal information from both the transcripts and the source audio. Visit this link for more details on how call analytics can redact audio. You can also redact personal information from the source audio yourself using the start and end timestamps that are provided in the redacted transcripts for each instance of an identified PII utterance. Please refer to this audio redaction solution for standard Transcribe APIs.
However, the specialized Amazon Transcribe Call Analytics APIs remove sensitive personal information from both the transcripts and the source audio. To learn more, review the Call Analytics audio redaction documentation.
No, automatic content redaction only works on audio as an input.
Automatic content redaction is designed to identify and remove personally identifiable information (PII), but due to the predictive nature of machine learning, it may not identify and remove all instances of PII in a transcript generated by the service. You should review any output provided by Automatic content redaction to ensure it meets your needs.
Yes, there are two additional capabilities supported by automatic content redaction for the streaming API that are not supported by the batch API. You can decide to only identify PII and not redact when using content redaction with streaming API. Also you have the ability to identify or redact specific PII types with streaming API. For example, you can redact just the social security number and credit card information and keep other PII like names and email addresses.
Please refer to the Amazon Transcribe documentation for information on the availability of automatic content redaction and PII redaction for batch and streaming APIs in the AWS Regions.
Automatic language identification is currently supported for batch and streaming APIs.
Amazon Transcribe can identify any of the languages supported by the batch and streaming APIs. Go here for details on supported languages and language-specific features.
Amazon Transcribe supports multi-language ID for batch. See this link for more details.
Yes, you can specify a list of languages that might be present in your media library. When you provide a list of languages, the identified language will be chosen from that list. If no languages are specified, the system will process the audio file against all the languages supported by Amazon Transcribe and select the most probable one. The accuracy of language identification is better when a select list of languages is provided. See this link for more details.
Pricing and availability
Open allRefer to the Amazon Transcribe Pricing page to learn more.
Please refer to the AWS Global Infrastructure Region Table. Go here for additional details on Amazon Transcribe endpoints and quotas.
Data Privacy
Open allAmazon Transcribe may store and use voice inputs processed by the service solely to provide and maintain the service and to improve and develop the quality of Amazon Transcribe and other Amazon machine-learning/artificial-intelligence technologies. Use of your content is important for continuous improvement of your Amazon Transcribe customer experience, including the development and training of related technologies. We do not use any personally identifiable information that may be contained in your content to target products, services, or marketing to you or your end users. Your trust, privacy, and the security of your content are our highest priority, and we implement appropriate and sophisticated technical and physical controls, including encryption at rest and in transit, designed to prevent unauthorized access to, or disclosure of, your content and ensure that our use complies with our commitments to you. Please see https://aws.amazon.com/compliance/data-privacy-faq/ for more information. You may opt out of having your content used to improve and develop the quality of Amazon Transcribe and other Amazon machine-learning/artificial-intelligence technologies by using an AWS Organizations opt-out policy. For information about how to opt out, see AI services opt-out policy.
Yes. You can use available Delete APIs to delete data and other artifacts associated with transcription jobs. If you have issues doing so, contact AWS support.
Only authorized employees will have access to your content that is processed by Amazon Transcribe. Your trust, privacy, and the security of your content are our highest priority, and we implement appropriate and sophisticated technical and physical controls, including encryption at rest and in transit, designed to prevent unauthorized access to, or disclosure of, your content and ensure that our use complies with our commitments to you. Please see https://aws.amazon.com/compliance/data-privacy-faq/ for more information.
You always retain ownership of your content, and we will only use your content with your consent.
When submitting text data that is used to train a dedicated model, you have ownership of the original text data and the generated custom model. The text data will neither be stored, nor used to improve our general speech recognition engine. Models produced by using CLM are self-contained and accessibly by only you.
There will be no transcription quality degradation resulting from our service not storing your training data. Once the training data is used to actually produce a custom language model, the model itself becomes available for repeated use at your discretion. The original training set you uploaded is expunged from our systems. The only drawback is if you require technical support. Because we do not retain your original training data, we would not have convenient access to those assets or related intermediate artifacts, should you require support team to investigate potential service issues. Support would still be available, but not as expedient because we may need to ask for additional information from you.
Since training data is not stored, the same data set and any additional data will have to be uploaded again to train new models. When there is an update to the base model provided by Amazon Transcribe, you will be notified. To take advantage of the latest base model, you should submit your data to train a new model. You will then have both the original custom model that you previously generated and also the new version to use.
You can delete any customer language model that you generated, at your discretion.
Any content processed by Amazon Transcribe is encrypted and stored at rest in the AWS region where you are using Amazon Transcribe. Some portion of content processed by Amazon Transcribe may be stored in another AWS region solely in connection with the continuous improvement and development of your Amazon Transcribe customer experience and other Amazon machine-learning/artificial-intelligence technologies. If you opt out of having your content used to develop the quality of Amazon Transcribe and other Amazon machine-learning/artificial-intelligence technologies by contacting AWS Support, your content will not be stored in another AWS region. You can request deletion of voice inputs associated with your account by contacting AWS Support. Your trust, privacy, and the security of your content are our highest priority and we implement appropriate and sophisticated technical and physical controls, including encryption at rest and in transit, designed to prevent unauthorized access to, or disclosure of, your content and ensure that our use complies with our commitments to you. Please see https://aws.amazon.com/compliance/data-privacy-faq/ for more information.
Yes, subject to your compliance with the Amazon Transcribe Service Terms, including your obligation to provide any required notices and obtain any required verifiable parental consent under COPPA, you may use Amazon Transcribe in connection with websites, programs, or other applications that are directed or targeted, in whole or in part, to children under age 13.
For information about the requirements of COPPA and guidance for determining whether your website, program, or other application is subject to COPPA, please refer directly to the resources provided and maintained by the United States Federal Trade Commission. This site also contains information regarding how to determine whether a service is directed or targeted, in whole or in part, to children under age 13.
Amazon Transcribe Call Analytics
Open allAmazon Transcribe Call Analytics is an AI-powered API that provides rich call transcripts and actionable conversation insights that you can add into call applications to improve customer experience and agent productivity. It combines powerful speech-to-text and custom natural language processing (NLP) models that are trained specifically to understand customer care and outbound sales calls. As a part of AWS Contact Center Intelligence (CCI) solutions, this API is contact center agnostic and makes it easier for customers and ISVs to add call analytics capabilities into their applications.