AssemblyAI - Transcribe and wait action ​

The Transcribe and wait action submits an audio file to AssemblyAI for transcription, then polls every 60 seconds until the transcript's status reaches completed, and returns the finished transcript.

SYNCHRONOUS SUBMIT, BOUNDED WAIT

This action polls for a maximum of 60 minutes and this setting is configurable in the Max poll minutes field, which is 60 by default. Use the Create record action with a Webhook URL and the Completed transcript trigger for longer running jobs.

Input ​

The following are a representative set of input fields. Refer to the AssemblyAI Transcript audio documentation for more information on available input fields.

Input fieldDescription
Audio URLEnter a publicly reachable audio URL, or the Upload URL output from the Upload audio file action.
Custom spellingOptional. Map custom spellings, each with one or more From terms and a single To replacement.
Redact static entitiesOptional. Enter static entity redaction labels and terms. Converted to a map in the API payload.
Language codeOptional. Select the audio's language. Don't use this field together with Language detection.
Language detectionOptional. Select Yes or No to automatically detect the audio's language. Don't use this field together with Language code.
Speaker labelsOptional. Select Yes or No to enable speaker diarization labels.
Auto highlightsOptional. Select Yes or No to enable key phrase detection. AssemblyAI supports similar toggles for entity detection, sentiment analysis, IAB category detection, and content safety.
Redact PIIOptional. Select Yes or No to redact personally identifiable information from the transcript.
Redact PII policiesSelect the PII categories to redact. Required when Redact PII is Yes, for newer accounts.
Max poll minutesOptional. Enter how many minutes to keep polling before the action times out. Default and maximum 60.

Output ​

The following output fields are common to every response. Refer to the AssemblyAI Transcript documentation for more information on available output fields.

Output fieldDescription
IDThe record's unique identifier.
StatusThe record's status. Always completed on this action's response.
TextThe transcribed text.
WordsA list of word-level objects.
Text (Words)The word's transcribed text.
Start (Words)The word's start time, in milliseconds.
End (Words)The word's end time, in milliseconds.
Confidence (Words)The model's confidence score for this word, between 0 and 1.
Speaker (Words)The identified speaker label for this word. The Speaker field is only available when Speaker labels is Yes.
Channel (Words)The audio channel this word came from, for multichannel audio recordings.

Last updated: