AssemblyAI icon

AssemblyAI

Transcribe audio and video files using AssemblyAI's speech-to-text and speech understanding AI models.

LLM Gateway → Speech Understanding

AI-generated

Overview

This node operation processes speech understanding tasks on existing transcripts using AssemblyAI's LLM Gateway. It supports three main types of tasks: translation of transcript text into multiple target languages, speaker identification by name or role, and custom formatting of dates, phone numbers, and emails within the transcript. This operation is useful for enhancing transcript data with translations, identifying speakers in multi-speaker audio, or applying specific formatting rules to recognized entities.

Use Case Examples

  1. Translating a transcript into Spanish and German with formal language style.
  2. Identifying speakers in a transcript by their roles such as host and guest.
  3. Applying custom date and phone number formats to a transcript for consistent presentation.

Properties

Name Meaning
Transcript ID ID of the transcript to process, required to specify which transcript to analyze.
Task Type Type of speech understanding task to perform, with options for translation, speaker identification, or custom formatting.
Target Languages Comma-separated list of target language codes for translation tasks (e.g., "es,de,fr"). Only shown for translation task type.
Use Formal Language Whether to use formal language style in translations. Only applicable for translation task type.
Match Original Utterance Whether to return translated text in the utterances array, including a translated_texts key for each target language. Requires Speaker Labels enabled. Only for translation task type.
Speaker Type Type of speaker identification to perform, either by name or role. Requires Speaker Labels enabled. Only for speaker identification task type.
Known Speaker Values Comma-separated list of known speaker values required for 'role' speaker type (max 35 chars each). Optional for 'name' type. Requires Speaker Labels enabled. Only for speaker identification task type.
Date Format Date format pattern to apply in custom formatting tasks (e.g., "mm/dd/yyyy" or "yyyy-mm-dd"). Only for custom formatting task type.
Phone Number Format Phone number format pattern to apply in custom formatting tasks (e.g., "(xxx)xxx-xxxx"). Only for custom formatting task type.
Email Format Email format pattern to apply in custom formatting tasks (e.g., "username@domain.com"). Only for custom formatting task type.

Output

JSON

  • json - The JSON response from AssemblyAI API containing the results of the speech understanding task, such as translations, speaker identifications, or formatted transcript data.

Dependencies

  • AssemblyAI API key credential required for authentication

Troubleshooting

  • Error if Known Speaker Values are missing when Speaker Type is 'role' in speaker identification task. Ensure to provide these values as a comma-separated list.
  • API errors may include detailed messages from AssemblyAI; check the error details for specific causes.
  • Ensure the transcript has Speaker Labels enabled when using speaker identification or match original utterance features, as these depend on speaker labeling.

Links

Discussion