Vertex AI Advanced icon

Vertex AI Advanced

Interact with Google Vertex AI models — multimodal text, image, audio, and video

Audio → Transcribe a Recording

AI-generated

Summary

Transcribe audio recordings into text by providing either URLs to audio files or binary audio data input.

Inputs

  • Project ID (required) — Google Cloud Platform project ID to use for the request.
  • Model (required) — The Vertex AI model ID to use for transcription, selectable from a filtered list or by entering a model ID directly.
  • Input Type (required) — Choose between providing Audio URL(s) or Binary File(s) as the audio input.
  • URL(s) — Comma-separated list of public URLs pointing to audio files to transcribe, required if Input Type is Audio URL(s).
  • Input Data Field Name(s) — Comma-separated binary property names containing audio files to transcribe, required if Input Type is Binary File(s).
  • Simplify Output — Whether to return a simplified transcription response or the full raw API response, defaults to simplified.
  • Options — Optional transcription parameters, including Start Time and End Time to specify audio segments in MM:SS or HH:MM:SS format.

Output shape

a list of transcription candidates when simplified; otherwise a single full response object

If "Simplify Output" is enabled, each candidate transcription is returned as a separate item. Otherwise, the complete API response is returned as one item. Multiple audio inputs are processed and returned in the same manner.

Examples

Example 1: Transcribe audio files by URL

Set Resource to Audio, Operation to Transcribe, enter Project ID, select Model, choose Input Type 'Audio URL(s)', and provide one or more comma-separated URLs of audio recordings.

Example 2: Transcribe audio binaries

Set Resource to Audio, Operation to Transcribe, enter Project ID, select Model, choose Input Type 'Binary File(s)', and specify the binary property name(s) from previous node data.

Links

Discussion