Vertex AI Advanced icon

Vertex AI Advanced

Interact with Google Vertex AI models — multimodal text, image, audio, and video

Audio → Analyze Audio

AI-generated

Summary

Analyze audio input via the specified Google Vertex AI model by providing audio URLs or binary audio files along with a prompt question. Return structured AI-generated analysis results, optionally simplified for easier consumption.

Inputs

  • Project ID (required) — The Google Cloud Platform project ID to use for the request.
  • Model (required) — The identifier of the Vertex AI audio model to use for analysis, selectable from a list or entered manually.
  • Text Input (required) — A prompt question or instruction to apply to the audio, e.g., 'What's in this audio?'.
  • Input Type (required) — Specifies whether to provide audio input as URLs or as binary files.
  • URL(s) — Comma-separated URLs pointing to the audio files to analyze (used if Input Type is 'Audio URL(s)').
  • Input Data Field Name(s) — Comma-separated binary property field names containing the audio data to analyze (used if Input Type is 'Binary File(s)').
  • Simplify Output — Whether to return a simplified response containing the top-level candidates only.
  • Options - Length of Description (Max Tokens) — Maximum token length to limit the detail level of the audio description output; lower values produce shorter, less detailed results.

Output shape

a list of JSON objects representing AI-generated analysis candidates, each corresponding to an input audio, or a single full response object if simplification is disabled.

The output is simplified by default, returning each candidate as a separate JSON item paired with the input. If 'Simplify Output' is false, the raw response containing all candidates is returned as a single JSON object. Multiple audio inputs can be specified either by URLs or multiple binary fields, and each results in corresponding output data.

Links

Discussion