Speak
AI-generatedSummary
Send a speak action to an active sipgate AI Flow session to have the AI assistant speak specified content using text or SSML with configurable speech synthesis and barge-in options.
Inputs
- Session ID (required) — The unique session ID from the AI Flow event identifying the active call session.
- Content Type (required) — Choose between plain text or SSML markup for speech content.
- Text (required) — The plain text for the AI assistant to speak (required if Content Type is Text).
- SSML (required) — The SSML markup controlling speech synthesis (required if Content Type is SSML).
- User Input Timeout (Seconds) — Timeout in seconds to wait for user input after speech ends; 0 disables timeout.
- TTS Provider (required) — The text-to-speech provider to use: Azure, ElevenLabs, or default configured provider.
- Language — Language code for Azure TTS provider (e.g., en-US); only applicable if Azure is selected.
- Voice — Voice name or ID for Azure or ElevenLabs TTS provider.
- Barge-In Options — Settings controlling when and how callers can interrupt the speech (barge-in), including strategy, minimum characters, and delay (ms).
Output shape
a list of action objects paired to input items, each describing the speak action sent to the AI Flow session
The output includes the action JSON sent with fields like session_id, type='speak', text or ssml content, optional user_input_timeout_seconds, tts configuration, and barge_in settings. No additional data is returned; errors are included per item if continueOnFail is enabled.