> ## Documentation Index
> Fetch the complete documentation index at: https://dify-6c0370d8-release-1-17-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Convert Text to Audio

> **Available for**: Chatflow, Workflow, New Agent, Chatbot, Agent, Text Generator apps.

Converts text to speech audio. Pass `text` to synthesize arbitrary text, or `message_id` to voice an existing message's answer.



## OpenAPI

````yaml /en/api-reference/openapi_service.json post /text-to-audio
openapi: 3.0.1
info:
  title: Dify Service API
  description: >-
    REST API for Dify applications and knowledge bases. Application endpoints
    authenticate with an app API key; knowledge endpoints authenticate with a
    dataset API key.
  version: 1.0.0
servers:
  - url: https://{api_base_url}
    description: >-
      Base URL of the Dify Service API. For self-hosted deployments, replace it
      with your own API base URL.
    variables:
      api_base_url:
        default: api.dify.ai/v1
        description: Host and path of the API base URL, without the `https://` prefix.
security:
  - ApiKeyAuth: []
tags:
  - name: Chat Messages
    description: Operations related to chat messages and interactions.
  - name: Files
    description: File upload and preview operations.
  - name: End Users
    description: Operations related to end user information.
  - name: Feedback
    description: User feedback operations.
  - name: Conversations
    description: Operations related to managing conversations.
  - name: Audio
    description: Text-to-Speech and Speech-to-Text operations.
  - name: Applications
    description: Operations to retrieve application settings and information.
  - name: Annotations
    description: Operations related to managing annotations for direct replies.
  - name: Human Input
    description: Endpoints for resuming paused workflows that require human input.
  - name: Workflow Runs
    description: Operations for executing and managing workflows.
  - name: Completion Messages
    description: Operations related to text generation and completion.
  - name: Knowledge Bases
    description: >-
      Operations for managing knowledge bases, including creation,
      configuration, and retrieval.
  - name: Documents
    description: >-
      Operations for creating, updating, and managing documents within a
      knowledge base.
  - name: Chunks
    description: Operations for managing document chunks and child chunks.
  - name: Metadata
    description: >-
      Operations for managing knowledge base metadata fields and document
      metadata values.
  - name: Tags
    description: Operations for managing knowledge base tags and tag bindings.
  - name: Models
    description: Operations for retrieving available models.
  - name: Knowledge Pipeline
    description: >-
      Operations for managing and running knowledge pipelines, including
      datasource plugins and pipeline execution.
paths:
  /text-to-audio:
    post:
      tags:
        - Audio
      summary: Convert Text to Audio
      description: >-
        **Available for**: Chatflow, Workflow, New Agent, Chatbot, Agent, Text
        Generator apps.


        Converts text to speech audio. Pass `text` to synthesize arbitrary text,
        or `message_id` to voice an existing message's answer.
      operationId: textToAudioChat
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/TextToAudioRequest'
            examples:
              textToAudioExample:
                summary: Request Example
                value:
                  text: Hello, welcome to our service.
                  user: abc-123
                  voice: alloy
                  streaming: false
      responses:
        '200':
          description: >-
            Returns the generated audio. The `Content-Type` header reflects the
            provider's audio container, verified from the response bytes when
            recognizable.


            The body can be AAC, FLAC, MP4, MP3, Ogg, WAV, or WebM. Output that
            cannot be recognized is labeled with the provider's declared type,
            or `audio/mpeg` when none is declared.


            Streamed provider output is delivered with chunked transfer
            encoding; the request `streaming` field does not control this.
          content:
            audio/aac:
              schema:
                type: string
                format: binary
            audio/flac:
              schema:
                type: string
                format: binary
            audio/mp4:
              schema:
                type: string
                format: binary
            audio/mpeg:
              schema:
                type: string
                format: binary
            audio/ogg:
              schema:
                type: string
                format: binary
            audio/wav:
              schema:
                type: string
                format: binary
            audio/webm:
              schema:
                type: string
                format: binary
        '400':
          description: >-
            - `app_unavailable` : The app is unavailable or misconfigured.

            - `invalid_param` : Text-to-speech is not enabled, `text` is
            missing, or no voice is available.

            - `provider_not_initialize` : No valid model provider credentials
            are configured.

            - `provider_quota_exceeded` : The model provider quota is exhausted.

            - `model_currently_not_support` : The current model does not support
            this operation.

            - `completion_request_error` : The text-to-speech request failed.
          content:
            application/json:
              examples:
                app_unavailable:
                  summary: app_unavailable
                  value:
                    status: 400
                    code: app_unavailable
                    message: App unavailable, please check your app configurations.
                invalid_param:
                  summary: invalid_param
                  value:
                    status: 400
                    code: invalid_param
                    message: TTS is not enabled
                provider_not_initialize:
                  summary: provider_not_initialize
                  value:
                    status: 400
                    code: provider_not_initialize
                    message: >-
                      No valid model provider credentials found. Please go to
                      Settings -> Model Provider to complete your provider
                      credentials.
                provider_quota_exceeded:
                  summary: provider_quota_exceeded
                  value:
                    status: 400
                    code: provider_quota_exceeded
                    message: >-
                      Your quota for Dify Hosted OpenAI has been exhausted.
                      Please go to Settings -> Model Provider to complete your
                      own provider credentials.
                model_currently_not_support:
                  summary: model_currently_not_support
                  value:
                    status: 400
                    code: model_currently_not_support
                    message: >-
                      Dify Hosted OpenAI trial currently not support the GPT-4
                      model.
                completion_request_error:
                  summary: completion_request_error
                  value:
                    status: 400
                    code: completion_request_error
                    message: Completion request failed.
        '500':
          description: '`internal_server_error` : Internal server error.'
          content:
            application/json:
              examples:
                internal_server_error:
                  summary: internal_server_error
                  value:
                    status: 500
                    code: internal_server_error
                    message: >-
                      The server encountered an internal error and was unable to
                      complete your request. Either the server is overloaded or
                      there is an error in the application.
components:
  schemas:
    TextToAudioRequest:
      type: object
      description: >-
        Request body for text-to-audio conversion. Provide either `message_id`
        or `text`.
      properties:
        message_id:
          type: string
          format: uuid
          description: >-
            ID of the message whose answer to voice. Takes priority over `text`
            when both are provided. Get message IDs from [List Conversation
            Messages](/en/api-reference/conversations/list-conversation-messages).
        text:
          type: string
          description: Text to synthesize into speech.
        user:
          type: string
          description: >-
            End-user identifier, defined by your app and unique within it. See
            [End User Identity](/en/api-reference/guides/end-user-identity).
        voice:
          type: string
          description: >-
            Voice to use for text-to-speech. Available voices depend on the TTS
            provider configured for this app. Use the `voice` value from [Get
            App Parameters](/en/api-reference/applications/get-app-parameters) →
            `text_to_speech.voice` for the default.
        streaming:
          type: boolean
          description: >-
            Accepted for backward compatibility but has no effect. Whether the
            audio is streamed is determined by the configured TTS provider's
            output, not by this field.
  securitySchemes:
    ApiKeyAuth:
      type: http
      scheme: bearer
      bearerFormat: API_KEY
      description: >-
        Every request authenticates with an API key: `Authorization: Bearer
        {API_KEY}`. App endpoints take an app API key; knowledge endpoints take
        a knowledge base API key ([Get
        Started](/en/api-reference/guides/get-started)).


        Keep keys server-side; never embed them in client code. Requests with a
        missing or invalid key fail with HTTP `401` (`unauthorized`).

````