You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 9dbd3da
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: .speakeasy/in.openapi.yaml
+80Lines changed: 80 additions & 0 deletions
Original file line number
Diff line number
Diff line change
@@ -20303,6 +20303,7 @@ components:
20303
20303
- 'top_p'
20304
20304
- 'max_tokens'
20305
20305
supports_implicit_caching: true
20306
+
supports_voice_cloning: false
20306
20307
tag: 'openai'
20307
20308
throughput_last_30m:
20308
20309
p50: 45.2
@@ -20404,6 +20405,10 @@ components:
20404
20405
type: 'array'
20405
20406
supports_implicit_caching:
20406
20407
type: 'boolean'
20408
+
supports_voice_cloning:
20409
+
default: false
20410
+
description: 'Whether this TTS endpoint accepts inline reference audio (`input_references`) for stateless voice cloning. Requests carrying reference audio are only routed to endpoints where this is true.'
description: 'Base64-encoded reference audio (optionally a data URI). Supported audio formats are provider-specific. Limited to 20 MiB of base64 (15 MiB of decoded audio).'
description: 'Audio format of the reference audio (e.g., wav, mp3). Optional; most providers detect the format from the audio bytes.'
21882
+
example: 'wav'
21883
+
type: 'string'
21884
+
required:
21885
+
- 'data'
21886
+
type: 'object'
21887
+
SpeechInputReferenceText:
21888
+
description: 'Transcript of the accompanying reference audio'
21889
+
example:
21890
+
text: 'I used to rule the world.'
21891
+
type: 'text'
21892
+
properties:
21893
+
text:
21894
+
description: 'Transcript of the accompanying reference audio.'
21895
+
example: 'I used to rule the world.'
21896
+
maxLength: 10000
21897
+
type: 'string'
21898
+
type:
21899
+
enum:
21900
+
- 'text'
21901
+
type: 'string'
21902
+
required:
21903
+
- 'type'
21904
+
- 'text'
21905
+
type: 'object'
21839
21906
SpeechRequest:
21840
21907
description: 'Text-to-speech request input'
21841
21908
example:
@@ -21849,6 +21916,17 @@ components:
21849
21916
description: 'Text to synthesize'
21850
21917
example: 'Hello world'
21851
21918
type: 'string'
21919
+
input_references:
21920
+
description: 'Reference content for stateless voice cloning: one `input_audio` part carrying the voice sample, optionally accompanied by one `text` part with its transcript. Only routed to endpoints that support voice cloning.'
0 commit comments