Repository navigation
0.4.0: speaking-style instructions for Gemini and OpenAI, carried by avatar.json - #4
Merged
Merged
Conversation
…avatar.json - Gemini takes `instructions` and `emotionInstructions`. With either set, the prompt is "### DIRECTOR'S NOTES" (the style) then "#### TRANSCRIPT" (the text), MulmoCast's form, which Gemini follows without reading it aloud; without them only the text is sent, as before. - OpenAI also takes `instructions`, a character placed before the emotion's style. - The character + emotion combining and the default emotion styles live in src/tts/style.ts, shared by both providers. - avatar.json `voice.<provider>` accepts `options`, so a character's speaking style travels with the avatar; the CLI passes them to createTts(). Its shape is shared with createTts()'s config. - Gemini requests answered 503/429 are retried after 1, 2 and 4 s. Closes #3 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e the voice does - Gemini text sent alone fails on gemini-2.5-flash-preview-tts when it is a question (HTTP 400 "Model tried to generate text", 3 of 3; MulmoCast's agent too). The prompt is now always director's notes (the emotion's style, after the character if any) then the transcript. - trimToVoice extends a phrase's last character over voice that runs on from it without a pause, up to the next spoken character. A long, sad "の?" sounded ~0.5 s past its aligned end, so createTts() rejected it as speech not in the text, and the mouth closed early (also in compileTimeline, which MulmoCast uses). Live check: 16 cases on gemini-3.8-flash-tts and gemini-2.5-flash-preview-tts; all pass on 3.8 with no re-request. The one failure is 2.5 garbling a line in a child's voice, which the check rightly rejects. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3.
What changes
gemini-2.5-flash-preview-ttsfails with HTTP 400 "Model tried to generate text" whenever the line is a question (3 of 3; MulmoCast'sttsGeminiAgentfails the same way without an instruction). The mismatch check (re-request, then error) stays as the safety net.src/tts/style.ts, used by both OpenAI and Gemini. OpenAI now takesinstructionstoo; it's ignored ontts-1*models, which don't accept instructions.avatar.jsonvoice.<provider>.options. A character's speaking style can live in the avatar file, and the CLI passes it tocreateTts(). Themodel/voice/optionsshape is defined once insrc/tts/types.tsand shared by thecreateTts()config andavatar.json.trimToVoiceextends a held final sound. When the voice runs on without a pause after a phrase's last character, that character is extended over it, up to the next spoken character. A long, sad 「の?」 kept sounding about 0.5 s past its aligned end, socreateTts()rejected it as "speech that is not in the text," and the mouth closed early. The early mouth close also happened incompileTimeline, which MulmoCast uses.gemini-3.8-flash-tts.Live test (2026-10-07)
8 cases × 2 models (boy = Puck, マルモ = Autonoe, Japanese and English, neutral/happy/sad/surprised, with and without instructions), every clip transcribed with gpt-4o-transcribe.
The one 2.5 failure: the model garbles the end of the line in a child's voice (transcribed as 「何をすれば」「何をするまん」), and the mismatch check correctly rejects it. Before these fixes, 2.5 returned 400 on the question line, and the held-vowel line failed on both models.
Compatibility
.tts-cache(made from the text alone) is still reused; delete the cache to regenerate it.avatar.jsonvoice entries with an emptyvoice/modelstring are now rejected, the same as increateTts().Checks
npm run format,lint,typecheckandtest(71) all pass, andscripts/smoke.sh avatars/anipassed.🤖 Generated with Claude Code
work in avatarscript