Skip to content

0.4.0: speaking-style instructions for Gemini and OpenAI, carried by avatar.json - #4

Merged
snakajima merged 2 commits into
mainfrom
feat/gemini-instructions
Oct 8, 2026
Merged

snakajima merged 2 commits into
mainfrom
feat/gemini-instructions

Conversation

@snakajima

@snakajima snakajima commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Closes #3.

What changes

  • Gemini prompt is always DIRECTOR'S NOTES.
    ### DIRECTOR'S NOTES
    <character instruction (instructions, if any)>
    <the line's emotion style (default or emotionInstructions)>
    
    #### TRANSCRIPT
    <text>
    
    This is MulmoCast's form. It is now used even when no options are set, because sending the text alone to gemini-2.5-flash-preview-tts fails with HTTP 400 "Model tried to generate text" whenever the line is a question (3 of 3; MulmoCast's ttsGeminiAgent fails the same way without an instruction). The mismatch check (re-request, then error) stays as the safety net.
  • Shared character + emotion logic. Combining the two and the default emotion styles now live in src/tts/style.ts, used by both OpenAI and Gemini. OpenAI now takes instructions too; it's ignored on tts-1* models, which don't accept instructions.
  • avatar.json voice.<provider>.options. A character's speaking style can live in the avatar file, and the CLI passes it to createTts(). The model / voice / options shape is defined once in src/tts/types.ts and shared by the createTts() config and avatar.json.
  • trimToVoice extends a held final sound. When the voice runs on without a pause after a phrase's last character, that character is extended over it, up to the next spoken character. A long, sad 「の?」 kept sounding about 0.5 s past its aligned end, so createTts() rejected it as "speech that is not in the text," and the mouth closed early. The early mouth close also happened in compileTimeline, which MulmoCast uses.
  • 503 / 429 retry for Gemini. Requests are retried after 1, 2 and 4 s. The default model stays gemini-3.8-flash-tts.

Live test (2026-10-07)

8 cases × 2 models (boy = Puck, マルモ = Autonoe, Japanese and English, neutral/happy/sad/surprised, with and without instructions), every clip transcribed with gpt-4o-transcribe.

gemini-3.8-flash-tts gemini-2.5-flash-preview-tts
Passed 8/8 (no re-requests) 7/8
Instructions read aloud 0 0

The one 2.5 failure: the model garbles the end of the line in a child's voice (transcribed as 「何をすれば」「何をするまん」), and the mismatch check correctly rejects it. Before these fixes, 2.5 returned 400 on the question line, and the held-vowel line failed on both models.

Compatibility

  • Gemini's prompt changes even without options (the default emotion style is now sent). The cache key doesn't include the prompt, so Gemini audio already in .tts-cache (made from the text alone) is still reused; delete the cache to regenerate it.
  • avatar.json voice entries with an empty voice / model string are now rejected, the same as in createTts().

Checks

  • npm run format, lint, typecheck and test (71) all pass, and scripts/smoke.sh avatars/ani passed.

🤖 Generated with Claude Code

work in avatarscript

snakajima and others added 2 commits October 7, 2026 21:50
…avatar.json

- Gemini takes `instructions` and `emotionInstructions`. With either set, the prompt is
  "### DIRECTOR'S NOTES" (the style) then "#### TRANSCRIPT" (the text), MulmoCast's form, which
  Gemini follows without reading it aloud; without them only the text is sent, as before.
- OpenAI also takes `instructions`, a character placed before the emotion's style.
- The character + emotion combining and the default emotion styles live in src/tts/style.ts,
  shared by both providers.
- avatar.json `voice.<provider>` accepts `options`, so a character's speaking style travels with
  the avatar; the CLI passes them to createTts(). Its shape is shared with createTts()'s config.
- Gemini requests answered 503/429 are retried after 1, 2 and 4 s.

Closes #3

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e the voice does

- Gemini text sent alone fails on gemini-2.5-flash-preview-tts when it is a question (HTTP 400
  "Model tried to generate text", 3 of 3; MulmoCast's agent too). The prompt is now always
  director's notes (the emotion's style, after the character if any) then the transcript.
- trimToVoice extends a phrase's last character over voice that runs on from it without a
  pause, up to the next spoken character. A long, sad "の?" sounded ~0.5 s past its aligned end,
  so createTts() rejected it as speech not in the text, and the mouth closed early (also in
  compileTimeline, which MulmoCast uses).

Live check: 16 cases on gemini-3.8-flash-tts and gemini-2.5-flash-preview-tts; all pass on 3.8
with no re-request. The one failure is 2.5 garbling a line in a child's voice, which the check
rightly rejects.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@snakajima
snakajima merged commit fadfc73 into main Oct 8, 2026
7 checks passed
@snakajima
snakajima deleted the feat/gemini-instructions branch October 8, 2026 15:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Gemini TTS: send speaking-style instructions (DIRECTOR'S NOTES form) and allow per-avatar voice instructions

1 participant