Skip to content

feat(media-backend/smt): add Qwen3-TTS support - #27

Merged
alex-spacemit merged 3 commits into
spacemit-com:mtmd-backendfrom
muggle-stack:agent/qwen3-tts-mtmd-backend
Jul 24, 2026
Merged

alex-spacemit merged 3 commits into
spacemit-com:mtmd-backendfrom
muggle-stack:agent/qwen3-tts-mtmd-backend

Conversation

@muggle-stack

Copy link
Copy Markdown

Summary

Integrates Qwen3-TTS into the SMT media backend, enabling llama-server to serve text-to-speech via the OpenAI-compatible /v1/audio/speech endpoint and return WAV audio. This PR implements the full inference pipeline, supports Chinese/English and mixed-language scenarios, and achieves RTF ≈ 0.89–0.91 on the K3 long-utterance test set.

Motivation

The SMT backend previously lacked end-to-end support for the Qwen3-TTS model family. This integration lets users get high-quality CN/EN TTS inside a single llama-server process with no extra subprocess or external service, while staying consistent with the existing OpenAI-compatible API surface to minimize adoption friction.

Changes

Interface

  • New launch flags: llama-server --media-backend smt --smt-config-dir <path>
  • Reuses the /v1/audio/speech endpoint; response body is WAV audio
  • Loads Talker / Code Predictor / codec models and config from the config dir

Architecture

Implemented as single-process, multi-threaded with no runner subprocess:

  • Inference and the HTTP server share a single address space, avoiding IPC and serialization overhead
  • Thread-level concurrent scheduling across the Talker, Code Predictor, and codec decode stages
  • Simplifies deployment and lifecycle management, and embeds cleanly into the existing llama-server runtime

Inference Pipeline

Full four-stage Qwen3-TTS pipeline:

  1. Text embedding — encodes text (including hotwords and mixed-language input)
  2. Talker — main generative model, produces discrete speech units
  3. Code Predictor — predicts codec-level tokens
  4. Codec decode — reconstructs PCM audio frames

Feature Coverage

Scenario Status
Chinese only ✅
English only ✅
CN/EN mixed ✅
Short / long utterances ✅
Hotwords ✅
Speaker reference (ref-bin) ✅

Scope

Included in this PR: Qwen3-TTS integration in the SMT backend, the inference pipeline, and /v1/audio/speech WAV output.

Explicitly out of scope (deferred to follow-up PRs):

  • Public TTS CLI
  • ref2bin tool
  • Benchmarks
  • Experimental feature flags / switches

Performance

On the K3 long-utterance test set, RTF ≈ 0.89–0.91 (generation time is slightly below audio duration, i.e. real-time threshold is met). Further headroom remains (scheduling, kernels, streaming time-to-first-byte, etc.) and is not in scope here.

Validation

  • ✅ Build: compiles cleanly on k3-1
  • ✅ Functional: all six voice test cases (CN / EN / mixed / short & long / hotword / ref-bin) return HTTP 200; WAV output plays correctly
  • ✅ Stability: no premature EOS, no frame-limit triggers, no NaN errors

Usage

llama-server \
  --media-backend smt \
  --smt-config-dir /path/to/qwen3-tts-config
curl -X POST http://localhost:8080/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-tts",
    "input": "你好,这是一段 Qwen3-TTS 测试。Hello, mixed content 测试。",
    "voice": "default"
  }' \
  --output speech.wav

Change Stats

  • Files changed: 21
  • Diff: +2103 / -18

Follow-ups

  • Public TTS CLI and the ref2bin tool
  • Systematic benchmarks and performance regression tracking
  • Streaming output (streaming /v1/audio/speech)
  • Experimental switches and multi-speaker management

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds Qwen3-TTS support to the SMT media backend and wires it into llama-server so it can serve OpenAI-compatible POST /v1/audio/speech requests returning WAV audio.

Changes:

  • Introduces a Qwen3-TTS runtime (text embedding + talker + code predictor + codec decode) and an SMT-side wrapper API returning WAV + synthesis stats.
  • Adds a TTS-focused server route set (/health, /v1/health, /v1/audio/speech) plus request validation and WAV response construction with timing headers.
  • Extends low-level support for the SMT/SpaceMIT path (CMake wiring, ggml spacemit kernels/core affinity helpers) and adjusts context padding behavior via LLAMA_CTX_PAD.

Reviewed changes

Copilot reviewed 21 out of 21 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
tools/smt-mtmd/smt/smt-tts-wrapper.h Declares SMT TTS wrapper interface and result struct (WAV + stats).
tools/smt-mtmd/smt/smt-tts-wrapper.cpp Implements backend selection and bridges to Qwen3-TTS runtime.
tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_talker.h Declares talker/code-predictor engine interface and callback frame type.
tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_talker.cpp Implements talker + code predictor execution and frame generation.
tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_runtime.h Declares Qwen3-TTS runtime and synthesis result/stats types.
tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_runtime.cpp Implements full Qwen3-TTS pipeline, segmentation, codec decoding, and WAV creation.
tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_gguf.h Adds a small GGUF mmap helper for reading auxiliary tensors efficiently.
tools/smt-mtmd/media-worker.h Improves error classification by preserving std::invalid_argument as client-style rejection.
tools/smt-mtmd/CMakeLists.txt Wires new SMT/Qwen3-TTS sources and required includes/libs; adds SpaceMIT-specific compile handling.
tools/server/server.cpp Adds a TTS server mode and conditionally routes startup into it when SMT TTS config matches.
tools/server/server-tts.h Declares TTS request validation and WAV response helpers.
tools/server/server-tts.cpp Implements request validation (size, format, speed, hotwords) and WAV response headers/body.
tools/server/server-media.h Adds a TTS result type and TTS init/synthesize APIs under LLAMA_SERVER_SMT_MTMD.
tools/server/server-media.cpp Implements SMT TTS init and synthesis execution via media_worker.
tools/server/CMakeLists.txt Adds server-tts sources when SMT/MTMD support is enabled.
src/models/qwen3.cpp Avoids building lm_head/logits when in embeddings mode.
src/llama-context.cpp Makes context padding configurable via LLAMA_CTX_PAD (power-of-two) for n_ctx and n_ctx_seq.
ggml/src/ggml-cpu/spacemit/ime2_kernels.cpp Adjusts SpaceMIT RVV kernel assembly loop structure/initialization.
ggml/src/ggml-cpu/spacemit/ime.h Exposes SpaceMIT preferred-core selection/get APIs.
ggml/src/ggml-cpu/spacemit/ime.cpp Implements preferred-core selection/get and improves CPU mask safety checks.
common/arg.cpp Adds server example to --tts-speaker-file help examples.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/llama-context.cpp
Comment on lines +262 to +266
const char * ctx_pad_env = getenv("LLAMA_CTX_PAD");
const int requested_ctx_pad = ctx_pad_env ? atoi(ctx_pad_env) : 0;
const uint32_t ctx_pad = requested_ctx_pad > 0 && (requested_ctx_pad & (requested_ctx_pad - 1)) == 0 ?
static_cast<uint32_t>(requested_ctx_pad) : 256;
cparams.n_ctx = GGML_PAD(cparams.n_ctx, ctx_pad);
Comment thread tools/server/server.cpp
Comment on lines +185 to +190
#if defined(LLAMA_SERVER_SMT_MTMD)
if (server_media_context::tts_config_matches(params)) {
common_params_print_info(params, false);
return run_tts_server(params);
}
#endif
- allow Qwen3 embedding-only contexts to skip the language model head
- support compact context padding for the TTS talker and code predictor
- reject malformed context padding overrides deterministically

Assisted-by: OpenAI Codex
@muggle-stack
muggle-stack force-pushed the agent/qwen3-tts-mtmd-backend branch from e6ed417 to 94799dd Compare July 22, 2026 12:05

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里的循环展开和指令重排让Qwen3-30B-A3B-Instruct-2507-Q4_0.gguf模型prefill tps下降了2个点,这个比较明显,因为稀疏模型很容易prefill阶段出发m1的kernel,同时也让dense模型的decode性能受到微小的负面影响,去掉这部分之后模型效果均能好一些。这个优化可以再考虑考虑,或者干脆去掉。

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

回归了一下,确实对tts无提升。同意去掉。

@muggle-stack
muggle-stack force-pushed the agent/qwen3-tts-mtmd-backend branch from 94799dd to 32d96c8 Compare July 23, 2026 07:35
Comment thread tools/server/server.cpp

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

直接干脆吧#include "server-media.h"放进9行下面吧,这样server-media.cpp也不需要包很多LLAMA_SERVER_SMT_MTMD了

- run text embedding, talker, code predictor, and codec decoding in process
- expose WAV synthesis through the OpenAI-compatible audio speech endpoint
- support request hotwords and startup-selected speaker reference binaries

Assisted-by: OpenAI Codex
@muggle-stack
muggle-stack force-pushed the agent/qwen3-tts-mtmd-backend branch from 32d96c8 to 24e8854 Compare July 23, 2026 08:39
- distribute Q4_0 and Q8_0 single-row output tiles across worker threads

Assisted-by: OpenAI Codex
@muggle-stack
muggle-stack force-pushed the agent/qwen3-tts-mtmd-backend branch from 24e8854 to c913ab8 Compare July 24, 2026 02:40
@co-seven

Copy link
Copy Markdown
Collaborator

审核通过

@alex-spacemit
alex-spacemit merged commit 6ad6d85 into spacemit-com:mtmd-backend Jul 24, 2026
14 of 19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants