Repository navigation
feat(media-backend/smt): add Qwen3-TTS support - #27
Merged
alex-spacemit merged 3 commits intoJul 24, 2026
Merged
alex-spacemit merged 3 commits into
alex-spacemit merged 3 commits into
Conversation
muggle-stack
force-pushed
the
agent/qwen3-tts-mtmd-backend
branch
from
July 22, 2026 05:49
d846791 to
8449793
Compare
muggle-stack
force-pushed
the
agent/qwen3-tts-mtmd-backend
branch
from
July 22, 2026 07:33
8449793 to
e6ed417
Compare
There was a problem hiding this comment.
Pull request overview
Adds Qwen3-TTS support to the SMT media backend and wires it into llama-server so it can serve OpenAI-compatible POST /v1/audio/speech requests returning WAV audio.
Changes:
- Introduces a Qwen3-TTS runtime (text embedding + talker + code predictor + codec decode) and an SMT-side wrapper API returning WAV + synthesis stats.
- Adds a TTS-focused server route set (
/health,/v1/health,/v1/audio/speech) plus request validation and WAV response construction with timing headers. - Extends low-level support for the SMT/SpaceMIT path (CMake wiring, ggml spacemit kernels/core affinity helpers) and adjusts context padding behavior via
LLAMA_CTX_PAD.
Reviewed changes
Copilot reviewed 21 out of 21 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| tools/smt-mtmd/smt/smt-tts-wrapper.h | Declares SMT TTS wrapper interface and result struct (WAV + stats). |
| tools/smt-mtmd/smt/smt-tts-wrapper.cpp | Implements backend selection and bridges to Qwen3-TTS runtime. |
| tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_talker.h | Declares talker/code-predictor engine interface and callback frame type. |
| tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_talker.cpp | Implements talker + code predictor execution and frame generation. |
| tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_runtime.h | Declares Qwen3-TTS runtime and synthesis result/stats types. |
| tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_runtime.cpp | Implements full Qwen3-TTS pipeline, segmentation, codec decoding, and WAV creation. |
| tools/smt-mtmd/smt/qwen3-tts/qwen3_tts_gguf.h | Adds a small GGUF mmap helper for reading auxiliary tensors efficiently. |
| tools/smt-mtmd/media-worker.h | Improves error classification by preserving std::invalid_argument as client-style rejection. |
| tools/smt-mtmd/CMakeLists.txt | Wires new SMT/Qwen3-TTS sources and required includes/libs; adds SpaceMIT-specific compile handling. |
| tools/server/server.cpp | Adds a TTS server mode and conditionally routes startup into it when SMT TTS config matches. |
| tools/server/server-tts.h | Declares TTS request validation and WAV response helpers. |
| tools/server/server-tts.cpp | Implements request validation (size, format, speed, hotwords) and WAV response headers/body. |
| tools/server/server-media.h | Adds a TTS result type and TTS init/synthesize APIs under LLAMA_SERVER_SMT_MTMD. |
| tools/server/server-media.cpp | Implements SMT TTS init and synthesis execution via media_worker. |
| tools/server/CMakeLists.txt | Adds server-tts sources when SMT/MTMD support is enabled. |
| src/models/qwen3.cpp | Avoids building lm_head/logits when in embeddings mode. |
| src/llama-context.cpp | Makes context padding configurable via LLAMA_CTX_PAD (power-of-two) for n_ctx and n_ctx_seq. |
| ggml/src/ggml-cpu/spacemit/ime2_kernels.cpp | Adjusts SpaceMIT RVV kernel assembly loop structure/initialization. |
| ggml/src/ggml-cpu/spacemit/ime.h | Exposes SpaceMIT preferred-core selection/get APIs. |
| ggml/src/ggml-cpu/spacemit/ime.cpp | Implements preferred-core selection/get and improves CPU mask safety checks. |
| common/arg.cpp | Adds server example to --tts-speaker-file help examples. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+262
to
+266
| const char * ctx_pad_env = getenv("LLAMA_CTX_PAD"); | ||
| const int requested_ctx_pad = ctx_pad_env ? atoi(ctx_pad_env) : 0; | ||
| const uint32_t ctx_pad = requested_ctx_pad > 0 && (requested_ctx_pad & (requested_ctx_pad - 1)) == 0 ? | ||
| static_cast<uint32_t>(requested_ctx_pad) : 256; | ||
| cparams.n_ctx = GGML_PAD(cparams.n_ctx, ctx_pad); |
Comment on lines
+185
to
+190
| #if defined(LLAMA_SERVER_SMT_MTMD) | ||
| if (server_media_context::tts_config_matches(params)) { | ||
| common_params_print_info(params, false); | ||
| return run_tts_server(params); | ||
| } | ||
| #endif |
- allow Qwen3 embedding-only contexts to skip the language model head - support compact context padding for the TTS talker and code predictor - reject malformed context padding overrides deterministically Assisted-by: OpenAI Codex
muggle-stack
force-pushed
the
agent/qwen3-tts-mtmd-backend
branch
from
July 22, 2026 12:05
e6ed417 to
94799dd
Compare
co-seven
reviewed
Jul 23, 2026
Collaborator
There was a problem hiding this comment.
这里的循环展开和指令重排让Qwen3-30B-A3B-Instruct-2507-Q4_0.gguf模型prefill tps下降了2个点,这个比较明显,因为稀疏模型很容易prefill阶段出发m1的kernel,同时也让dense模型的decode性能受到微小的负面影响,去掉这部分之后模型效果均能好一些。这个优化可以再考虑考虑,或者干脆去掉。
muggle-stack
force-pushed
the
agent/qwen3-tts-mtmd-backend
branch
from
July 23, 2026 07:35
94799dd to
32d96c8
Compare
co-seven
reviewed
Jul 23, 2026
Collaborator
There was a problem hiding this comment.
直接干脆吧#include "server-media.h"放进9行下面吧,这样server-media.cpp也不需要包很多LLAMA_SERVER_SMT_MTMD了
- run text embedding, talker, code predictor, and codec decoding in process - expose WAV synthesis through the OpenAI-compatible audio speech endpoint - support request hotwords and startup-selected speaker reference binaries Assisted-by: OpenAI Codex
muggle-stack
force-pushed
the
agent/qwen3-tts-mtmd-backend
branch
from
July 23, 2026 08:39
32d96c8 to
24e8854
Compare
- distribute Q4_0 and Q8_0 single-row output tiles across worker threads Assisted-by: OpenAI Codex
muggle-stack
force-pushed
the
agent/qwen3-tts-mtmd-backend
branch
from
July 24, 2026 02:40
24e8854 to
c913ab8
Compare
Collaborator
|
审核通过 |
alex-spacemit
merged commit Jul 24, 2026
6ad6d85
into
spacemit-com:mtmd-backend
14 of 19 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Integrates Qwen3-TTS into the SMT media backend, enabling
llama-serverto serve text-to-speech via the OpenAI-compatible/v1/audio/speechendpoint and return WAV audio. This PR implements the full inference pipeline, supports Chinese/English and mixed-language scenarios, and achievesRTF ≈ 0.89–0.91on the K3 long-utterance test set.Motivation
The SMT backend previously lacked end-to-end support for the Qwen3-TTS model family. This integration lets users get high-quality CN/EN TTS inside a single
llama-serverprocess with no extra subprocess or external service, while staying consistent with the existing OpenAI-compatible API surface to minimize adoption friction.Changes
Interface
llama-server --media-backend smt --smt-config-dir <path>/v1/audio/speechendpoint; response body is WAV audioArchitecture
Implemented as single-process, multi-threaded with no runner subprocess:
llama-serverruntimeInference Pipeline
Full four-stage Qwen3-TTS pipeline:
Feature Coverage
ref-bin)Scope
Included in this PR: Qwen3-TTS integration in the SMT backend, the inference pipeline, and
/v1/audio/speechWAV output.Explicitly out of scope (deferred to follow-up PRs):
ref2bintoolPerformance
On the K3 long-utterance test set, RTF ≈ 0.89–0.91 (generation time is slightly below audio duration, i.e. real-time threshold is met). Further headroom remains (scheduling, kernels, streaming time-to-first-byte, etc.) and is not in scope here.
Validation
k3-1HTTP 200; WAV output plays correctlyUsage
Change Stats
+2103 / -18Follow-ups
ref2bintool/v1/audio/speech)