This project is a Python scaffold for a live voice changer that uses RVC-style .pth model files and .index feature-index files.
Important: RVC
.pth+.indexmodels are voice conversion models. They convert one audio signal into another voice. They do not accept raw text directly. The practical low-latency flow is:
microphone audio -> optional speech-to-text captions/logging -> RVC voice conversion -> virtual audio device -> Discord/game
If you truly want microphone -> speech-to-text -> generated voice, that is a speech-to-text plus text-to-speech app, and the .pth/.index RVC files are not the part that reads the text. This scaffold keeps speech-to-text optional and routes the original microphone audio through an RVC backend.
- A desktop UI entry point:
ai-voicechanger-gui - A command-line app entry point:
ai-voicechanger - Device listing for microphone/output selection
- A stream pipeline that captures microphone audio and writes processed audio to an output device
- Optional speech-to-text worker for captions/debugging
- Validation for
.pthand.indexfiles - A pluggable RVC converter interface with:
passthroughmode for testing audio routingexternal-rvcmode that calls a user-provided command for conversion
Install a virtual audio cable and set it as the app output, then choose that virtual cable as the microphone in Discord or your game:
- Windows: VB-CABLE, VoiceMeeter, SteelSeries Sonar, or similar
- macOS: BlackHole or Loopback
- Linux: PulseAudio/PipeWire virtual sink/source
python -m venv .venv
source .venv/bin/activate
pip install -e .[stt,dev]If you only want audio routing without transcription:
pip install -e .Place your model files somewhere on disk, for example:
models/my_voice.pth
models/added_IVF.index
Large model files should not be committed to git.
ai-voicechanger-guiUse the UI to:
- Pick your microphone input.
- Pick the output device that Discord or your game should hear, usually a virtual audio cable.
- Start with
passthroughto confirm routing works. - Switch to
external-rvcand browse for your.pthand.indexfiles. - Enter the RVC command template that converts
{input}into{output}.
The UI also shows a checklist of what you need: microphone, virtual cable, RVC model/index files, optional STT model, and headphones.
ai-voicechanger --list-devicesThis verifies microphone capture and virtual audio output before adding RVC inference:
ai-voicechanger \
--input-device "Your Microphone" \
--output-device "CABLE Input" \
--backend passthroughai-voicechanger \
--input-device "Your Microphone" \
--output-device "CABLE Input" \
--backend external-rvc \
--pth models/my_voice.pth \
--index models/added_IVF.index \
--rvc-command "python path/to/rvc_infer_cli.py --input {input} --output {output} --model {pth} --index {index}"external-rvc writes each audio chunk to a temporary WAV file, calls your command, and reads the converted WAV back. This is simple and backend-agnostic, but it is not the lowest-latency architecture. For production realtime use, replace ExternalRvcConverter with a direct in-process RVC inference implementation.
ai-voicechanger \
--input-device "Your Microphone" \
--output-device "CABLE Input" \
--backend passthrough \
--stt-model small.enSpeech-to-text is for captions/logging in this app. It is intentionally not fed into the RVC .pth/.index files because those models expect audio, not text.
For Discord or games, latency matters. Start with:
- Wired headphones to avoid feedback
- A virtual cable as the app output
--block-ms 80or lower if your machine can keep up- GPU-backed RVC inference if you add a direct backend
The external command backend is best for proving the integration. A direct backend is recommended for real-time usage.