Hi, thanks for the great work on openWakeWord!
I'd like to clarify the licensing of the two shared feature extraction models: melspectrogram.onnx and embedding_model.onnx.
The README states that:
All code in the repository is licensed under Apache 2.0.
All included pre-trained models are licensed under CC BY-NC-SA 4.0 due to training data with unknown or restrictive licensing.
However, I understand that melspectrogram.onnx and embedding_model.onnx are distinct from the wake word classification models — specifically:
melspectrogram.onnx was exported from a torchlibrosa-based PyTorch implementation (as shown in notebooks/converting_google_speech_embedding_model.ipynb), which is Apache 2.0 code.
embedding_model.onnx is a manual Keras re-implementation of Google's speech_embedding model, which Google releases under Apache 2.0. The README also explicitly notes this (under Model Architecture).
My questions are:
Are melspectrogram.onnx and embedding_model.onnx covered by the repository's Apache 2.0 license (not CC BY-NC-SA 4.0), given their origin from Apache 2.0 sources?
Is it permitted to use these two models in a commercial product, as long as we train our own wake word classification model on top of them (i.e., we do not use the pre-trained alexa, hey_jarvis, etc. classification models)?
Any clarification would be greatly appreciated. Thank you!
Hi, thanks for the great work on openWakeWord!
I'd like to clarify the licensing of the two shared feature extraction models: melspectrogram.onnx and embedding_model.onnx.
The README states that:
All code in the repository is licensed under Apache 2.0.
All included pre-trained models are licensed under CC BY-NC-SA 4.0 due to training data with unknown or restrictive licensing.
However, I understand that melspectrogram.onnx and embedding_model.onnx are distinct from the wake word classification models — specifically:
melspectrogram.onnx was exported from a torchlibrosa-based PyTorch implementation (as shown in notebooks/converting_google_speech_embedding_model.ipynb), which is Apache 2.0 code.
embedding_model.onnx is a manual Keras re-implementation of Google's speech_embedding model, which Google releases under Apache 2.0. The README also explicitly notes this (under Model Architecture).
My questions are:
Are melspectrogram.onnx and embedding_model.onnx covered by the repository's Apache 2.0 license (not CC BY-NC-SA 4.0), given their origin from Apache 2.0 sources?
Is it permitted to use these two models in a commercial product, as long as we train our own wake word classification model on top of them (i.e., we do not use the pre-trained alexa, hey_jarvis, etc. classification models)?
Any clarification would be greatly appreciated. Thank you!