Skip to content

License clarification: Are melspectrogram.onnx and embedding_model.onnx under Apache 2.0 and suitable for commercial use? #338

Description

@RellyLee

Hi, thanks for the great work on openWakeWord!

I'd like to clarify the licensing of the two shared feature extraction models: melspectrogram.onnx and embedding_model.onnx.

The README states that:

All code in the repository is licensed under Apache 2.0.
All included pre-trained models are licensed under CC BY-NC-SA 4.0 due to training data with unknown or restrictive licensing.
However, I understand that melspectrogram.onnx and embedding_model.onnx are distinct from the wake word classification models — specifically:

melspectrogram.onnx was exported from a torchlibrosa-based PyTorch implementation (as shown in notebooks/converting_google_speech_embedding_model.ipynb), which is Apache 2.0 code.
embedding_model.onnx is a manual Keras re-implementation of Google's speech_embedding model, which Google releases under Apache 2.0. The README also explicitly notes this (under Model Architecture).
My questions are:

Are melspectrogram.onnx and embedding_model.onnx covered by the repository's Apache 2.0 license (not CC BY-NC-SA 4.0), given their origin from Apache 2.0 sources?
Is it permitted to use these two models in a commercial product, as long as we train our own wake word classification model on top of them (i.e., we do not use the pre-trained alexa, hey_jarvis, etc. classification models)?
Any clarification would be greatly appreciated. Thank you!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions