diff --git a/README.md b/README.md index 0ec330c..6b43be5 100644 --- a/README.md +++ b/README.md @@ -18,6 +18,33 @@ Validated audio pipeline · Multi-backend matching · Flask web UI · Terminal o ## Overview DIY Shazam captures audio from the CLI microphone/file path or the Flask browser UI, normalizes it through one bounded audio pipeline, and identifies tracks using RapidAPI/Shazam, AcoustID, AudD, or local spectrogram peaks and constellation hash pairs. FFT output is a diagnostic visualization only; it is not the recognition algorithm. Flask serves the complete browser UI and JSON API from one origin. +## Implemented + +- Flask is the supported browser application; `web/app.py` serves the UI and JSON API from one origin. +- CLI and web inputs use the documented bounded normalization pipeline. The CLI accepts WAV/PCM; the web path converts the documented browser upload formats through FFmpeg. +- Provider dispatch uses RapidAPI/Shazam, AcoustID, AudD, and the local constellation-hash backend with stable public statuses and safe diagnostics. +- Production configuration includes Gunicorn, `/healthz`, `/readyz`, bounded uploads and FFmpeg work, atomic Supabase quota operations, trusted-proxy controls, and debug-off defaults outside explicit development mode. +- The current checkout has 193 passing Python tests. CI also defines Ruff, coverage, dependency-audit, and secret-scanning gates. + +## Known limitations + +The repository does not currently claim general recognition accuracy or latency. The real-world benchmark corpus, 90 microphone clips, provider comparison, credentialed smoke test, and browser-engine state matrix are still incomplete. Provider catalog coverage, noise, re-encoding, volume or pitch changes, partial clips, live performances, covers, remixes, and alternate releases can all change the result. + +Supabase authentication, persistent user history, RLS-backed history, account deletion, user settings, and protected user routes are not implemented product features. History is session-only and optional in the browser UI. + +## Planned + +- Assemble and run the lawful 30-track / 90-clip benchmark, then import only generated results. +- Add browser/component coverage for recording, permission denial, unsupported `MediaRecorder`, cleanup, and all public result states. +- Complete browser-engine screenshots/recordings and credentialed provider smoke evidence. +- Evaluate the local constellation-hash contribution against provider baselines before changing the distinctiveness score. + +## Evaluation evidence + +- Reproducible benchmark entry points are documented in [`evaluation/README.md`](evaluation/README.md), but no complete result is present in this checkout. +- The committed FFT image at [`docs/screenshots/fft-output.png`](docs/screenshots/fft-output.png) is diagnostic output from `shazam_project.fft_analyze.analyze_audio`; the original capture command was not preserved in Git. Recreate it through the CLI's `python main.py` path after choosing `mic` or `file`. +- The current browser smoke check was run against `python web/app.py` on a local development port with debug disabled. It verified page load, status rendering, and the unsupported-upload error state; it did not use provider credentials or record microphone audio. + ## Performance | Metric | Result | @@ -156,6 +183,8 @@ Supported env vars (see `shazam_project.config.load_config()`): `AUDD_API_TOKEN` ## Web UI `python web/app.py` serves `/`, `/static/*`, `/api/match`, and `/api/status` from the same origin. CLI file mode accepts WAV/PCM files. Web uploads support WAV, MP3, M4A, AAC, OGG, FLAC, and WEBM; non-WAV web uploads require FFmpeg on `PATH` and are converted before decoding. The browser also supports microphone recording, manual stop, waveform visualization, loading/error/no-match states, light/dark theme persistence, and session-only recognition history. +The durations are intentionally different by path: CLI microphone mode defaults to 8 seconds and accepts an interactive override; the RapidAPI/Shazam adapter sends at most the first 5 seconds of the normalized clip; browser recording auto-stops after 10 seconds but can be stopped manually; and the reproducible benchmark uses separate 4-second, 8-second, and 15-second microphone clips. + Every input is downmixed to mono float32 samples in `[-1, 1]` and resampled to 44,100 Hz by default. Provider adapters receive temporary mono 16-bit PCM WAV files. Inputs shorter than 1 second, longer than 30 seconds, or larger than 10 MiB are rejected by default; all three limits are configurable. The deployment endpoints are: diff --git a/TODO.md b/TODO.md index 6ac1dea..8e5a158 100644 --- a/TODO.md +++ b/TODO.md @@ -4,13 +4,13 @@ This is the working checklist derived from the repository review on 2026-07-31. ## Current assessment -Strict evidence-based score at review time: **5/9**. +Strict evidence-based score at review time: **6/9**. | Area | Current score | Main reason | |---|---:|---| -| Correctness | 1.5/3 | Recognition works through provider integrations, but accuracy is unmeasured and several runtime paths are fragile. | -| Engineering | 2/3 | The module split and basic CI are good, but production hardening and deeper browser-state coverage remain open. | -| Documentation | 1/2 | The structure is strong, but the documents mix planned product architecture with implemented functionality. | +| Correctness | 1.5/3 | Recognition and input validation are covered by the current test suite, but real-world accuracy and browser-engine failure paths remain unmeasured. | +| Engineering | 2.5/3 | Production readiness, atomic quotas, bounded uploads, WSGI configuration, CI, and 193 local tests are evidenced; browser-state depth remains open. | +| Documentation | 1.5/2 | Implemented behavior, limitations, and evaluation gates are now separated; benchmark evidence and screenshot provenance remain open. | | Distinctiveness | 0.5/2 | The project is a capable Shazam-style integration, not yet a novel recognition system. | Do not raise these scores based on screenshots or placeholder metrics alone. Update them only after the evidence gates below are satisfied. @@ -35,6 +35,11 @@ Complete these main tasks in order. A main task may be ticked only after its sub - [ ] Replace README placeholder accuracy and timing values with generated, reviewable results. - [ ] Tick Main task 1 only after the benchmark can be reproduced from a clean checkout. +The benchmark runner now writes cache, JSON, and Markdown artifacts atomically +so interrupted runs cannot leave a truncated provider result that looks valid +on the next replay. This improves reproducibility but does not replace the +missing lawful corpus, credentials, or real benchmark execution. + ### Main task 2 — Audio-pipeline validation - [x] Decide and document that FFT is diagnostic only; spectrogram peaks and constellation hash pairs perform local recognition. @@ -55,9 +60,9 @@ Complete these main tasks in order. A main task may be ticked only after its sub ### Main task 4 — Production configuration and security - [x] Reconcile `SUPABASE_SERVICE_ROLE_KEY` with the server-only service-role security model. -- [ ] Decide whether `INTERNAL_API_SECRET` is required for the supported browser flow. -- [ ] Standardize response fields and statuses across providers and the CLI/browser consumers. -- [ ] Add a startup/health check for provider, Supabase, FFmpeg, and `fpcalc` configuration. +- [x] Decide that `INTERNAL_API_SECRET` is not required for the supported same-origin browser flow; if set, it is reserved for deliberate server-to-server callers and the browser flow is intentionally unauthorized. +- [x] Standardize response fields and statuses across providers and the CLI/browser consumers. +- [x] Add a startup/health check for provider, Supabase, FFmpeg, `fpcalc`, and writable temporary storage configuration. - [x] Make quota checks and increments atomic, and define fail-closed behavior for Supabase failures. - [x] Bound cooldown/client state and document the trusted proxy model. - [x] Add `Retry-After` headers and disable debug mode outside explicit local development. @@ -66,7 +71,7 @@ Complete these main tasks in order. A main task may be ticked only after its sub ### Main task 5 — Test-depth expansion -- [ ] Test rate-limit responses and quota accounting through public routes. +- [x] Test rate-limit responses and quota accounting through public routes. - [x] Test configuration loading, missing keys, invalid `FP_CALC_PATH`, and provider combinations. - [x] Add mocked microphone tests for invalid duration/sample rate, capture failure, and cleanup. - [ ] Add browser/component tests for recording, upload, loading, matched, no-match, unauthorized, rate-limited, and network-error states. @@ -88,10 +93,10 @@ Complete these main tasks in order. A main task may be ticked only after its sub ### Main task 7 — Documentation reconciliation -- [ ] Split documentation into Implemented, Known limitations, Planned, and Evaluation evidence sections. -- [ ] Mark unsupported Supabase/auth/history/RLS/account/settings claims as planned or remove them. +- [x] Split documentation into Implemented, Known limitations, Planned, and Evaluation evidence sections. +- [x] Mark unsupported Supabase/auth/history/RLS/account/settings claims as planned or remove them. - [x] Reconcile README content with the actual Flask source tree. -- [ ] Reconcile documented 8-second CLI, 5-second RapidAPI trim, and 10-second browser-recording behavior. +- [x] Reconcile documented 8-second CLI, 5-second RapidAPI trim, and 10-second browser-recording behavior. - [ ] Document the exact source and command for each screenshot, and add failure-state screenshots where useful. - [x] Remove unsupported platform claims and document API statuses, backend order, environment variables, and security boundaries. - [ ] Tick Main task 7 only after documentation describes shipped behavior rather than aspiration. @@ -171,7 +176,7 @@ Complete these main tasks in order. A main task may be ticked only after its sub - [x] Reconcile `SUPABASE_SERVICE_ROLE_KEY` with the documented server-only security model. - [x] Require `X-API-Secret` directly when configured; Origin/Referer never authenticates browser clients. - [x] Keep real server secrets out of browser configuration and static assets. -- [x] Make the Flask-served UI work when `INTERNAL_API_SECRET` is enabled, or remove that mode from the supported flow. Configured same-origin/allowlisted browser requests are accepted without exposing the secret. +- [x] Keep `INTERNAL_API_SECRET` out of the supported browser deployment; when deliberately configured for server-to-server callers, requests without the secret are rejected and same-origin headers never bypass authentication. - [x] Configure optional cross-origin API access from an allowlist; same-origin browser access is the default. - [x] Add RapidAPI configuration to `/api/status`; report the actual active backend order. - [x] Standardize response fields and statuses across all providers and the CLI/browser consumers, including local `no_match` responses. @@ -268,6 +273,7 @@ Record evidence here as work lands: | Date | Task/check | Evidence | Result | |---|---|---|---| +| 2026-08-17 | Current checkout audit and documentation reconciliation | Dedicated branch `codex/audio-recognition-p0-audit`; supported `.venv-pipeline` ran 193 tests; FFmpeg and fpcalc were available; local browser smoke checked page load, status rendering, and unsupported-upload handling; `/readyz`/quota/WSGI behavior is covered by repository tests. | Production/configuration and documentation checkboxes updated. No provider values, source catalog, microphone clips, benchmark results, or credentialed smoke evidence are present locally, so the real benchmark and release gates remain open. | | 2026-08-01 | Production rate limits | Added `production_rate_limits` migration through the Supabase CLI; private row-locked quota RPC, RLS with no public policies, server-only service-role access, HMAC client identifiers, fail-closed 503 handling, development fallback, trusted-proxy configuration, direct API-secret authentication, and Retry-After responses. | 93 tests passed; 68% total branch coverage; compileall and diff checks passed. Local/linked SQL execution remains unavailable: Docker is not running and the linked `shazam-project` is inactive; linked advisors returned no lints and migration listing timed out. | | 2026-08-01 | Prompt 3 cleanup and review fixes | Flask status display restored RapidAPI and Supabase fields; local `no_match` uses the shared `result: null` shape; all providers receive normalized mono float32 audio; provider diagnostics are safe; rate limits run before upload processing; fixed 16-bit provider WAV encoding is documented; README/TODO record the validated contract and Prompt 4 production blockers. | 73 tests passed; 66% total branch coverage; compileall and diff checks passed locally; CI and `coverage.xml` artifact are pending this push; Codecov upload previously reported `Repository not found` and remains deferred | | 2026-08-01 | CI/test hardening | Added configuration, microphone, Flask, generated-WAV, provider-fallback, safe-rendering, dependency, Ruff, and secret-scan gates; removed committed desktop.ini files, generated FFT artifact, and obsolete development plan. | CI run 73 passed; 146 tests passed; 74% total branch coverage; Python 3.10/3.11/3.12, Ruff, dependency audit, and Gitleaks passed; provider credentials, Supabase, FFmpeg, fpcalc, microphone hardware, and real-world benchmark are intentionally excluded | diff --git a/evaluation/README.md b/evaluation/README.md index f49ebbd..256edb2 100644 --- a/evaluation/README.md +++ b/evaluation/README.md @@ -4,6 +4,8 @@ This evaluation is deliberately gated before any recording or provider call. Sou The target corpus is 30 legally reusable or user-owned source tracks, with one 4-second, one 8-second, and one 15-second speaker-to-microphone clip per track. Record at least three conditions and document genre, era, title, artist, and provenance/license for every source row. +Current gate: this checkout contains only `sources.example.csv` and an ignored metadata cache; it does not contain the 30-track source manifest, source audio, microphone clips, or generated benchmark results. Do not treat metadata-only rows as audio permission. For Free Music Archive material, retain the individual track page and track-level license in `provenance_or_license_note`; the dataset metadata license does not replace the artist-selected audio license. The FMA dataset's official distribution and licensing notes are maintained in the [dataset repository](https://github.com/mdeff/fma). + ## 1. Prepare and validate the source catalog Copy the example and edit it with real local paths. Do not add the source files or the completed `sources.csv` to Git. diff --git a/scripts/benchmark.py b/scripts/benchmark.py index 5b59f1b..9e4ab77 100644 --- a/scripts/benchmark.py +++ b/scripts/benchmark.py @@ -159,6 +159,15 @@ def _cache_file(cache_dir: Path, key: str) -> Path: return cache_dir / f"{key}.json" +def _atomic_write_text(path: Path, content: str) -> None: + """Write benchmark artifacts without leaving truncated cache records.""" + + path.parent.mkdir(parents=True, exist_ok=True) + temporary = path.with_name(f".{path.name}.tmp") + temporary.write_text(content, encoding="utf-8") + temporary.replace(path) + + def _cache_state(cache_hits: int, cache_misses: int) -> str: if cache_hits and cache_misses: return "mixed" @@ -188,7 +197,6 @@ def _write_cache( timeout: int, result: dict[str, Any], ) -> None: - cache_dir.mkdir(parents=True, exist_ok=True) payload = { "cache_schema": CACHE_SCHEMA_VERSION, "cache_key": key, @@ -198,8 +206,9 @@ def _write_cache( "settings": _backend_settings(backend, config, timeout), "result": _safe_result(result), } - _cache_file(cache_dir, key).write_text( - json.dumps(payload, indent=2, sort_keys=True), encoding="utf-8" + _atomic_write_text( + _cache_file(cache_dir, key), + json.dumps(payload, indent=2, sort_keys=True), ) @@ -621,10 +630,9 @@ def run( "backend_summary": summaries, "records": records, } - output_path.parent.mkdir(parents=True, exist_ok=True) - output_path.write_text(json.dumps(output, indent=2, sort_keys=True), encoding="utf-8") + _atomic_write_text(output_path, json.dumps(output, indent=2, sort_keys=True)) report_path = report_path or output_path.with_suffix(".md") - report_path.write_text(render_markdown(output), encoding="utf-8") + _atomic_write_text(report_path, render_markdown(output)) return output diff --git a/showcase/audio-recognition/case-study.md b/showcase/audio-recognition/case-study.md new file mode 100644 index 0000000..c64c8c0 --- /dev/null +++ b/showcase/audio-recognition/case-study.md @@ -0,0 +1,134 @@ +# DIY Shazam — case study + +## Project overview + +DIY Shazam is an audio-recognition project with a command-line entry point and a Flask browser application. It accepts microphone input and audio files, normalizes them through a shared pipeline, and attempts recognition through configured external providers or a local constellation-hash index. + +**Current status:** Source-verified on audit branch `codex/audio-recognition-p0-audit` at `6eb46cf`, with 193 local pytest tests, Ruff, and `git diff --check` passing. PR #11 is pushed, open, and draft; no current-branch CI status entries were reported by the connected GitHub check. No live demo URL, complete real-world benchmark, or credentialed provider result is currently available. + +## Problem and audience + +The project addresses a familiar user need—identify an unknown song—but treats it as an engineering problem rather than a single API call. Audio can arrive as a microphone recording or several file formats, and every route introduces failure modes: malformed data, unsupported formats, missing dependencies, provider timeouts, no matches, quota limits, and missing configuration. + +The repository does not define a formal customer persona. For portfolio purposes, I frame the audience as developers or learners who want a transparent, self-hosted recognition workflow and a small educational implementation of audio fingerprinting. That audience framing is an interpretation, not a measured user claim. + +## Goals and constraints + +The implementation goals were to: + +- provide one supported browser path through Flask; +- share audio normalization across CLI, browser, providers, and local matching; +- keep public response fields stable and safe; +- fall through between configured backends when a provider errors or returns no match; +- bound upload size, duration, FFmpeg work, and request quotas; +- make benchmark output reproducible without exposing credentials or committing audio. + +The important constraints are external credentials, FFmpeg and `fpcalc` availability, microphone hardware, provider quotas, and legal source audio. The supported `.venv-pipeline` interpreter could not be launched without scoped elevation in this sandbox because its target Python process was inaccessible; the requested pytest command then completed successfully under that scoped launch. + +## Solution and key user flow + +The supported browser flow is: + +1. The user opens the Flask-served page. +2. The user uploads a supported format or records through the browser microphone. +3. The server checks request size and rate limits before expensive upload processing. +4. Non-WAV web input is converted through bounded FFmpeg. +5. The shared loader validates the audio, downmixes it to mono float32, resamples it, and enforces duration limits. +6. The matcher tries RapidAPI/Shazam, AcoustID, AudD, and the local index according to configuration. +7. The UI renders a safe result, retry path, no-match state, rate-limit response, or configuration/error message. + +The public response vocabulary is explicit: `matched`, `no_match`, `not_configured`, `invalid_audio`, `rate_limited`, and `error`. Provider payloads, credentials, local paths, and stack traces are not part of the browser response. + +## Design and technical decisions + +### One audio contract + +`shazam_project/recorder.py` defines the shared `AudioClip` and normalization behavior. Input is validated, downmixed, resampled to the configured internal rate—44.1 kHz by default—and bounded by duration and upload-size limits. Provider temporary files use fixed 16-bit PCM WAV. This prevents each backend from silently interpreting a different input representation. + +### Fallback instead of single-provider fragility + +`shazam_project/matcher.py` records backend attempts and distinguishes no match, missing configuration, timeout, HTTP failure, malformed response, and fingerprint errors. Eligible failures fall through to the next configured backend. The public response is filtered to a safe schema while detailed diagnostics stay internal. + +### Local fingerprinting as an educational contribution + +The local backend in `shazam_project/fingerprint.py` extracts spectral peaks, creates paired hashes, and matches consistent time offsets against a JSON index. The repository explicitly labels the FFT plot as diagnostic only. The local method is a technical learning path, not a claim of production-scale Shazam equivalence. + +### Bounded production behavior + +`web/app.py`, `Dockerfile`, `render.yaml`, and the Supabase migration define a harder production boundary: `/healthz` is liveness, `/readyz` checks readiness prerequisites, `/api/status` exposes non-secret capability flags, quotas can fail closed, and the container runs as a non-root user with read-only filesystem expectations and a temporary filesystem. + +### Reproducible evaluation before published numbers + +The benchmark runner requires 30 source tracks, one 4-, 8-, and 15-second microphone clip per source, at least three recording conditions, stable identifiers where available, operator metadata, and all selected backends. It caches safe normalized results and refuses incomplete output through the README updater. This is intentionally conservative: no accuracy or latency number is published until the data exists. + +## Implementation highlights + +- Flask `web/` is the single supported browser application. +- The browser supports upload, recording, waveform visualization, light/dark theme, retry states, and optional session-only history. +- Web uploads support WAV, MP3, M4A, AAC, OGG, FLAC, and WEBM; FFmpeg conversion is bounded by duration and output-size limits. +- The current audit branch (`6eb46cf`) includes the `b0b8d79` audio-pipeline/browser-recovery hardening baseline plus the documentation and evaluation-gate reconciliation. +- CI runs Python tests on 3.10, 3.11, and 3.12, Ruff formatting/lint, branch coverage, dependency auditing, secret scanning, Render Blueprint validation, and a production-container smoke path. + +## Validation and results + +### Verified in repository source + +- The supported routes are `/`, `/healthz`, `/readyz`, `/api/match`, and `/api/status`. +- The response status vocabulary and audio contract are documented and covered by Python tests. +- The repository contains a reproducible benchmark command and refuses incomplete README imports. +- The checked-in FFT image is a real project artifact, but it is a diagnostic spectrum and not evidence of recognition quality. + +### Verified locally on the current audit branch + +- `.venv-pipeline\Scripts\python.exe -m pytest -q` completed with `193 passed in 9.56s`; the initial non-elevated launcher failed before test collection because it could not create the configured Python process. +- `.venv-pipeline\Scripts\ruff.exe check .` passed, and `git diff --check` produced no output. +- A local development HTTP smoke returned 200 for `/` and `/healthz`, exposed upload and record controls, returned non-secret `/api/status`, returned the stable `invalid_audio` response for a malformed upload, and returned `/readyz` 503 because no recognition backend is configured. +- `sounddevice.query_devices()` listed 33 host devices, including input and output devices. No microphone recording or provider call was made. + +### Verified remotely on merged `main` + +The merged [PR #8](https://github.com/icecold009/Audio-Recognition/pull/8) records GitHub Actions run `30738959312`: 190 tests passed in each Python 3.10, 3.11, and 3.12 matrix job; branch coverage reached 81% against a 70% threshold; Ruff, `pip-audit`, Gitleaks, Render schema validation, production image build, and external container smoke passed. The smoke path exercised `/`, `/healthz`, the expected production `/readyz` and `/api/status` failures without configuration, malformed upload rejection, and a mocked successful recognition response. + +### Not yet verified + +- The current branch has not received a remote CI run; local checks are not equivalent to the full hosted matrix. +- No complete real-world corpus has been recorded. +- No credentialed provider comparison, recognition accuracy, p95 latency, catalog coverage, or live deployment smoke is claimed. +- The in-app browser could not reach the elevated local listener, so no browser-engine/permission screenshot or real browser microphone capture is claimed. Source and mocked tests cover the recording fallback and cleanup paths. + +## Challenges and tradeoffs + +The project favors explicit boundaries over optimistic demos. That means the README intentionally shows no benchmark number until the corpus and metadata are complete. The same choice makes the setup heavier: legal source provenance, microphone recording, external provider quotas, FFmpeg, `fpcalc`, and operator metadata all become prerequisites. + +The local matcher is valuable because it is inspectable and can run without a provider call, but it is not a substitute for a measured production catalog. The provider fallback improves resilience, but it also makes performance and correctness backend-dependent. Session history is convenient, but it remains browser-session state rather than a persistent account feature. + +## What I learned + +The main lesson is that “it returned a song once” is not a useful quality claim. A trustworthy showcase needs traceable input data, exact denominators, failure categories, latency methodology, and clear separation between source behavior, CI evidence, live verification, and planned work. I also learned that the audio contract should be centralized early; otherwise each provider becomes a separate interpretation of the same recording. + +## Future improvements + +1. Assemble and document the legal 30-source corpus and 90 clips per backend. +2. Run the benchmark with reviewed credentials and publish generated results only after the completeness gate passes. +3. Add browser-engine smoke coverage for recording, upload, permission denial, unsupported `MediaRecorder`, result states, and resource cleanup. +4. Capture a clean screenshot set from the supported Flask app and attach commands/captions to each image. +5. Make the supported pipeline interpreter launch without elevation in the developer environment and run the full release gate on the current branch. +6. Verify a deployed instance separately if a live URL becomes available. + +## Technologies used + +Python, Flask, NumPy, SciPy, SoundFile, SoundDevice, Requests, Matplotlib, FFmpeg, Chromaprint/fpcalc, Gunicorn, Supabase/Postgres quota RPCs, Docker, Render Blueprint, pytest, coverage.py, Ruff, pip-audit, and Gitleaks. + +## Evidence sources + +- [`README.md`](../../README.md) +- [`TODO.md`](../../TODO.md) +- [`evaluation/README.md`](../../evaluation/README.md) +- [`web/app.py`](../../web/app.py) +- [`web/templates/index.html`](../../web/templates/index.html) +- [`web/static/app.js`](../../web/static/app.js) +- [`shazam_project/recorder.py`](../../shazam_project/recorder.py) +- [`shazam_project/matcher.py`](../../shazam_project/matcher.py) +- [`shazam_project/fingerprint.py`](../../shazam_project/fingerprint.py) +- [`Dockerfile`](../../Dockerfile) +- [`render.yaml`](../../render.yaml) diff --git a/showcase/audio-recognition/diy-shazam-showcase.pptx b/showcase/audio-recognition/diy-shazam-showcase.pptx new file mode 100644 index 0000000..b0286c7 Binary files /dev/null and b/showcase/audio-recognition/diy-shazam-showcase.pptx differ diff --git a/showcase/audio-recognition/evidence-checklist.md b/showcase/audio-recognition/evidence-checklist.md new file mode 100644 index 0000000..cb6fe09 --- /dev/null +++ b/showcase/audio-recognition/evidence-checklist.md @@ -0,0 +1,58 @@ +# DIY Shazam — screenshot and evidence checklist + +**Evidence rule:** Use actual repository output, test output, or a real running app. Do not fabricate product screenshots, recognition results, accuracy, latency, users, or deployment claims. + +## Current evidence inventory + +| Evidence | Status | Use in showcase | +|---|---|---| +| Repository source and documentation | Verified at `6eb46cf` | Supports the architecture, UI, API contract, limits, matcher behavior, evaluation gates, and limitations. | +| Existing FFT image | Verified: [`docs/screenshots/fft-output.png`](../../docs/screenshots/fft-output.png) | Technical visual only; caption it as “Diagnostic frequency spectrum; FFT is not the recognition algorithm.” | +| Merged-main CI | Verified remotely in [PR #8](https://github.com/icecold009/Audio-Recognition/pull/8) | Supports 190 tests per Python 3.10–3.12 job, 81% branch coverage, static/security gates, Render validation, and container smoke on merged main. | +| Current audit branch CI | Not reported by connected GitHub check | PR #11 is pushed and draft; local pytest (193), Ruff, and `git diff --check` passed, but local checks are not the hosted matrix/release gate. | +| Live demo | Missing | Supply a public URL, then run a browser smoke check and label it live evidence. | +| Real product screenshots | Missing | Capture from the supported Flask app; no browser screenshot was available in this task. | +| Real-world benchmark results | Missing by design | Assemble the legal corpus and run `scripts/benchmark.py`; do not infer quality from synthetic tests or CI. | +| Credentialed provider smoke | Missing | Configure credentials outside Git and record a redacted, reproducible smoke result. | +| Fresh local test run | Verified with scoped launch | `.venv-pipeline\Scripts\python.exe -m pytest -q` reported `193 passed in 9.56s`; the initial non-elevated launch failed before collection because the configured Python process was inaccessible. | +| Host audio device enumeration | Verified, capture not attempted | `sounddevice.query_devices()` listed 33 devices, including microphone inputs and speaker outputs. This does not prove a real recording or provider recognition result. | + +## Screenshot plan + +| Suggested file | What to capture | Source/command | Caption | Status | +|---|---|---|---|---| +| `screenshots/01-home-empty.png` | Flask home page with upload, record, and status controls | Start the supported app with `python web/app.py`, then capture `/` | “The single supported Flask browser entry point.” | **[NEEDS EVIDENCE]** | +| `screenshots/02-status-development.png` | Non-secret `/api/status` capability flags | `GET /api/status` in the local app | “Runtime capability status without secrets.” | **[NEEDS EVIDENCE]** | +| `screenshots/03-invalid-upload.png` | Malformed or unsupported upload response | `POST /api/match` with a controlled invalid fixture | “Invalid audio is rejected with a stable public status.” | **[NEEDS EVIDENCE]** | +| `screenshots/04-mock-match.png` | Successful mocked recognition response | Reproduce the CI container smoke fixture or a local deterministic mock | “The UI renders a matched response without exposing provider internals.” | **[NEEDS EVIDENCE]** | +| `screenshots/05-no-match-or-retry.png` | No-match or provider-error recovery state | Use a deterministic mocked response | “Failure states offer a clear retry path.” | **[NEEDS EVIDENCE]** | +| `screenshots/06-fft-diagnostic.png` | Existing frequency spectrum | [`docs/screenshots/fft-output.png`](../../docs/screenshots/fft-output.png), generated by `main.py`/`fft_analyze.py` | “Diagnostic spectrum; not song identification.” | Verified existing artifact | + +## Claims checklist + +### Safe to state now + +- Flask `web/` is the supported browser implementation. +- CLI and browser inputs share a normalization contract. +- Web supports the documented formats and bounded FFmpeg conversion. +- Matcher order and public response statuses are documented in source. +- The local backend uses spectral peaks, hash pairs, and time-offset consensus. +- Merged-main remote CI recorded 190 tests per Python version, 81% branch coverage, lint/audit/secret gates, Render schema validation, and container smoke. +- The benchmark tooling refuses incomplete result imports and keeps source audio/provider credentials out of the repository. + +### Must remain labeled as missing or unverified + +- Recognition accuracy or latency on real recordings. +- Provider superiority, catalog coverage, or robustness under noise/re-encoding. +- Production uptime, user adoption, testimonials, or deployment readiness. +- A live hosted URL or browser-engine compatibility. +- Current-branch remote CI status after `6eb46cf`. +- Browser-engine permission state, real microphone capture, and device-to-provider recognition. + +## Follow-up capture procedure + +1. Make the supported pipeline interpreter launch without elevation, or document the scoped launch requirement for the developer environment. +2. Run the full release gate, coverage, compile, Ruff, and diff checks; record the exact commit and output. +3. Start `python web/app.py` in development mode and capture the home, status, invalid-upload, no-match, and matched-mock states. +4. If a live URL is provided, verify the same journey against the deployed app and label screenshots as live rather than local. +5. Assemble the legally reusable evaluation corpus with track-level provenance/license notes, run the documented benchmark, review generated JSON/Markdown, and import results only if the completeness gate passes. diff --git a/showcase/audio-recognition/presentation-outline.md b/showcase/audio-recognition/presentation-outline.md new file mode 100644 index 0000000..9368c1c --- /dev/null +++ b/showcase/audio-recognition/presentation-outline.md @@ -0,0 +1,142 @@ +# DIY Shazam — presentation outline + +**Status:** Source-verified showcase draft. The repository is a locally validated prototype, not a verified live production demo. This audit PR branch is `codex/audio-recognition-p0-audit` at `6eb46cf`; local pytest (193), Ruff, and diff checks passed, while the latest remote CI evidence applies to merged `origin/main`/PR #8 and no current-branch CI status entries are reported. + +**Audience:** General portfolio audience + +**Purpose:** Portfolio / technical project walkthrough + +**Length:** 7 slides, approximately 2 minutes spoken + +**Live demo:** None supplied; browser screenshots remain an evidence gap. + +## Slide 1 — DIY Shazam + +**Main message:** A self-hosted song-recognition workflow that accepts a microphone recording or audio file and routes it through multiple matching backends. + +**On-slide points:** + +- Identify a song from a microphone or upload +- Flask browser UI plus CLI entry point +- RapidAPI/Shazam, AcoustID, AudD, and local fingerprint backends + +**Recommended visual:** A real screenshot of the Flask home screen showing the upload and recording controls. **[NEEDS EVIDENCE]** No live demo or current browser screenshot was supplied. Use `docs/screenshots/fft-output.png` only as a technical fallback, labeled as a diagnostic spectrum rather than a product screenshot. + +**Speaker notes:** “This is DIY Shazam: a song-recognition project built around a practical input pipeline and several interchangeable backends. The interesting part is not just calling an API; it is making microphone input, uploaded files, validation, fallback, and safe failure states behave consistently.” + +**Evidence:** [`README.md`](../../README.md), [`web/templates/index.html`](../../web/templates/index.html), [`main.py`](../../main.py) + +## Slide 2 — The problem and the user + +**Main message:** Recognition is easy to demo with one clean clip, but reliable behavior requires controlling input formats, duration, dependencies, provider failures, and honest evaluation. + +**On-slide points:** + +- Audio arrives from different devices and formats +- Providers can be missing, slow, unavailable, or return no match +- Upload and quota limits must protect the server +- The target user is a developer or learner exploring a DIY recognition workflow + +**Recommended visual:** A simple input-to-decision flow: microphone/file → validation → matching → result or safe failure. + +**Speaker notes:** “The project is aimed at a developer or learner who wants to understand the full path from audio input to recognition. The hard engineering problem is the boundary between messy audio and a dependable response: malformed files, unsupported formats, provider errors, rate limits, and no-match outcomes all need explicit handling.” + +**Evidence:** [`README.md`](../../README.md), [`TODO.md`](../../TODO.md), [`docs/01-product-requirements.md`](../../docs/01-product-requirements.md) + +## Slide 3 — The primary user journey + +**Main message:** The browser keeps the interaction simple while the server owns normalization, matching, and safe public responses. + +**On-slide points:** + +- Upload WAV, MP3, M4A, AAC, OGG, FLAC, or WEBM, or record in-browser +- Convert non-WAV web uploads through bounded FFmpeg +- Normalize to mono float32 at the configured internal rate +- Return `matched`, `no_match`, `not_configured`, `invalid_audio`, `rate_limited`, or `error` + +**Recommended visual:** A six-step journey diagram with the six public statuses as the outcome set. + +**Speaker notes:** “From the user’s perspective, the flow is deliberately short: choose a file or record, submit, and read the result. The server converts supported web formats, validates size and duration, normalizes the signal, tries the configured matcher order, and exposes a stable response shape without leaking provider payloads or secrets.” + +**Evidence:** [`web/templates/index.html`](../../web/templates/index.html), [`web/app.py`](../../web/app.py), [`shazam_project/recorder.py`](../../shazam_project/recorder.py), [`README.md`](../../README.md) + +## Slide 4 — Product walkthrough + +**Main message:** The Flask UI is a focused recognition surface with recovery-oriented states, not a collection of disconnected demos. + +**On-slide points:** + +- Recording controls and waveform visualizer +- File upload and server capability status +- Matched, no-match, invalid-input, rate-limit, and retry states +- Light/dark theme and optional session-only history + +**Recommended visual:** A three-panel screenshot sequence: empty home, result/error state, and history/details. **[NEEDS EVIDENCE]** Capture these from the actual Flask app after restoring a usable runtime or supplying a live URL. + +**Speaker notes:** “The UI is intentionally small: the main job is to capture audio and communicate the next state clearly. The current branch also hardens browser recovery by making storage optional and stopping microphone tracks and visualizer resources when recording ends or fails.” + +**Evidence:** [`web/templates/index.html`](../../web/templates/index.html), [`web/static/app.js`](../../web/static/app.js), [`tests/test_web.py`](../../tests/test_web.py), [`tests/test_ci_hardening.py`](../../tests/test_ci_hardening.py) + +## Slide 5 — Technical approach + +**Main message:** One shared audio contract feeds a fallback dispatcher and an educational local constellation-hash implementation. + +**On-slide points:** + +- Downmix, validate, resample, and bound audio once +- Try RapidAPI/Shazam → AcoustID → AudD → local index as configured +- Local matching uses spectral peaks, hash pairs, and time-offset consensus +- Production path adds FFmpeg, Gunicorn, health/readiness checks, and quota controls + +**Recommended visual:** + +```mermaid +flowchart LR + input["Mic or file"] --> web["Flask UI / CLI"] + web --> pipeline["Validate + normalize"] + pipeline --> dispatch["Matcher dispatcher"] + dispatch --> rapid["RapidAPI / Shazam"] + dispatch --> acoustid["AcoustID"] + dispatch --> audd["AudD"] + dispatch --> local["Local spectral peaks + hash pairs"] + dispatch --> response["Safe public response"] +``` + +**Speaker notes:** “The central design decision is the shared `AudioClip` contract: mono floating-point audio at one internal sample rate, with explicit duration and size limits. The dispatcher can fall through after a provider error, no match, or missing configuration. The local backend is a transparent learning implementation based on spectral peaks and hash-pair offsets; FFT output is diagnostic only.” + +**Evidence:** [`shazam_project/recorder.py`](../../shazam_project/recorder.py), [`shazam_project/matcher.py`](../../shazam_project/matcher.py), [`shazam_project/fingerprint.py`](../../shazam_project/fingerprint.py), [`Dockerfile`](../../Dockerfile), [`render.yaml`](../../render.yaml) + +## Slide 6 — Validation and current proof + +**Main message:** The engineering gates are substantial, but recognition quality is not yet proven by a real-world corpus. + +**On-slide points:** + +- Merged-main CI: 190 tests per Python 3.10–3.12 matrix job +- Merged-main branch coverage: 81%, above the 70% gate +- Ruff, pip-audit, Gitleaks, Render schema, and container smoke passed remotely +- Benchmark tooling exists, but no complete corpus or credentialed provider run is imported +- Current audit branch has local checks but is not yet separately remote-CI-verified + +**Recommended visual:** A two-column “verified / still open” scorecard. Do not show invented accuracy or latency values. + +**Speaker notes:** “The strongest remote evidence is engineering validation on merged main: the CI run covered tests, coverage, linting, dependency and secret scans, Render schema validation, and production-container smoke. This audit branch also passes 193 local tests, Ruff, and diff checks, but that does not replace remote CI or prove song-recognition quality. The benchmark is deliberately gated until there are 30 legally reusable source tracks, 90 microphone clips per backend, three recording conditions, operator metadata, and provider configuration.” + +**Evidence:** [Merged PR #8](https://github.com/icecold009/Audio-Recognition/pull/8), [`evaluation/README.md`](../../evaluation/README.md), [`scripts/benchmark.py`](../../scripts/benchmark.py), [`TODO.md`](../../TODO.md) + +## Slide 7 — Lessons, limitations, and next steps + +**Main message:** The next milestone is evidence quality, not another feature. + +**On-slide points:** + +- No real-world accuracy, latency, or catalog-coverage claims yet +- Browser-engine tests and screenshot evidence remain open +- Provider credentials, FFmpeg/fpcalc host setup, and a clean local runtime are external prerequisites +- Next: run the reproducible benchmark, add browser smoke coverage, re-run CI on this branch, then publish only generated results + +**Recommended visual:** A short “now → next” roadmap. + +**Speaker notes:** “The project taught me to separate a convincing demo from a trustworthy claim. The next useful work is not to add another backend; it is to collect legal, reproducible evidence, verify the browser journey, and make the current branch pass the same release gates. Until then, the honest status is: technically hardened and CI-validated on merged main, but not quality-benchmarked or live-demo verified.” + +**Evidence:** [`TODO.md`](../../TODO.md), [`evaluation/README.md`](../../evaluation/README.md), [`evidence-checklist.md`](evidence-checklist.md) diff --git a/showcase/audio-recognition/presentation-script.md b/showcase/audio-recognition/presentation-script.md new file mode 100644 index 0000000..61eeaa6 --- /dev/null +++ b/showcase/audio-recognition/presentation-script.md @@ -0,0 +1,17 @@ +# DIY Shazam — spoken presentation script + +**Target length:** Approximately 2 minutes + +**Status:** Evidence-backed script. Claims about merged-main CI remain separate from the current audit PR branch, which has local checks but no current-branch CI status entries. + +This is DIY Shazam, a song-recognition project that accepts either a microphone recording or an audio file and routes it through multiple matching backends. The project is designed as a practical end-to-end workflow: input handling, normalization, provider fallback, safe errors, and evaluation all matter as much as the happy-path match. + +The problem is that audio is messy. A user may upload a different format, provide a clip that is too short or too long, lose microphone access, or hit a provider that is unavailable or has no result. The application therefore makes those states explicit instead of treating every failure as a zero or a server crash. + +The main journey is simple. In the Flask browser UI, the user uploads WAV, MP3, M4A, AAC, OGG, FLAC, or WEBM, or records directly from the microphone. Web formats that need it go through bounded FFmpeg conversion. The shared pipeline downmixes to mono float32, resamples to the configured internal rate, enforces size and duration limits, and then passes the normalized clip to a matcher dispatcher. + +The dispatcher tries configured backends in order: RapidAPI/Shazam, AcoustID, AudD, and an educational local fingerprint index. The local path extracts spectral peaks, pairs them into hashes, and looks for consistent time offsets. FFT output is only a diagnostic visualization; it is not presented as the identification algorithm. Public responses stay deliberately small and use stable statuses such as `matched`, `no_match`, `not_configured`, `invalid_audio`, `rate_limited`, and `error`. + +The engineering evidence is stronger than the recognition-quality evidence. Merged main has remote CI evidence showing 190 tests on each Python 3.10, 3.11, and 3.12 job, 81% branch coverage against a 70% gate, plus Ruff, dependency and secret scans, Render schema validation, and container smoke. The current audit branch `codex/audio-recognition-p0-audit` at `6eb46cf` also passes 193 local tests, Ruff, and `git diff --check`, but has not yet received a separate remote CI run. There is also no complete real-world benchmark yet: the repository still needs a legally reusable 30-track corpus, 90 microphone clips per backend, three recording conditions, provider configuration, and operator metadata before accuracy or latency can be claimed. + +The main lesson is to make the evidence boundary visible. The next milestone is not another feature; it is a reproducible benchmark, browser smoke coverage, screenshot evidence, and a fresh CI run on this branch. Today, the honest status is a technically hardened, CI-validated project on merged main, with live-demo and real-world recognition quality still unverified.