Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,33 @@ Validated audio pipeline · Multi-backend matching · Flask web UI · Terminal o
## Overview
DIY Shazam captures audio from the CLI microphone/file path or the Flask browser UI, normalizes it through one bounded audio pipeline, and identifies tracks using RapidAPI/Shazam, AcoustID, AudD, or local spectrogram peaks and constellation hash pairs. FFT output is a diagnostic visualization only; it is not the recognition algorithm. Flask serves the complete browser UI and JSON API from one origin.

## Implemented

- Flask is the supported browser application; `web/app.py` serves the UI and JSON API from one origin.
- CLI and web inputs use the documented bounded normalization pipeline. The CLI accepts WAV/PCM; the web path converts the documented browser upload formats through FFmpeg.
- Provider dispatch uses RapidAPI/Shazam, AcoustID, AudD, and the local constellation-hash backend with stable public statuses and safe diagnostics.
- Production configuration includes Gunicorn, `/healthz`, `/readyz`, bounded uploads and FFmpeg work, atomic Supabase quota operations, trusted-proxy controls, and debug-off defaults outside explicit development mode.
- The current checkout has 193 passing Python tests. CI also defines Ruff, coverage, dependency-audit, and secret-scanning gates.

## Known limitations

The repository does not currently claim general recognition accuracy or latency. The real-world benchmark corpus, 90 microphone clips, provider comparison, credentialed smoke test, and browser-engine state matrix are still incomplete. Provider catalog coverage, noise, re-encoding, volume or pitch changes, partial clips, live performances, covers, remixes, and alternate releases can all change the result.

Supabase authentication, persistent user history, RLS-backed history, account deletion, user settings, and protected user routes are not implemented product features. History is session-only and optional in the browser UI.

## Planned

- Assemble and run the lawful 30-track / 90-clip benchmark, then import only generated results.
- Add browser/component coverage for recording, permission denial, unsupported `MediaRecorder`, cleanup, and all public result states.
- Complete browser-engine screenshots/recordings and credentialed provider smoke evidence.
- Evaluate the local constellation-hash contribution against provider baselines before changing the distinctiveness score.

## Evaluation evidence

- Reproducible benchmark entry points are documented in [`evaluation/README.md`](evaluation/README.md), but no complete result is present in this checkout.
- The committed FFT image at [`docs/screenshots/fft-output.png`](docs/screenshots/fft-output.png) is diagnostic output from `shazam_project.fft_analyze.analyze_audio`; the original capture command was not preserved in Git. Recreate it through the CLI's `python main.py` path after choosing `mic` or `file`.
- The current browser smoke check was run against `python web/app.py` on a local development port with debug disabled. It verified page load, status rendering, and the unsupported-upload error state; it did not use provider credentials or record microphone audio.

## Performance

| Metric | Result |
Expand Down Expand Up @@ -156,6 +183,8 @@ Supported env vars (see `shazam_project.config.load_config()`): `AUDD_API_TOKEN`
## Web UI
`python web/app.py` serves `/`, `/static/*`, `/api/match`, and `/api/status` from the same origin. CLI file mode accepts WAV/PCM files. Web uploads support WAV, MP3, M4A, AAC, OGG, FLAC, and WEBM; non-WAV web uploads require FFmpeg on `PATH` and are converted before decoding. The browser also supports microphone recording, manual stop, waveform visualization, loading/error/no-match states, light/dark theme persistence, and session-only recognition history.

The durations are intentionally different by path: CLI microphone mode defaults to 8 seconds and accepts an interactive override; the RapidAPI/Shazam adapter sends at most the first 5 seconds of the normalized clip; browser recording auto-stops after 10 seconds but can be stopped manually; and the reproducible benchmark uses separate 4-second, 8-second, and 15-second microphone clips.

Every input is downmixed to mono float32 samples in `[-1, 1]` and resampled to 44,100 Hz by default. Provider adapters receive temporary mono 16-bit PCM WAV files. Inputs shorter than 1 second, longer than 30 seconds, or larger than 10 MiB are rejected by default; all three limits are configurable.

The deployment endpoints are:
Expand Down
30 changes: 18 additions & 12 deletions TODO.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,13 @@ This is the working checklist derived from the repository review on 2026-07-31.

## Current assessment

Strict evidence-based score at review time: **5/9**.
Strict evidence-based score at review time: **6/9**.

| Area | Current score | Main reason |
|---|---:|---|
| Correctness | 1.5/3 | Recognition works through provider integrations, but accuracy is unmeasured and several runtime paths are fragile. |
| Engineering | 2/3 | The module split and basic CI are good, but production hardening and deeper browser-state coverage remain open. |
| Documentation | 1/2 | The structure is strong, but the documents mix planned product architecture with implemented functionality. |
| Correctness | 1.5/3 | Recognition and input validation are covered by the current test suite, but real-world accuracy and browser-engine failure paths remain unmeasured. |
| Engineering | 2.5/3 | Production readiness, atomic quotas, bounded uploads, WSGI configuration, CI, and 193 local tests are evidenced; browser-state depth remains open. |
| Documentation | 1.5/2 | Implemented behavior, limitations, and evaluation gates are now separated; benchmark evidence and screenshot provenance remain open. |
| Distinctiveness | 0.5/2 | The project is a capable Shazam-style integration, not yet a novel recognition system. |

Do not raise these scores based on screenshots or placeholder metrics alone. Update them only after the evidence gates below are satisfied.
Expand All @@ -35,6 +35,11 @@ Complete these main tasks in order. A main task may be ticked only after its sub
- [ ] Replace README placeholder accuracy and timing values with generated, reviewable results.
- [ ] Tick Main task 1 only after the benchmark can be reproduced from a clean checkout.

The benchmark runner now writes cache, JSON, and Markdown artifacts atomically
so interrupted runs cannot leave a truncated provider result that looks valid
on the next replay. This improves reproducibility but does not replace the
missing lawful corpus, credentials, or real benchmark execution.

### Main task 2 — Audio-pipeline validation

- [x] Decide and document that FFT is diagnostic only; spectrogram peaks and constellation hash pairs perform local recognition.
Expand All @@ -55,9 +60,9 @@ Complete these main tasks in order. A main task may be ticked only after its sub
### Main task 4 — Production configuration and security

- [x] Reconcile `SUPABASE_SERVICE_ROLE_KEY` with the server-only service-role security model.
- [ ] Decide whether `INTERNAL_API_SECRET` is required for the supported browser flow.
- [ ] Standardize response fields and statuses across providers and the CLI/browser consumers.
- [ ] Add a startup/health check for provider, Supabase, FFmpeg, and `fpcalc` configuration.
- [x] Decide that `INTERNAL_API_SECRET` is not required for the supported same-origin browser flow; if set, it is reserved for deliberate server-to-server callers and the browser flow is intentionally unauthorized.
- [x] Standardize response fields and statuses across providers and the CLI/browser consumers.
- [x] Add a startup/health check for provider, Supabase, FFmpeg, `fpcalc`, and writable temporary storage configuration.
- [x] Make quota checks and increments atomic, and define fail-closed behavior for Supabase failures.
- [x] Bound cooldown/client state and document the trusted proxy model.
- [x] Add `Retry-After` headers and disable debug mode outside explicit local development.
Expand All @@ -66,7 +71,7 @@ Complete these main tasks in order. A main task may be ticked only after its sub

### Main task 5 — Test-depth expansion

- [ ] Test rate-limit responses and quota accounting through public routes.
- [x] Test rate-limit responses and quota accounting through public routes.
- [x] Test configuration loading, missing keys, invalid `FP_CALC_PATH`, and provider combinations.
- [x] Add mocked microphone tests for invalid duration/sample rate, capture failure, and cleanup.
- [ ] Add browser/component tests for recording, upload, loading, matched, no-match, unauthorized, rate-limited, and network-error states.
Expand All @@ -88,10 +93,10 @@ Complete these main tasks in order. A main task may be ticked only after its sub

### Main task 7 — Documentation reconciliation

- [ ] Split documentation into Implemented, Known limitations, Planned, and Evaluation evidence sections.
- [ ] Mark unsupported Supabase/auth/history/RLS/account/settings claims as planned or remove them.
- [x] Split documentation into Implemented, Known limitations, Planned, and Evaluation evidence sections.
- [x] Mark unsupported Supabase/auth/history/RLS/account/settings claims as planned or remove them.
- [x] Reconcile README content with the actual Flask source tree.
- [ ] Reconcile documented 8-second CLI, 5-second RapidAPI trim, and 10-second browser-recording behavior.
- [x] Reconcile documented 8-second CLI, 5-second RapidAPI trim, and 10-second browser-recording behavior.
- [ ] Document the exact source and command for each screenshot, and add failure-state screenshots where useful.
- [x] Remove unsupported platform claims and document API statuses, backend order, environment variables, and security boundaries.
- [ ] Tick Main task 7 only after documentation describes shipped behavior rather than aspiration.
Expand Down Expand Up @@ -171,7 +176,7 @@ Complete these main tasks in order. A main task may be ticked only after its sub
- [x] Reconcile `SUPABASE_SERVICE_ROLE_KEY` with the documented server-only security model.
- [x] Require `X-API-Secret` directly when configured; Origin/Referer never authenticates browser clients.
- [x] Keep real server secrets out of browser configuration and static assets.
- [x] Make the Flask-served UI work when `INTERNAL_API_SECRET` is enabled, or remove that mode from the supported flow. Configured same-origin/allowlisted browser requests are accepted without exposing the secret.
- [x] Keep `INTERNAL_API_SECRET` out of the supported browser deployment; when deliberately configured for server-to-server callers, requests without the secret are rejected and same-origin headers never bypass authentication.
- [x] Configure optional cross-origin API access from an allowlist; same-origin browser access is the default.
- [x] Add RapidAPI configuration to `/api/status`; report the actual active backend order.
- [x] Standardize response fields and statuses across all providers and the CLI/browser consumers, including local `no_match` responses.
Expand Down Expand Up @@ -268,6 +273,7 @@ Record evidence here as work lands:

| Date | Task/check | Evidence | Result |
|---|---|---|---|
| 2026-08-17 | Current checkout audit and documentation reconciliation | Dedicated branch `codex/audio-recognition-p0-audit`; supported `.venv-pipeline` ran 193 tests; FFmpeg and fpcalc were available; local browser smoke checked page load, status rendering, and unsupported-upload handling; `/readyz`/quota/WSGI behavior is covered by repository tests. | Production/configuration and documentation checkboxes updated. No provider values, source catalog, microphone clips, benchmark results, or credentialed smoke evidence are present locally, so the real benchmark and release gates remain open. |
| 2026-08-01 | Production rate limits | Added `production_rate_limits` migration through the Supabase CLI; private row-locked quota RPC, RLS with no public policies, server-only service-role access, HMAC client identifiers, fail-closed 503 handling, development fallback, trusted-proxy configuration, direct API-secret authentication, and Retry-After responses. | 93 tests passed; 68% total branch coverage; compileall and diff checks passed. Local/linked SQL execution remains unavailable: Docker is not running and the linked `shazam-project` is inactive; linked advisors returned no lints and migration listing timed out. |
| 2026-08-01 | Prompt 3 cleanup and review fixes | Flask status display restored RapidAPI and Supabase fields; local `no_match` uses the shared `result: null` shape; all providers receive normalized mono float32 audio; provider diagnostics are safe; rate limits run before upload processing; fixed 16-bit provider WAV encoding is documented; README/TODO record the validated contract and Prompt 4 production blockers. | 73 tests passed; 66% total branch coverage; compileall and diff checks passed locally; CI and `coverage.xml` artifact are pending this push; Codecov upload previously reported `Repository not found` and remains deferred |
| 2026-08-01 | CI/test hardening | Added configuration, microphone, Flask, generated-WAV, provider-fallback, safe-rendering, dependency, Ruff, and secret-scan gates; removed committed desktop.ini files, generated FFT artifact, and obsolete development plan. | CI run 73 passed; 146 tests passed; 74% total branch coverage; Python 3.10/3.11/3.12, Ruff, dependency audit, and Gitleaks passed; provider credentials, Supabase, FFmpeg, fpcalc, microphone hardware, and real-world benchmark are intentionally excluded |
Expand Down
2 changes: 2 additions & 0 deletions evaluation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,8 @@ This evaluation is deliberately gated before any recording or provider call. Sou

The target corpus is 30 legally reusable or user-owned source tracks, with one 4-second, one 8-second, and one 15-second speaker-to-microphone clip per track. Record at least three conditions and document genre, era, title, artist, and provenance/license for every source row.

Current gate: this checkout contains only `sources.example.csv` and an ignored metadata cache; it does not contain the 30-track source manifest, source audio, microphone clips, or generated benchmark results. Do not treat metadata-only rows as audio permission. For Free Music Archive material, retain the individual track page and track-level license in `provenance_or_license_note`; the dataset metadata license does not replace the artist-selected audio license. The FMA dataset's official distribution and licensing notes are maintained in the [dataset repository](https://github.com/mdeff/fma).

## 1. Prepare and validate the source catalog

Copy the example and edit it with real local paths. Do not add the source files or the completed `sources.csv` to Git.
Expand Down
20 changes: 14 additions & 6 deletions scripts/benchmark.py
Original file line number Diff line number Diff line change
Expand Up @@ -159,6 +159,15 @@ def _cache_file(cache_dir: Path, key: str) -> Path:
return cache_dir / f"{key}.json"


def _atomic_write_text(path: Path, content: str) -> None:
"""Write benchmark artifacts without leaving truncated cache records."""

path.parent.mkdir(parents=True, exist_ok=True)
temporary = path.with_name(f".{path.name}.tmp")
temporary.write_text(content, encoding="utf-8")
temporary.replace(path)


def _cache_state(cache_hits: int, cache_misses: int) -> str:
if cache_hits and cache_misses:
return "mixed"
Expand Down Expand Up @@ -188,7 +197,6 @@ def _write_cache(
timeout: int,
result: dict[str, Any],
) -> None:
cache_dir.mkdir(parents=True, exist_ok=True)
payload = {
"cache_schema": CACHE_SCHEMA_VERSION,
"cache_key": key,
Expand All @@ -198,8 +206,9 @@ def _write_cache(
"settings": _backend_settings(backend, config, timeout),
"result": _safe_result(result),
}
_cache_file(cache_dir, key).write_text(
json.dumps(payload, indent=2, sort_keys=True), encoding="utf-8"
_atomic_write_text(
_cache_file(cache_dir, key),
json.dumps(payload, indent=2, sort_keys=True),
)


Expand Down Expand Up @@ -621,10 +630,9 @@ def run(
"backend_summary": summaries,
"records": records,
}
output_path.parent.mkdir(parents=True, exist_ok=True)
output_path.write_text(json.dumps(output, indent=2, sort_keys=True), encoding="utf-8")
_atomic_write_text(output_path, json.dumps(output, indent=2, sort_keys=True))
report_path = report_path or output_path.with_suffix(".md")
report_path.write_text(render_markdown(output), encoding="utf-8")
_atomic_write_text(report_path, render_markdown(output))
return output


Expand Down
Loading
Loading