Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@
* Added stems conversion: to mix several recordings into one reconstruction.
* Changed drive to reach for louder instructions while a recording converts; reconvert anything converted at a drive other than `1.00`.
* Improved clarity of reconstructions.
* Fixed the lowest triangle notes wobbling in pitch after a conversion; libraries are built again for it.
* Optimized the size of reconstructions.
* Bumped the reconstruction data-version to `2.2` with backward compatibility for `2.1`.
* Bumped the library data-version to `2.1`.
Expand Down
40 changes: 30 additions & 10 deletions docs/concepts/reconstruction.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,7 +96,7 @@ rate:
|----------|---------------------------|-----------------------------------------|---------------------------|
| `fft` | linear | uniform, `Δf ≈ sample_rate / N ≈ 27 Hz` | one short window (~37 ms) |
| `logfft` | logarithmic, floored at `Δf` | the FFT's `Δf`, on a musical axis | one short window (~37 ms) |
| `cqt` | logarithmic (constant-Q) | constant *relative* (fine low end) | long for low notes (~300 ms) |
| `cqt` | logarithmic (constant-Q) | constant *relative* (fine low end) | long for low notes (~600 ms) |

They sit at different points of the **time–frequency trade-off** (the Gabor limit:
sharper frequency resolution requires a longer time window, and vice versa):
Expand All @@ -113,7 +113,17 @@ sharper frequency resolution requires a longer time window, and vice versa):
sharply in time.
- **CQT** (constant-Q transform) places bins geometrically and gives
every musical interval the same number of bins, so it resolves low pitches finely.
It is the default.
It is the default. Its lowest bin sits at the lowest note the chip sounds: the
triangle's, an octave below the pulse's, since the triangle steps through its wave
at half the pulse's rate. A bass line on the triangle is therefore read from its
fundamental, which is where the pitch of a bent note is read from too.
This floor is a measured choice. Each bin's wavelet depends on its own frequency alone,
so every bin above the pulse's lowest note is the same at either floor, and so is what
the conversion makes of the music there. The lower floor adds the triangle's lowest
octave. Starting at the pulse's lowest note (54.6 Hz), a pulse took a triangle sliding
from 37 Hz at its third harmonic, and a 37 Hz bass under a melody went unplayed (44.1 kHz
audio at a 60 Hz frame rate, converted onto pulse 1, the triangle and noise). The lower
floor costs about a tenth more conversion time and a larger library.
The price is time support: its low-frequency basis functions are long (hundreds of
milliseconds), so brief events are smeared in time at the low end. _SampleToNES_
computes the CQT **once over the whole signal** with a hop of one frame, so each
Expand Down Expand Up @@ -323,13 +333,19 @@ and the spectrum discards. A partial standing between two bin centers still adva
rate. Comparing that advance across two columns against the rate the bin itself turns at gives the
partial's frequency far more finely than the bins are spaced. The reading takes the first few harmonics of
the note the decoder chose. It weights each by the energy behind it and settles each against the
fundamental the harmonics below it agreed on. This places the note **within a tenth of a cent** across the
whole range.
fundamental the harmonics below it agreed on. A harmonic counts where its partial stands within half a
semitone of where that fundamental puts it, the room a note owns, so a partial another voice sounds a bin
away stays out of the reading. The advance across two columns repeats every `sample_rate / hop` hertz
(60 Hz at the defaults), which from about 1 kHz up is narrower than a note's room. There a second reading,
of the advance over a sixteenth of a hop, names the repeat the first harmonic stands on, and the advance
across two columns keeps its precision inside it. This places the note **within a tenth of a cent** across
the whole range.

The reading also says how much of the frame stands behind it: the share of the column's energy its
harmonics hold. A pitched frame reads around 0.5, a frame sharing the channel with another tone around
0.3, and noise around 0.04. One threshold therefore separates the frames worth bending from the frames
with no pitch to read.
harmonics hold, with every bin measured on the scale the features use. On that scale a bass takes the
share its level gives it in every register, so a melody over a low bass keeps its reading. A pitched frame
reads around 0.6, a frame sharing the channel with another tone around 0.3, and noise around 0.02. One
threshold therefore separates the frames worth bending from the frames with no pitch to read.

### 6.2 Landing the note, and holding it

Expand All @@ -344,13 +360,14 @@ chases. The per-frame proposals are therefore settled by a change-penalized walk
Viterbi decoder uses to settle a note contour. The cost of a bend is how far it is from that frame's
reading, plus a toll on changing at all. The states a frame may take are the bends its neighborhood
proposed, together with no bend. That keeps the walk to a handful of states even where a note owns tens of
dividers.
dividers. A bend counts divider steps from its own note, so each note's frames are settled on their own: a
frame that reads nothing keeps a bend its own note read, and a new note starts from its own reading.

### 6.3 What it costs, and what it leaves alone

The refinement enumerates no candidate and rescores nothing. It leaves the library, the per-frame matching
and the decoder's lattice exactly as they were. It adds one transform per recording and a small walk per
channel.
and the decoder's lattice exactly as they were. It adds one transform per recording, two short ones over
the bins from about 1 kHz up for the second reading, and a small walk per channel.

The transform's cost depends on the machine. On a CUDA build it is too small to measure. On a CPU build it
is a tenth or more of a short conversion, because the reading needs a handful of bins per frame and the
Expand Down Expand Up @@ -394,6 +411,9 @@ be shown and played on a common scale.
and the coefficient is one global scalar. Material whose *useful* content spans a wider range than that
cannot be fully captured. A long crescendo and a very quiet passage under a loud one are examples.
Content far below the working level falls under the quietest playable note and is rendered as silence.
- **The triangle's fixed level.** The triangle plays at one volume. A bass a few decibels quieter than that
level is left out, and the calibration referees score the result closer to the recording than the same
conversion with the bass played too loud. A bass near that level is played.
- **CQT time resolution.** Constant-Q analysis needs long windows at low frequencies, so low-pitched
transients are smeared in time under `cqt`. `fft` and `logfft` localize time better at the cost of
low-frequency resolution.
Expand Down
25 changes: 21 additions & 4 deletions docs/development/bugs-and-todos.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,10 +48,18 @@ dimension the import starts carrying.
which is a tenth or more of a short conversion on a CPU build.
* Keeping the recordings a stopped folder scan has found, so stopping a long walk keeps the count the reader
watched climb.
* Calibrating the pitch refinement. The settings that decide how a pitch reading bends a note (a
confidence threshold, a change weight and a window) are chosen by hand, and
[the calibration](../tools/calibration.md) could measure them. The change weight trades vibrato against
jitter.
* Calibrating the pitch refinement's change weight and window. The confidence threshold was measured on
bending probes; the change weight and the window are chosen by hand, and
[the calibration](../tools/calibration.md) could measure them. The change weight counts divider steps,
which span about 2 cents at 110 Hz and about 27 cents at 1760 Hz. One weight therefore lets noise played
as low notes change its bend often while it holds high vibrato back. Counting it in cents weighs every
register alike, and its value then needs choosing again.
* A setting for the lowest frequency the analysis reads. The analysis starts at the triangle's lowest
note, about 27.3 Hz, so every note the chip plays is read from its fundamental, and its longest window
spans over half a second. Music that stays above the triangle's lowest octave would convert about a
tenth faster with the floor an octave higher. In lower music, a pulse then plays those notes at their
third harmonic. A floor below the triangle's lowest note needs longer library samples, since each must
last three windows.

### Technical

Expand All @@ -67,6 +75,9 @@ dimension the import starts carrying.
project sample. An edit to such a document is undoable nowhere ([undo](application/undo.md)), so an edit
that silences a channel or lets a recording go is reversible only by reloading the file.
* Improve performance of the browser's favorite scan of the entire tree per click
* The `fft` and `logfft` analysis floors. Their window spans two cycles of the pulse's lowest note, so the
triangle's lowest octave reaches them by its harmonics alone. Covering it doubles their window and
coarsens their timing.

## Architecture

Expand Down Expand Up @@ -98,6 +109,12 @@ currently out of line. An entry leaves when the code meets the contract again.

## Bugs

* A silent triangle renders the middle of its wave, where the console holds the step it stopped on. A
faithful hold needs a DC-blocking output stage at every mix (the reconstruction's render, the sequencer
and the song render), since nothing drains a held level today.
* Another voice inside a harmonic's bin pulls the pitch reading. A partial of another voice within half a
semitone of one of a note's harmonics shares that harmonic's bin and adds to its reading: a 65 Hz bass
under a steady 330 Hz pulse reads about 4 cents off, and within a cent alone.
* The Sample column shows no sample on a frame's first rows, since its reading starts over at each frame. It
offers no transpose or volume there, while playback applies them to the sample the previous frame left
sounding.
9 changes: 5 additions & 4 deletions docs/formats/instruction-libraries.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,9 @@ Each entry contains:
* **pitch** (33–119) for pulse and triangle, or **period** (0–15) for noise;
* **volume** (0–15) for pulse and noise;
* **duty_cycle** (0–3) for pulse, or the **short** (0–1) flag for noise;
* **waveform** — one full period of the rendered wave (the longest noise samples
are trimmed to one second);
* **waveform** — one full period of the rendered wave. The longest noise samples
are trimmed to two seconds, which holds the whole stretch a constant-Q spectrum
is read over;
* **spectrum** — the waveform's precomputed frequency content.

### Configuration key
Expand All @@ -40,14 +41,14 @@ A file holds a deflated [MessagePack](https://msgpack.org/) payload, with the fr
[Reconstructions](reconstructions.md#storage-and-export), and has its configuration in the file name:

```
sr_44100_nf_60_ws_13579_tg_0_sm_cqt_ch_384e710987cb958adf2b214df1267d10.ins
sr_44100_nf_60_ws_27157_tg_0_sm_cqt_ch_384e710987cb958adf2b214df1267d10.ins
```

| Fragment | Meaning |
| --- | --- |
| `sr_44100` | sample rate 44100 Hz |
| `nf_60` | NES frequency 60 Hz |
| `ws_13579` | FFT window size (samples) |
| `ws_27157` | FFT window size (samples) |
| `tg_0` | transformation gamma 0 |
| `sm_cqt` | spectrum method (`fft` / `logfft` / `cqt`) |
| `ch_384e…` | a hash of the library configuration section |
Expand Down
6 changes: 4 additions & 2 deletions docs/glossary.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,15 +20,17 @@ and the classes that implement it.
### Pulse (square)

A channel that plays a square wave. Its duty cycle is selectable and it has 15 volume levels. The chip
has two independent pulse channels: `pulse1` and `pulse2`.
has two independent pulse channels: `pulse1` and `pulse2`. The chip silences a pulse whose
[divider](#divider) is below 8, so a bend that goes higher than that at the top notes goes quiet.

### Triangle

A channel that plays a triangle wave of fixed shape and volume. Only its pitch varies. Its timer divides
the APU clock by 32 where the pulse timers divide by 16, and all three read the same period table. A
triangle note therefore sounds an octave below the pulse note with the same period, so a triangle
instruction of pitch P sounds at pitch P−12. FamiTracker uses the same convention, so an exported note
plays at the pitch _SampleToNES_ played it.
plays at the pitch _SampleToNES_ played it. Below [divider](#divider) 2 the chip's triangle steps above
hearing, and its output rests at the middle of the wave.

### Noise

Expand Down
20 changes: 18 additions & 2 deletions src/sampletones_application/logic/instruction/library.py
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,7 @@ def __init__(
self._library_manager = library_manager
self._is_operation_active = is_operation_active
self._eta_estimator: Optional[ETAEstimator] = None
self._replaced_library_path: Optional[Path] = None

self._lock_function: Optional[VoidCallback] = None
self._unlock_function: Optional[VoidCallback] = None
Expand Down Expand Up @@ -278,14 +279,19 @@ def rebuild_library(self, library_key: InstructionLibraryKey) -> None:
it was built for, the exclusive-operation gate permitting.

Those settings become the configuration's before the generation starts, which writes the
library in the place of the one it replaces.
library in the place of the one it replaces. Where this build names that library's file
differently, the replaced file is removed once the new one is written.
"""
if self._is_operation_active():
logger.warning("A conversion or library generation is already in progress")
return

self.call(self.on_apply_library_config, library_key, self._library_manager.stored_config(library_key))
replaced = self._library_manager.get_path(library_key)
rebuilt = self._library_manager.get_path(self._config_manager.key)
self.generate_library()
if self._library_manager.is_generating() and replaced != rebuilt:
self._replaced_library_path = replaced

def generate_library(self) -> None:
if self._library_manager.is_generating():
Expand Down Expand Up @@ -511,21 +517,31 @@ def _update_progress_state(self, task_progress: TaskProgress) -> None:
)

def _on_generation_completed(self) -> None:
"""Closes the generation and reads the catalog again, which lists the library it wrote."""
"""Closes the generation, removes the file a rebuild replaced under another name, and reads
the catalog again, which lists the library it wrote."""
self.call(self.on_generation_completed)
self._close_generation()
self._remove_replaced_library()
self.refresh_libraries(load_if_needed=False)

def _on_generation_error(self, exception: Exception) -> None:
self.call(self.on_generation_error, exception)
self._close_generation()
self._replaced_library_path = None
self.update_status()

def _on_generation_canceled(self) -> None:
self.call(self.on_generation_canceled)
self._close_generation()
self._replaced_library_path = None
self.update_status()

def _remove_replaced_library(self) -> None:
"""Removes the file a rebuild replaced under another name, where it still stands."""
replaced, self._replaced_library_path = self._replaced_library_path, None
if replaced is not None and replaced.is_file():
remove_path(replaced)

def _close_generation(self) -> None:
"""Lets the creator go along with the tree lock the generation took when it was asked for,
which loading a library and rebuilding the tree both yield to."""
Expand Down
2 changes: 1 addition & 1 deletion src/sampletones_core/configs/generation.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,6 @@ decoder:
on_off_weight: 0.2

refinement:
confidence: 0.15
confidence: 0.10
change_weight: 2.0
window: 4
2 changes: 1 addition & 1 deletion src/sampletones_core/constants/algorithm.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@
# Library creation

MIN_SAMPLE_LENGTH: Final[float] = 0.05
MAX_SAMPLE_LENGTH: Final[float] = 1.0
MAX_SAMPLE_LENGTH: Final[float] = 2.0
LIBRARY_PHASES_PER_SAMPLE: Final[int] = 100

# Calculation methods
Expand Down
5 changes: 5 additions & 0 deletions src/sampletones_core/constants/general.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,13 +7,18 @@
APU_CLOCK: Final[float] = 1789773.0
TIMER_CYCLE_DIVIDER: Final[int] = 16
MIN_TIMER: Final[int] = 1
MIN_SOUNDING_PULSE_TIMER: Final[int] = 8
MIN_SOUNDING_TRIANGLE_TIMER: Final[int] = 2
MAX_TIMER: Final[int] = 0x7FF
MIN_PITCH: Final[int] = 33
MAX_PITCH: Final[int] = 119
MIN_PLAYED_PITCH: Final[int] = LIMIT_MIN_PITCH
PITCH_RANGE: Final[int] = MAX_PITCH - MIN_PITCH

TRIANGLE_PHASE_INCREMENT: Final[float] = 0.5

MIN_FREQUENCY: Final[float] = APU_CLOCK / (TIMER_CYCLE_DIVIDER * (MAX_TIMER + 1))
MIN_TRIANGLE_FREQUENCY: Final[float] = MIN_FREQUENCY * TRIANGLE_PHASE_INCREMENT
MAX_FREQUENCY: Final[float] = APU_CLOCK / TIMER_CYCLE_DIVIDER

NOTE_NAMES: Tuple[str, ...] = (
Expand Down
5 changes: 3 additions & 2 deletions src/sampletones_core/constants/spectrum.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,11 @@

from sampletones_shared.constants.music import OCTAVE_SEMITONES

from .general import MIN_FREQUENCY, QUIETEST_VOLUME_LEVEL
from .general import MIN_FREQUENCY, MIN_TRIANGLE_FREQUENCY, QUIETEST_VOLUME_LEVEL

BINS_PER_OCTAVE: Final[int] = OCTAVE_SEMITONES
CQT_CUTOFF_FREQUENCY: Final[float] = MIN_FREQUENCY
CQT_CUTOFF_FREQUENCY: Final[float] = MIN_TRIANGLE_FREQUENCY
LOG_FFT_CUTOFF_FREQUENCY: Final[float] = MIN_FREQUENCY
SPECTRUM_FLOOR: Final[float] = QUIETEST_VOLUME_LEVEL**2

CQT_REFERENCE_CONTEXT_FACTOR: Final[int] = 3
Expand Down
2 changes: 1 addition & 1 deletion src/sampletones_core/data/document.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
from sampletones_shared.types.path import Pathlike

DOCUMENT_MAGIC: Final[bytes] = b"\x1f\x8b"
COMPRESSION_LEVEL: Final[int] = 9
COMPRESSION_LEVEL: Final[int] = 6
STATED_TIMESTAMP: Final[int] = 0


Expand Down
41 changes: 31 additions & 10 deletions src/sampletones_core/fft/cqt/transform.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
from typing import Optional
from typing import Final, Optional

import numpy as np

Expand All @@ -13,29 +13,50 @@
from ..utils import calculate_n_bins
from .kernel import CQTKernel, build_cqt_kernel

MAX_GATHERED_SAMPLES: Final[int] = 1 << 24

def _framed_signal(audio: np.ndarray, frame_length: int, hop_length: int) -> Array:
"""Stack centered frames of ``audio`` on the compute device, one column per hop.

``audio`` is zero-padded by half a frame on each side so column ``t`` is centered on sample
``t * hop_length``; the number of columns is ``1 + len(audio) // hop_length``.
"""
def _padded_signal(audio: np.ndarray, frame_length: int) -> Array:
"""``audio`` on the compute device, zero-padded by half a frame on each side, which centers column
``t`` on sample ``t * hop_length``."""
left = frame_length // 2
right = frame_length - left
device_audio: Array = xp.asarray(audio, dtype=xp.complex64)
padded: Array = xp.concatenate(
[xp.zeros(left, dtype=xp.complex64), device_audio, xp.zeros(right, dtype=xp.complex64)]
)
frame_count = 1 + len(audio) // hop_length
starts = xp.arange(frame_count) * hop_length
return padded


def _framed_columns(
padded: Array,
frame_length: int,
hop_length: int,
first: int,
count: int,
) -> Array:
"""Stack ``count`` centered frames of the padded signal, one column per hop from column ``first``."""
starts = (xp.arange(count) + first) * hop_length
indices = starts[:, None] + xp.arange(frame_length)[None, :]
framed: Array = padded[indices].T
return framed


def _transform(audio: np.ndarray, kernel: CQTKernel, hop_length: int) -> np.ndarray:
frames = _framed_signal(audio, kernel.frame_length, hop_length)
coefficients: Array = kernel.matrix @ frames
"""Correlates the kernel with every frame of ``audio``, ``1 + len(audio) // hop_length`` columns.

Every frame spans the lowest bin's whole wavelet, so the frames are gathered a block of columns
at a time, each block holding at most ``MAX_GATHERED_SAMPLES`` samples. That keeps the memory a
recording takes to one block, however long the recording runs.
"""
padded = _padded_signal(audio, kernel.frame_length)
frame_count = 1 + len(audio) // hop_length
block = max(1, MAX_GATHERED_SAMPLES // kernel.frame_length)
columns = [
kernel.matrix @ _framed_columns(padded, kernel.frame_length, hop_length, first, min(block, frame_count - first))
for first in range(0, frame_count, block)
]
coefficients: Array = xp.concatenate(columns, axis=1)
return to_numpy(coefficients)


Expand Down
Loading
Loading