Manual · Part II · In production
On the pedal: fitting the engine in 1.333 ms
The Elk Stomp measurements: five harmony voices plus chord intelligence inside the audio callback, the three builds it took to get there, and the six disciplines that made the numbers trustworthy.
Updated Oct 2, 2026
The question was empirical: does a deterministic harmony engine plus five voices of pitch shifting fit inside a 1.333 millisecond audio callback on a dual-core Cortex-A7? The hardware is Elk Audio’s Stomp: STM32MP157, two cores at 400 to 800 MHz, 48 kHz, a 64-sample buffer, Elk Audio OS 1.2.2, Sushi 1.3.0. The plugin is HORN SXTN, the THIRI VST, running headless.
The answer
With 1.0 meaning the whole deadline, measured on the board, three sweeps of six points at 60 seconds each:
| Voices | Average | Worst block |
|---|---|---|
| 0 (engine and tracker only) | 0.568 | 0.726 |
| 1 | 0.577 | 0.749 |
| 3 | 0.588 | 0.757 |
| 5 | 0.592 | 0.788 |
The fifth voice costs almost exactly what the first did: 0.005 of the block per voice. The finding is the flatness, not the headroom.
Three builds, one CSV
The data file is append-only and keeps the failures. These are the five-voice rows for each build in that file (the README prints 5.94 for the first build from a different row). Read top to bottom:
| Build | Average at 5 voices | Worst block | Verdict |
|---|---|---|---|
| Phase vocoder, unsliced tracker | 2.08 | 5.99 | one block in eight six times over deadline |
| Phase vocoder, tau-sliced tracker | 1.83 | 2.93 | over real time at one voice |
| PSOLA, tau-sliced tracker | 0.59 | 0.79 | ships |
Three lessons came out of that table and they apply to any real-time DSP, not only ours.
A worst-case spike is not a CPU problem. The first build looked healthy on average while a single O(n²) analysis fired once per hop and landed in one block. Spreading the identical arithmetic across the hop’s eight blocks cut the worst block eightfold with no change in output.
No buffer size fixes an average. The phase vocoder exceeded real time at one voice. Doubling the buffer doubles the deadline and the work; the fraction stays put.
The cheap algorithm won by 52×. PSOLA costs 0.005 per voice, the phase vocoder 0.254. The board was never the constraint; the algorithm was.
Two findings worth stealing
Pin the control surface. This one is a single paired session recorded in the README and the port map, not a sweep in the data file; the clean baseline across sweeps is 0.74 to 0.83 worst block. The audio thread lives on CPU 1. An unpinned hardware daemon and control app landed on it and took the worst block from 0.770 to 2.273 while the average sat at a reassuring 0.60. taskset -c 0 put it back to 0.760. Elk’s reference launcher does not pin.
NEON is not free on an A7. The expected four-times gain on the tracker’s multiply-accumulate loop measured as 10.7 cycles per vector operation in accumulation chains, about scalar VFP speed. GCC will not auto-vectorise float loops without -funsafe-math-optimizations. Budget from measurement.
The six disciplines
- The bit-identity gate. Render a fixed input offline, hash it, compare to a committed baseline. A change is either provably output-identical or it is a change.
- Void the rows. A found confound voids every row before it. Re-sweep.
- The confound checklist. Competing load, governor, affinity and priority, NEON attributes in the deployed binary, denormals. Run before every sweep; record per row.
- Hash, don’t timestamp. The board has no real-time clock. Every row carries the SHA-256 of the deployed binary.
- Figures from data. The figure script asserts the claims before it draws. It caught three errors already published in slides.
- The average lies. Sushi has no xrun counter; a block maximum at or over 1.0 is the proxy.
How the engine got there
The C++ core, thiri_core, is a port of the JavaScript engine, not a fork, pinned to it by golden vectors generated from the JS: 16,353 in the published engine’s README, with larger sets on the plugin repository’s current branches (the exact count differs by branch). The arpeggiator engine (Helix) is pinned separately, 18,080 of 18,080 vectors bit-for-bit, so the pedal and the web app play the identical arpeggio for the same chart, seed and macros. One trap worth recording: the engine must not normalise voice order, because the vectors reproduce the JavaScript’s own inconsistency on purpose; consumers sort.
Still open
The live-input Sushi configuration on the main branch still selects the phase vocoder by default; only the file-source configuration selects PSOLA. A live horn through the harmonizer with harmony engaged is not yet on disk. That is the gate before any demo video, and the next measurement to add to the CSV.
Next
Google Lyria, Magenta and Gemini: what THIRI can steer →
The 2026-10-01 experiments: which Google music models exist, what THIRI can enforce in each, what was measured, the Lyria scale value that does nothing, and a melody generator with no model that beat Magenta on the study cue.