Skip to content
THIRI logo
build.thiri.ai
Music theory for agents

Case study · Elk Audio

THIRI VST on the Elk Stomp

Does a real-time music-intelligence layer fit inside a 1.333 millisecond audio callback on a few-hundred-pound pedal? It does. This is the harness, the data and the method, so you can disbelieve it properly.

Updated 2026-10-02

0.592

of the deadline, average, 5 voices

0.788

worst block, 5 voices

0.005

marginal cost per voice

52×

cheaper per voice than the phase vocoder

The constraint

Agentic music tools need two things they mostly do not have yet. A theory of correctness: a layer where "flat 9 over a dominant" is a thing that is either right or wrong, inspectable and constrainable before it becomes audio. And a real-time substrate: an agent that only runs in a cloud GPU is not in the loop with a player, it is a rendering service. If musical intelligence is going to sit in the signal path, it has to survive inside the audio callback, on the hardware people actually put on a stage.

The first is a design problem. The second is empirical, and it is the one this project answers. The engine is THIRI's C++ core, a port of the JavaScript engine pinned to it by golden test vectors (16,353 per the published engine's README; the plugin repository's current branches pin more, and the count differs by branch). The hardware is the Elk Audio Stomp: STM32MP157, two Cortex-A7 cores at 400 to 800 MHz, 32-bit ARMv7, 48 kHz, a 64-sample buffer, Elk Audio OS 1.2.2 and Sushi 1.3.0. The plugin, HORN SXTN, takes a horn or a voice in, tracks its pitch, asks THIRI for the voicing against the chart, and plays up to five formant-preserving harmony voices back out.

Signal flow on the Stomp: audio in, pitch tracker, THIRI voicing engine, five PSOLA voices, audio out
Signal flow. Figures are generated from the data file, never hand-typed.

The answer

Five pitch-shifted harmony voices, plus full chord intelligence and pitch tracking, on a stock Stomp. 1.0 is the 1.333 ms deadline. Every row was measured on the hardware: three sweeps of six points, 60 seconds per point.

VoicesAverageWorst block
0 (engine and tracker only)0.5680.726
10.5770.749
30.5880.757
50.5920.788

The fifth voice costs almost exactly what the first did: marginal cost is 0.005 of the block period per voice. That flatness is the finding, not the headroom.

Cost versus voices: average and worst block stay nearly flat from zero to five voices

Three regimes, visible in one CSV

The data file is append-only and contains every run, including the failures. Reading it top to bottom is the actual story. The rows below are the five-voice rows of each build in that file (the README prints 5.94 for the first build, which is a different row).

BuildAverage at 5 voicesWorst blockVerdict
Phase vocoder, unsliced tracker2.085.99Tracker spike, six times over deadline
Phase vocoder, tau-sliced tracker1.832.93Over real time at one voice (1.07 average)
PSOLA, tau-sliced tracker0.590.79Ships
  • A worst-case spike is not a CPU problem. The first build averaged 0.67 to 0.80, healthy by any CPU meter, while one block in eight ran 5.94 times over deadline, because an O(n²) analysis fired once per hop and landed in a single block. Spreading the identical arithmetic across the hop's eight blocks dropped the worst block eightfold with no change in results.
  • No buffer size fixes an average. The phase vocoder exceeds real time at one voice. Doubling the buffer doubles the deadline and the work; the fraction does not move.
  • The cheap algorithm won by 52×. Marginal cost per voice: PSOLA 0.005, phase vocoder 0.254, over the same zero-to-five span. The board was never the constraint. Choosing an algorithm that suited the constraint was.
Optimization journey: worst block falls from 5.94 to 0.79 across the three builds

Two findings that generalise beyond us

Pin your control surface, or it will eat your deadline silently. The audio thread is pinned to CPU 1 by the driver; CPU 0 sits idle. A hardware daemon and a control app, left unpinned, land on the audio core.

Worst block
Audio host alone0.770
Plus control surface, unpinned2.273 (missed deadlines)
Plus control surface, taskset -c 00.760

This A/B is one paired session on our rig, recorded in the README and the port map rather than in the sweep file; the board's clean baseline across the sweeps sits between 0.74 and 0.83 worst block. The average never moved, about 0.60 in all three cases. Watching a CPU meter you would have seen a comfortable 60% while audio dropped out. Elk's reference launcher does not pin, so this is worth checking in any Elk app.

Vectorising is not free money on an A7. We assumed NEON would give about four times on the tracker's multiply-accumulate loop. Measured with a known-instruction-count calibration: 10.7 cycles per vector operation in accumulation chains, about what scalar VFP does. GCC also will not auto-vectorise float loops without -funsafe-math-optimizations. Budget from measurement, not from the ISA datasheet.

The method is the contribution

  1. The bit-identity gate. A change is either provably output-identical or it is a change. Render a fixed input offline, hash it, compare to a committed baseline. Every optimisation passed it or was reverted.
  2. Void the rows. When you find a confound, every measurement taken before the fix is void. Re-sweep; do not patch numbers.
  3. The confound checklist. Competing load, CPU governor, affinity and real-time priority, NEON attributes in the deployed binary, denormals. Run it before every sweep, record the answers per row.
  4. Hash, don't timestamp. The board has no real-time clock and reports 2025. Every row records the SHA-256 of the deployed binary.
  5. Figures are generated from the data. The figure script reads the CSV, asserts the claims and emits the SVGs. That caught three errors that had already been published in slides.
  6. The average lies. Measure the worst block. Sushi has no xrun counter; a maximum at or over 1.0 is the proxy. Every headline failure in this project was invisible to average CPU.

The harness, the 52-row data file and the six disciplines in full are public:

reproduce it (needs an Elk board and Elk's desktop devkit)
tools/preflight.sh mind@<board-ip>            # confound checklist: all PASS/INFO
tools/sushi_board.sh start mind@<board-ip>    # SIGINT-only lifecycle
python3 tools/profile_sweep.py ...            # three sweeps of six points, 60 s each

github.com/BluesPrince/thiri-meets-elk (MIT). The engine is not in that repo and does not need to be; the measurement discipline is what transfers.

Context

This is our project on Elk Audio's hardware: not by, and not endorsed by, Elk Audio.

HORN SXTN took first place in the Elk Audio challenge at the MUTEK Montréal hackathon (22 to 23 August 2026) and second in Roland's Project LYDIA challenge. The winning demo ran the desktop plugin in Ableton Live; the board measurements above came in the days after, once Elk's SDK was in hand. The prize was a Stomp dev kit and mentoring. The relationship is a technical one: a working prototype, not a product.

What is still open

  • The live-input configuration on the main branch still selects the phase vocoder by default; only the file-source configuration selects PSOLA. A live-horn run through the harmonizer with harmony engaged has not yet been recorded on disk. That is the gate before a demo video.
  • No public download of the plugin exists yet. The Ableton install video on the home page shows the desktop path.
  • The Elk Stomp SDK is not public; reproduction needs Elk's devkit.

Next

The manual chapter on the pedal →

The same measurements with the method and the open questions.