Skip to content
THIRI logo
build.thiri.ai
Music theory for agents

Manual · Part III · Building on it

How we measure

The six disciplines behind every number THIRI publishes, from the bit-identity gate to measuring the worst block, and how they transfer from a pedal to a film score to an API.

Updated Oct 2, 2026

Every number on this site is supposed to survive disbelief. That is a method, not a mood, and it was written down in the open while the Elk Stomp measurements were being taken. The six disciplines below are from that public write-up. They apply to a plugin on a pedal, a score against picture, and an API response alike.

1. The bit-identity gate

A change is either provably output-identical or it is a change. Render a fixed input offline, hash the output, and compare it to a committed baseline. If the hash matches, the optimisation changed nothing you can hear and the old measurements still stand. If it does not, you have a new build and every number for it starts at zero.

On the pedal this is how PSOLA replaced the phase vocoder without the voicing engine’s output moving. In the studio it is how a snare’s tuning knob was proven byte-identical at its default before the knob was allowed to exist. For the API it is the reason golden response files are safe: no sampling, so a snapshot taken today is still correct until the changelog says the grid changed.

2. Void the rows

When you find a confound, every measurement taken before the fix is void. Do not patch the numbers; re-sweep. It feels expensive and it is cheaper than one wrong decision made on a number that was quietly contaminated.

The data file in the public measurement repo keeps the voided rows on purpose. Reading it top to bottom is the real story, failures included.

3. The confound checklist

Before every sweep, answer the same questions and record the answers per row: is anything else competing for the CPU; what governor is the processor on; which core is the audio thread pinned to and at what priority; does the deployed binary actually carry the vector instructions you think it does; are denormals flushed. On the Stomp, an unpinned control app took the worst block from 0.770 to 2.273 of the deadline while the average sat at a reassuring 0.60. The checklist is what catches that before it reaches a slide.

4. Hash, don’t timestamp

The board has no real-time clock and reports a year that is wrong. File modification times are meaningless there. So every row records the SHA-256 of the deployed binary, which is the only reliable statement of what produced a number. The same rule reaches the film work: a delivered mix is identified by its hash, and the picture stream is checked bit-identical against the original before the mux is trusted.

5. Generate figures from the data

The figure script reads the data file, asserts the claims the prose makes, and only then draws. A figure that cannot be regenerated from the file is a drawing, not evidence. Doing it this way caught three errors that had already been published in slides, which is the point.

6. The average lies. Measure the worst block.

Real-time audio fails on its worst block, not its average. Sushi has no dropout counter, so a block maximum at or over the deadline is the proxy. Every headline failure in the Elk work was invisible to average CPU. The same instinct carries to picture: a mix can sit at a healthy loudness while one foley hit lands 30 ms late, so each sound is measured against its own frame and the worst one is the number that gets reported.

What this looks like in practice

  • Units are stated and labelled. The Stomp numbers are fractions of the 1.333 ms block, and a fraction is never printed with “ms” after it.
  • A claim names its row. “0.592 average at five voices” points at a specific line in a specific file with a specific binary hash.
  • A disagreement between two sources is written down as a disagreement, not resolved by picking the nicer number. Where this manual says “unverified”, that is the method speaking.
  • A measurement is not a recommendation. The first phase-vocoder build looked fine on average and was never going to ship; the data said so before anyone tuned it.

Where to see the receipts

The harness, the 52-row data file and the generated figures are in the public measurement repository, under an MIT licence, so you can run the same discipline on your own DSP and disbelieve us properly. The scoring work’s verify script and level tables follow the same rules; the studio’s listening tools measure a render before anyone claims it sounds right.

Next

The Builders playbook →

How T.H.I.R.I. Builders runs week to week, what the paid tiers include and how they reach your key, how to bring a build or a bug so it gets fixed, where each kind of writing is published, and the six steps of the Get Started Guide.