Skip to content
THIRI logo
build.thiri.ai
Music theory for agents

Manual · Part II · In production

Google Lyria, Magenta and Gemini: what THIRI can steer

The 2026-10-01 experiments: which Google music models exist, what THIRI can enforce in each, what was measured, the Lyria scale value that does nothing, and a melody generator with no model that beat Magenta on the study cue.

Updated Oct 2, 2026

For each Google music model the question is narrow: can THIRI’s output reach it as a constraint, and does the result follow? This chapter records the 2026-10-01 study, run on the carving-film cue from Scored to picture.

Why measure instead of trust

An earlier falsification on Stable Audio 3 found that text chords scored 2 of 12 and that audio seeding was bleed, not conditioning. The rule from it: an audio model is not “THIRI-steered” until its chord-following beats a control with no conditioning. Steering has to be symbolic, or it has to be measured.

What each model can hold

Ranked by what THIRI can enforce, from Google’s own documentation as read on 2026-10-01.

LevelModelWhat reaches it
NotesMagenta RealTime 2, released 2026-06-04A mask over 128 pitches every 40 ms. No key, tempo or chord input.
NotesGemini 3.1 ProNotes as text or structured data, rendered by Csound.
NotesCoconet, coconet/bachKeeps fixed voices, fills the rest in four parts.
ChordsMagenta.js 1.23.1: ImprovRNN, MusicVAE mel_chords, two multitrack modelsA chord symbol per step.
Key and tempoLyria RealTimeA scale enum of twelve major and relative-minor sets; bpm, beats per minute, 60 to 200.
Key and tempoLyria 3.5, Lyria 3 Pro and ClipPrompt text only.
NothingLyria 2, Magenta RealTime 1, MusicFX, Dream TrackNo musical input THIRI can hold.

Lyria RealTime’s scale enum lines up with THIRI’s parent-major key context: G Phrygian is E_FLAT_MAJOR_C_MINOR, B♭ Lydian is F_MAJOR_D_MINOR. It sets the seven notes, not the home note. Changing scale or tempo needs a context reset, a hard cut with up to two seconds of lag. Google’s own page calls its key and tempo “imprecise”.

What ran on this machine

An Intel Core i5 Mac on macOS 13. Magenta RealTime 2 cannot install there: magenta-rt 2.0.3 needs JAX 0.9 or newer, and jaxlib’s last Intel-Mac wheel is 0.4.38. Magenta.js runs under bun. Lyria RealTime is remote. Gemini ran headless through Antigravity.

Method

One cue for everything: Gm7 | E♭maj7 | D7♯9 | Cm7, twice, at 120 beats per minute, in G Phrygian. THIRI’s analyze_chord and voice-led generate_voicing output is the shared input. One scorer labels every note as chord tone, THIRI chord-scale tone, or outside. One renderer reuses the carving film’s theremin and strings. Each model gets controls: no chords, the wrong key, or a tritone-shifted progression.

Gemini: colour, not correctness

Gemini 3.1 Pro, three runs in each of three conditions. Without THIRI it wrote 82 to 85 percent chord tones and hit a chord tone on every beat 1 and 3. With THIRI’s analysis in the prompt, 86 to 90 percent and no outside notes. With THIRI as a Model Context Protocol tool the model could call, 71 to 90 percent at about twice the tokens and time. The memory records the top run as 91; the report rounds the same 90.5 percent share to 90.

That is a ceiling effect: the model already knows the chord tones. Over D7♯9 the THIRI runs used E♭, B and A♭ from the half-whole diminished scale the engine named, where the plain runs used B♭. THIRI changed the colour, not the correctness.

Magenta.js: conditioning that works

Ten takes per condition under bun. MusicVAE mel_chords landed 82 percent chord tones on THIRI’s chords, 35 percent with no chords, 14 percent with a tritone shift. ImprovRNN: 73, 52 and 26 percent, p at or below 0.003. All 60 takes took 73 seconds on the i5. The controls show the conditioning is the mechanism. One limit: mel_chords reads only root and triad, so D7♯9 arrives as D major.

Lyria RealTime: eleven scales work, one does nothing

Twenty-one sessions of 32 seconds. Eleven of the twelve scale values put 76 to 83 percent of the audio’s energy in their own seven notes, up to 47 points more than Lyria’s own pick for the prompt. The wrong set, A major, dropped G Phrygian’s share from 92 to 36 percent. Tempo held: 119.7 on most takes for a request of 120. Chords in the prompt text were not followed: the D7♯9 needs F♯ and A, which got 2 percent of the energy; 55 percent sat in key overall.

E_FLAT_MAJOR_C_MINOR is a no-op. In four of four same-seed pairs in the study it returned audio bit-identical to SCALE_UNSPECIFIED, including a bright prompt whose default pick held A naturals. The client library sends the value correctly, so the evidence points at the server. That remains unverified until Google answers. A standalone repro script added a fifth identical pair, which is the count the drafted bug report carries. The value is exactly THIRI’s G Phrygian set. The workaround for G-minor cues is B_FLAT_MAJOR_G_MINOR, which shares six of the seven notes.

Exp4: a melody generator with no model

Exp4 asks whether THIRI alone can write a better line than Magenta: rules plus a seeded beam search over THIRI’s analysis. The per-bar palette is THIRI’s primary chord-scale intersected with the key; a borrowed chord keeps the engine’s own scale, so D7♯9 gets half-whole diminished. Strong beats land on chord tones, guide tones preferred. Non-chord tones resolve by step and leaps are recovered. The form is a four-plus-four period with bars 5 and 6 restating bars 1 and 2. A hit marks the step that must carry the peak. No training, no weights.

Ten takes, scored against the real chords: 86 percent chord tones, 0 percent outside, 100 percent of strong beats on chord tones. Steps 51 percent, leap recovery 100 percent, motif recurrence 56 percent. All ten takes peaked on step 96, on F♯5 then F5, the ♯9. Controls: key only 58 percent, no chords 38, tritone-shifted 15, p 0.00005. Each take composes in 0.12 seconds.

Same cue, same scorer: Magenta mel_chords 82 percent chord tones, 5 percent outside, 22 percent steps, no motif recurrence, a 39-semitone range; Gemini with THIRI in the prompt 88 percent, 45 percent steps, 60 percent motif recurrence. Exp4 beat Magenta on chord tones, outside notes, strong beats, steps, leap recovery and motif recurrence, and tied Gemini on correctness. The one measure it lost, semitone rubs against the pad, counts doubling a pad note as a rub.

Exp5: a Lyria bed under THIRI’s arrangement

A percussion-and-texture bed, bass muted, scale B_FLAT_MAJOR_G_MINOR, replaced the carving film’s drum loop under the delivered stems. The take came out at exactly 120.0 beats per minute, sits on B♭ and D over the chords, and is levelled section by section to the old drum stem’s energy: 18 to 27 dB under the music by the memory’s reading, with the build report recording applied gains of 18 to 29 dB. The mix measures −16 LUFS, loudness units relative to full scale; the video stream is bit-identical.

The timing number changed; the change is the lesson. The first reading put the bed’s strong onsets within 3.3 ms median of the film’s beat grid. It came from the spectral-flux envelope that had placed the bed, which, calibrated on clicks and drum hits, leads audible onsets by about 13.5 ms. Measured against ground truth on 2026-10-02, the bed plays 14.5 ms late as a signed median. Measure against ground truth, never against your own placement envelope. Also: the same seed, recorded again on 2026-10-02 in the film studio, came out at 120.97 for 120, so a seed does not pin the tempo; stretch to lock. And pad the bed’s front with silence, since Lyria’s downbeat can land before the bar line.

Caveats

  • One cue. Every number above is for one eight-bar progression at one tempo.
  • Exp4 wins partly by construction: its rules are what the musicality measures reward. Its lines are plainer, about 10 semitones of range and 6.5 pitch classes a take.
  • Gemini was near ceiling before THIRI arrived; a harder cue might separate the conditions.
  • The Lyria no-op is five of five with one client; Google’s confirmation would close it.
  • THIRI’s own conduct_band ignored an explicit mode: asked for G Phrygian, it returned a G major ii–V–I. Steer with analyze_chord, generate_voicing and resolve_chord.

What is open

  • The blind listen of Exp4 against Magenta and Gemini.
  • More progressions for Exp4, then range and tension tuned by ear, before it becomes a THIRI tool.
  • The Lyria bug report, drafted for the Google AI Developers Forum, unposted.
  • Exp5 shifted 14.5 ms and rebuilt.
  • Magenta RealTime 2, the only model that takes THIRI’s exact pitches, untested for want of a machine.

Next

Composer C: an agent that writes Csound with THIRI holding the harmony →

A chat-driven studio where a language model edits a typed song spec, THIRI supplies every voicing, Csound renders deterministically, and nothing is called finished until it has been measured.