ModelsArchitectures & capability
Claude Opus 5.5 invents a music notation and browser synthesizer to compose adventure-game tracks for Simon Willison's Scrimshaw Jukebox
From a two-sentence prompt, Claude Opus 5.5 designed a plain-text score format, wrote a Web Audio synthesizer to play it, and composed six original adventure-game tracks. Simon Willison judged the music surprisingly good and asked whether competent composition is a recent model capability.

The model never generates audio itself. It writes a text score, with settings, voices, patterns and an arrangement, plus a synthesizer that plays the score live in the browser. It offers more than 40 synthesized voices and uses no audio samples, and anyone can edit the score directly. [1] [4]
A follow-up pull request, co-authored by Opus 5.5 and Willison, adds a long prompt that teaches other models the Scrimshaw format, including counting rules and a worked example. That makes the setup a reusable template: have one model design a small domain-specific format, then hand its spec to others. It also makes cross-model comparisons easy to run, though none has been published. [10] [1]
Willison frames competent composition as a possibly new ability and says only careful experiments across models could confirm it. Composing in text notation has precedent: ChatMusician (2024) produced structured pieces in ABC notation after further training on music data. Scrimshaw differs because a general-purpose model invented its own notation and also wrote the player. [1] [3] [5] [6]
The output leaned heavily on the Secret of Monkey Island soundtrack named in the prompt, more than Willison intended. He guesses that training data included text renditions of game music from online forums, but that is unconfirmed. For commercial teams, naming a franchise in a prompt creates intellectual-property risk. [1] [2]
The evidence so far is one demo judged by one listener. If the ability generalizes, teams could prototype editable, version-controlled game audio without licensing samples, and the published format prompt makes cross-model comparisons cheap to run.
Executive brief
Simon Willison gave Claude Opus 5.5 a two-sentence prompt: invent a text format for music, build a player for it, and write game music "of the quality of the original Secret of Monkey Island." The model produced Scrimshaw Jukebox, a browser synthesizer with six original adventure-game tracks. Willison called the results "surprisingly good" (post). He also wondered whether competent composition is a capability text models gained only in the past few months. He says confirming that would take careful comparisons across models. No such comparison exists in the reviewed sources. For now this is one practitioner's demo, not a benchmark result.
What changed and event timeline
Music as a text language
ChatMusician continued pre-training LLaMA2 on ABC notation, a plain-text music format. It reported generating structured, full-length pieces that beat a GPT-4 baseline, plus a new music theory benchmark (;).
Opus 5.5 ships
Anthropic released Claude Opus 5.5 for long-running agentic coding and knowledge work. METR and Frontier Design tested it externally before release (;).
24: Music-video demos
Users showed Opus 5.5 building music videos from a user-chosen audio track, using JavaScript animation and outside APIs. The article is promotional and contains no independent evaluation ().
Scrimshaw Jukebox posted
Willison published the tool and the exact prompt he used. He noted the model "leaned a lot harder" into Monkey Island than he intended ().
LLM composing prompt added
PR #353, co-authored by Opus 5.5 and Willison, added a "Compose with an LLM" section. It contains a long prompt that teaches other models the Scrimshaw score format ().
Capabilities and access
- Model: Claude Opus 5.5 (
claude-opus-5-5family). It has a 1M-token context window and 128K maximum output. - Output: the model generates no audio itself.
- Access: the jukebox is public, and you can edit any score in the browser.
Read the full section
- Model: Claude Opus 5.5 (
claude-opus-5-5family). It has a 1M-token context window and 128K maximum output. It is available through Anthropic, AWS, Google Cloud and Azure (Let's Data Science). - Output: the model generates no audio itself. It writes a text score plus a Web Audio synthesizer that renders the score live, with no audio samples (tool).
- Access: the jukebox is public, and you can edit any score in the browser.
Technical analysis for researchers and developers
- Format design: the score lists settings, voice declarations, pattern blocks and an arrangement, in that order (tool).
- Rhythm: every bar has a fixed number of steps, and each dash extends a note by one step.
- Supported time signatures: 4/4, 6/8, 3/4 and 12/8, plus a swing setting.
Read the full section
- Format design: the score lists settings, voice declarations, pattern blocks and an arrangement, in that order (tool).
- Rhythm: every bar has a fixed number of steps, and each dash extends a note by one step.
- Supported time signatures: 4/4, 6/8, 3/4 and 12/8, plus a swing setting.
- Pattern tools: repeats (
B*2) and transpositions (A+5). - Instruments: more than 40 synthesized voices. Each voice has settings for volume, pan, reverb, envelope, glide and brightness.
- Reproducibility: the original prompt is published, and the format prompt in PR #353 gives any other model the same spec, including counting rules and a worked example. That makes cross-model tests straightforward to run, but nobody has published any.
Claims and evidence
- "Surprisingly good" output: this is Willison's personal judgment. No listener study or musicological review backs it (post).
- Composition is a new capability: Willison poses this as a question, not a finding (X post). Independent corroboration is not available.
- Opus 5.5 is cheaper and safer: these claims come from Anthropic.
Read the full section
- "Surprisingly good" output: this is Willison's personal judgment. No listener study or musicological review backs it (post).
- Composition is a new capability: Willison poses this as a question, not a finding (X post). Independent corroboration is not available.
- Opus 5.5 is cheaper and safer: these claims come from Anthropic. Reported price figures conflict:
- One report says it costs 40% less to run than Opus 5, priced at $4/$20 per million input/output tokens (Let's Data Science).
- Thurrott cites both the 40% figure and a 20% cut in token prices.
Context and prior work
Writing music as text is an established research direction. ChatMusician (2024) showed that a text-only LLM could compose in ABC notation once trained on a music corpus (repo). Scrimshaw differs in two ways: Willison compares this to text models recently getting competent at 3D graphics, which suggests a broader pattern of code-mediated creative output. The music-video demos fit the same pattern (Rad Neurons).
Read the full section
Writing music as text is an established research direction. ChatMusician (2024) showed that a text-only LLM could compose in ABC notation once trained on a music corpus (repo). Scrimshaw differs in two ways:
- A general-purpose model invented its own notation instead of using a standard one.
- The same model wrote both the score and the player.
Willison compares this to text models recently getting competent at 3D graphics, which suggests a broader pattern of code-mediated creative output. The music-video demos fit the same pattern (Rad Neurons).
Limitations, safety and contested findings
- Sample size: one prompt, one model and one evaluator. Willison says only "careful experiments" across models could show whether the ability is new (post).
- Imitation of copyrighted style: the output followed the named soundtrack closely. Willison asks how Claude "knows" what Monkey Island sounds like.
- Weak evidence elsewhere: the music-video coverage has no benchmarks and does not discuss limitations (Rad Neurons).
Read the full section
- Sample size: one prompt, one model and one evaluator. Willison says only "careful experiments" across models could show whether the ability is new (post).
- Imitation of copyrighted style: the output followed the named soundtrack closely. Willison asks how Claude "knows" what Monkey Island sounds like. His guess is that training data included text renditions of game music from online forums (X post). That is speculation, not documented fact.
- Weak evidence elsewhere: the music-video coverage has no benchmarks and does not discuss limitations (Rad Neurons).
Business and practitioner implications
- Prototyping: an AI-designed text format plus a synthesizer can produce editable, version-controllable game audio with no sample licensing.
- Pattern worth copying: have the model design a small domain-specific format (a DSL), then hand the format's spec to other models.
- IP risk: when prompts name an existing franchise, output can lean heavily toward it.
Read the full section
- Prototyping: an AI-designed text format plus a synthesizer can produce editable, version-controllable game audio with no sample licensing. That suits prototypes and jam builds.
- Pattern worth copying: have the model design a small domain-specific format (a DSL), then hand the format's spec to other models. PR #353 is a ready template.
- IP risk: when prompts name an existing franchise, output can lean heavily toward it. Teams shipping commercially should avoid reference-by-name prompts and review melodies before release.
- Evaluation gap: music needs human listening tests. Teams should build their own before relying on vendor or anecdotal claims.
The source trail.
Sources (10)
Simon Willison: Tool: Scrimshaw Jukebox
simonwillison.netOne thing I'd LOVE to understand is how Claude "knows" what the music from Secret of Monkey Island sounds like Maybe Anthropic crawled a whole bunch of online music discu
Related coverage; assess separately
x.comAnyone know when the models started being able to compose competent music? I had Opus 5.5 "write some computer game music for me" and the results were much better than I
Related coverage; assess separately
x.com