Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Claude Opus 5.5 Composes Game Music Competently - OpenSmartRoute
Simon Willison tested Claude Opus 5.5 on composing Monkey Island-style game music. The model produced surprisingly high-quality results in a text-based format.
Key points
Claude Opus 5.5 composed competent computer game music.
The output matched the quality of the original Monkey Island soundtrack.
Simon Willison tested this capability on October 6, 2026.
The model designed a simple text-based format for playback.
Why it matters: Developers can now expect agents to generate creative assets like music without manual composition tools.
By OpenSmartRoute editorial · written through the router by writer-small
Screenshot of a retro pixel-art web music player. Header in blackletter type reads "Scrimshaw Jukebox", with the text "Six original adventure-game tracks, written as plain text and played by a synthes. Image: Simon Willison on LLMs (original)
Simon Willison tested Claude Opus 5.5 on composing music. He wanted to see if the model could create game scores. The test focused on a specific retro style called Monkey Island. This style features adventure games from the late 1990s. The model produced text-based music formats that sound good.
Simon Willison is a well-known figure in the AI community. He often posts detailed breakdowns of how models work. His blog serves as a primary source for engineers and researchers. Readers follow him to understand practical limitations and capabilities. This specific test explores a new frontier for language models.
The goal was to move beyond simple text generation. The model needed to handle audio-like structures through text. Simon Willison asked for a design that could play out loud. He sought an artifact containing example tracks within the output.
Claude Opus 5.5 is a large language model released recently. It belongs to the family of models known as generative AI. These systems can create new content based on training data. The test checked if it could mimic musical composition skills.
Simon Willison asked for music matching the quality of Monkey Island. This refers to the original secret game by Ron Gilbert and Dan Connors. The game is famous for its catchy, chiptune-style soundtrack. Players remember the themes from the 1990s adventure series.
The model leaned heavily into the Monkey Island theme. It did not follow the initial instructions perfectly at first. The results were surprisingly good despite this drift. Simon Willison noted the unexpected quality of the output.
He wondered if competent music composition is a new capability. This ability might be similar to advances in 3D graphics. Text models have evolved rapidly over the past few months. Some features emerged suddenly rather than gradually.
Simon Willison suggested running careful experiments with other models. He compared recent releases with not-so-recent versions from earlier years. This comparison would confirm if the skill is truly new. Or if older models could already do this task.
We are approaching a date set as October 3rd, 2026. This date marks a potential shift in budget caps for AI services. Developers will need to manage costs more strictly soon. OpenAI DevDay 2026 took place on September 29th of that year.
Tiiuae released Falcon-Emirati-7B to handle Emirati Arabic dialect nuances. It scores 84.83% on the Alyah benchmark, beating larger competitors.
The blog post itself was published on October 6th, 2026. It is part of a series tracking developments in large language models. Readers can find related posts from late September for context. The timeline shows how quickly the field is moving forward.
Simon Willison offered sponsorship options for his newsletter. He asked readers to consider supporting him for ten dollars a month. This funding helps curate an email digest of monthly LLM updates. The digest covers important developments in the industry.
The test results hint at broader trends in model capabilities. Engineers should watch for similar shifts in other creative domains. Music composition might be one of many new skills appearing. Other areas like video generation or complex reasoning could follow.
Understanding these shifts helps engineers plan their infrastructure better. They need to prepare for models that do more than chat. Budget caps mean costs will rise if usage increases significantly. Teams must evaluate the cost per token carefully now.
The prompt used by Simon Willison was quite specific about format. He requested a text-based design for the music structure. The artifact needed to include tracks playable out loud. This requirement pushed the model beyond standard text generation tasks.
Claude Opus 5.5 handled the request with surprising creativity. It generated code or data structures representing musical notes. These structures could theoretically be rendered by audio software. The output resembled a score rather than just lyrics.
The Monkey Island theme is distinct in its use of melody. It often uses simple, repetitive patterns that are easy to remember. Simon Willison noted the model leaned too hard into this style. The result was very recognizable as that specific game's music.
This suggests the model has access to training data about the game. It likely saw descriptions or transcriptions of the original soundtrack. The model combined these details with its own compositional abilities.
High-quality output in a text format is still rare for models. Most systems struggle to produce coherent, structured creative work. This test showed that Claude Opus 5.5 can handle such tasks. It produced something usable without human editing in many cases.
Simon Willison called this a new capability emerging recently. He compared it to the sudden rise of good 3D graphics. Both fields saw breakthroughs that seemed to appear overnight. Text models are now catching up in these creative areas.
The implications for agent systems are significant if true. Agents could compose music or generate other media autonomously. This would change how content creators work with AI tools. They might replace human composers or designers entirely.
Engineers need to test similar prompts with their own models soon. Running experiments is the only way to verify these claims. Different models have different strengths and weaknesses in creativity. Some might excel where others fail at music generation.
The date of October 3rd, 2026, highlights an upcoming industry change. Budget caps will force developers to be more efficient with tokens. This means every generated token counts toward the total cost limit. Teams must optimize their prompts and model choices accordingly.
OpenAI DevDay 2026 provided a live blog covering recent announcements. The event took place on September 29th, 2026. It set the stage for these new capabilities and budget discussions. The blog post builds on themes introduced at that conference.
Simon Willison's newsletter serves as a curated source of information. Readers pay ten dollars monthly to get the most important updates. This supports independent reporting in an industry dominated by big tech.
The test results are concrete evidence of growing model intelligence. They show that text models can understand abstract creative concepts. They can then translate those concepts into structured outputs. This translation is key to building useful AI tools.
Engineers should look for similar benchmarks when evaluating new models. Music composition is a good proxy for general creativity. Other tests might include writing poetry or generating code art. These tasks require understanding context and style, not just facts.
The ability to compose music competently changes the landscape of development. It means agents can create full experiences without human help. This could lead to automated game development pipelines in the future.
Simon Willison remains a key voice for the open-source community. His work often bridges the gap between research and practice. Engineers rely on his insights to navigate the rapid changes.
The prompt instructions were clear about the desired output format. The model followed them well enough to produce playable music. This adherence to constraints is crucial for building reliable systems.
Cost management will become a priority as budget caps tighten. Teams must balance quality with the rising price of tokens. Efficient prompting can reduce costs significantly in creative workflows.
The comparison to 3D graphics illustrates the pace of innovation. Both fields moved from being niche to mainstream quickly. Text models are now following this same trajectory in creativity.
Engineers should expect more complex tasks from models soon. Simple text generation will no longer be enough for many use cases. Creative generation and multi-modal outputs will become standard requirements.
The blog post ends with a call to action for readers. They are encouraged to support the newsletter or explore further resources. Simon Willison provides links to related posts on his website.
In summary, Claude Opus 5.5 demonstrated surprising musical composition skills. It produced text-based music that matched a specific retro style. The results suggest a new era of generative AI capabilities. Engineers must prepare for these changes in their workflows and budgets.
What Changed - Simon Willison tested Claude Opus 5.5 on composing music
Simon Willison ran a specific test with Claude Opus 5.5. He asked the model to write computer game music. The output resembled themes from Monkey Island. This marks a shift in what text models can do. They are moving beyond simple text generation. Music composition is now within reach of these systems.
The Prompt - Specific instructions given to the model for game music
The user gave clear instructions about the desired output. The request asked for a simple text-based format. It also required an artifact that could play the music aloud. The prompt specified the quality needed for the secret of Monkey Island. These constraints guided the model's creative process. The model leaned heavily into the Monkey Island theme. This happened even though the user did not intend such a strong focus.
The Results - High-quality output resembling Monkey Island themes
The generated music was surprisingly good in quality. It captured the specific style of the original game. The results showed competence in musical structure and tone. Simon Willison noted the unexpected strength of the output. He wondered if this ability is new to text models. Or perhaps other models could do this for a while already. Further experiments are needed to confirm this trend.
Background Context - How text models have evolved into creative generators
Text models started by processing written language only. They learned to follow instructions and generate coherent paragraphs. Recently, they have expanded their capabilities significantly. This evolution includes handling abstract creative concepts well. The shift resembles the rise of 3D graphics in computing. Both fields moved from niche tools to mainstream applications quickly. Text models are now following a similar path in creativity.
Why it matters - New capabilities for agent systems and content creation
This change allows agents to build full experiences alone. It removes the need for human help in some tasks. Automated game development pipelines could emerge soon. Content creators will have new tools at their disposal. The ability to compose music competently changes the landscape of development. It opens doors for complex, multi-modal outputs from AI systems.
What to do - How engineers can test similar prompts with their own models
Engineers should look for benchmarks that measure creativity. Music composition is a good proxy for general intelligence. Other tests might include writing poetry or generating code art. These tasks require understanding context and style deeply. Engineers must evaluate models on these new capabilities. Simple text generation will no longer be enough for many use cases. Creative generation and multi-modal outputs will become standard requirements soon.