The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Apple Music Playlist Curator listing page.

English | 简体中文 | 繁體中文 | 日本語 | 한국어 | Español | Português do Brasil | Deutsch | Français
Deep playlist curation for Apple Music: describe a feeling, scene, era, tension, or narrative arc; your agent turns it into a catalog-grounded selection whose versions, pacing, and transitions hold together as a listening experience.
See the 12-track curation demo: 22 grounded candidates become a three-act story, including a case where the numerically cheaper order damaged the narrative.
Pure Python standard library — no pip install required to run, and no Apple Developer Program
membership needed. Requires Python 3.10+ and works on Windows / macOS / Linux.
The primary interface is the local am-mcp stdio server. The language model already running in
your MCP client interprets the brief and chooses candidates; this project searches the Apple Music
catalog, resolves exact tracks, and performs account operations. There is no bundled model, LLM
API key, artist list, or fixed theme. The CLI remains available for login, diagnostics, scripting,
audits, and advanced sequencing.
Recommended — install the MCP stdio service:
Already use uv? Run the published package without a permanent install:
For an MCP client, the equivalent Registry-aligned configuration is:
Register am-mcp in the client. The common configuration shape is:
Then describe the result, not the implementation:
Create a 25-track late-night driving playlist: atmospheric alternative R&B and electronic, mostly from the last ten years, no live versions, with a calm landing.
Clients with MCP Prompt support can select create_playlist_from_description. In every other
client, send the same request in chat: the server instructions and typed tools expose the same
status → candidate pool → catalog grounding → direct comparison → dry-run → create workflow.
From a clone (nothing to install):
Installation also puts the CLI and MCP commands on your PATH:
For development, pip install -e . from a clone makes edits take effect without reinstalling.
See SETUP.en.md for the credential walkthrough (three ways to get the user token, including a zero-dependency one).
This is deliberately harder than “make me a workout playlist.” The brief asks music to carry a plot, and some of its constraints cannot be expressed as tempo or mood sliders:
Build a 12-track, three-act story in which a machine wakes in a city, mistakes attention for intimacy, asks to be touched, becomes vulnerable, and sees dawn. Cross electronic music and art pop from the late 1970s to the present; use one track per artist, studio recordings only, and let the voices become progressively more human. The ending must feel quiet and earned, not merely low-energy.
The run below used the public US Apple Music catalog on 2026-09-23. It did not read or write a private library.
Each title below opens the exact US catalog recording returned by the resolver, so the sequence can be auditioned rather than taken on trust.
| Act | Grounded order | What the sequence is doing |
|---|---|---|
| I — Boot | The Robots — Kraftwerk Technopolis — Yellow Magic Orchestra Kid A — Radiohead | A body, then a city, then an unstable first-person voice. |
| II — Desire | Oblivion — Grimes Digital Witness — St. Vincent Is It Cold In The Water? — SOPHIE Touch — Daft Punk & Paul Williams All Is Full of Love — Björk | Public attention becomes bodily risk, transformation, a request for contact, and finally an answer. |
| III — Re-entry | Cellophane — FKA twigs Retrograde — James Blake Long Road Home — Oneohtrix Point Never An Ending (Ascent) — Brian Eno | The synthetic shell fails; retreat becomes return, and the story lands at dawn. |
The interesting failure happened during ordering. With only the three acts locked, the numerical
optimizer cut the measured cost from 48.28 to 14.88 — but put Retrograde after the dawn and
made All Is Full of Love answer a request that had not happened yet. That is cheaper and worse.
The host model therefore added semantic beat boundaries (request → answer, return → dawn) and
let am_optimize_order make only local changes inside those boundaries. Selection and story stayed
linguistic; BPM, key, energy, and valence remained supporting evidence.
This is the normal MCP workflow: the host model interprets the brief and proposes more candidates
than it needs; am_resolve_candidates grounds them; the model chooses tracks and interprets narrative roles only when the brief calls for them;
am_optimize_order optionally checks local flow without crossing semantic boundaries; and
am_create_playlist(dry_run=true) verifies the exact final recordings before the write.
To compare that workflow with a one-shot baseline, use the six difficult multilingual briefs and blind-listening protocol in examples/. The suite provides auditable constraints, not predetermined “correct” songs.
| File | Purpose |
|---|---|
am_playlist.py | Core: token management, catalog search, create / edit / delete playlists, track resolution |
am_mcp_server.py | Primary MCP stdio service: one description-to-playlist prompt plus 13 tools |
playlist_audit.py | Metadata audit: length, artist concentration, genres, eras, durations, duplicates, interludes |
playlist_flow.py | Audio-feature audit: BPM / key / loudness / energy / valence, adjacency checks, arc shape |
playlist_optimize.py | Simulated-annealing track ordering against the measured rules |
listening_stats.py | Listening history: recently played, and per-track/album/artist play counts (Apple Music Replay backend) |
profile_library.py | Descriptive sample profile: measured BPM / energy / valence spread and sample-relative quadrants; it does not define the user's taste |
am_library.py | Your library: paged export of every catalog-backed song you own, enriched with ISRC / year / genre |
build_pool.py | Listening-evidence pool: merge recent plays with multiple Replay years, retaining dates and play counts for the LLM to interpret |
Supporting modules:
| File | Purpose |
|---|---|
playlist_core.py | Platform-neutral core: Camelot, BPM folding, the four adjacency rules, the six narrative shapes. Imports nothing from this project and nothing third-party |
am_paths.py | Platform-neutral paths and version — where config and cache live |
am_meta.py | The single catalog_meta implementation (batched catalog lookups) |
All three analysis modules are importable as libraries:
The standard prompt create_playlist_from_description asks for a natural-language brief and
optional name, track count, and response language. Curation stays in the host model; the tools are
the grounded Apple Music execution layer:
am_status · am_search_songs · am_resolve_candidates · am_list_playlists · am_show_playlist ·
am_create_playlist · am_add_tracks · am_delete_playlist ·
am_audit_playlist · am_analyze_flow · am_optimize_order ·
am_recently_played · am_top_played
am_resolve_candidates grounds a generous LLM-proposed pool in real catalog metadata, flags
duplicates and suspicious versions, and deliberately does not score theme fit. The host model
compares candidates directly with the user's words and explains their playlist roles.
am_analyze_flow diagnoses transitions; am_optimize_order can optionally refine ordering inside
already chosen narrative blocks. It never decides which songs belong in the playlist.
Mount it in a Cordis agent preset with the template in preset/, or wire it into
any other MCP client with:
Complete tested examples for Codex, Claude, Cursor, VS Code/Copilot, Gemini CLI, Windsurf, Docker, Cordis/DSH, and Harness are in docs/client-setup.md.
Build the non-root local container with:
The client must run it attached with docker run --rm -i; mount only the app config directory and
a writable cache as shown in the client guide. Config stays writable so token refresh can persist.
Never bake Apple credentials into the image.
The suite is entirely offline — no network, no credentials. Two kinds:
catalog_meta, no non-None --storefront default, cache outside the
repo, every MCP tool wired to a handler, and the optimizer free of platform imports.They earn their keep immediately — the suite caught a syntax error in a file written minutes earlier, before it was ever run.
Apple's catalog API exposes no audio features at all — no tempo, key, loudness, energy, or
valence. Spotify's audio-features endpoint was shut off for new apps on 2024-11-27, and
AcousticBrainz retired in 2022.
playlist_flow.py bridges the gap with a free, key-less chain built on ISRC, which Apple does
return:
Results are cached locally, so the network cost is paid once per playlist.
Coverage is reported, never silently dropped. Every consumer prints a funnel saying why each track could not be measured:
That distinction is the point. No ISRC means the chain cannot start at all — a different feature source will not help. Not in the source means the ISRC is fine and switching sources (or analysing the audio locally) would fix it. Collapsing both into "missing features" discards the only information that tells you what to do next.
It matters more than it looks: tempo / key / energy / valence are the only things the
adjacency rules and the arc can act on, so coverage is the ceiling on how good an ordering can
be. At 60% coverage, four positions in ten were never evaluated — while the cost number still
looks excellent. The optimizer therefore prints coverage above its cost lines, and warns below 90%.
Derived from the research collected in docs/ — including a PLOS ONE study in which
130 music professionals sequenced albums, and a randomized trial on mood-adaptive music ordering.
Hard adjacency rules
These four have exactly one definition, in playlist_core.check_pair(), and both the audit and
the optimizer call it. They used to be implemented twice, and the copies disagreed: the audit
called a track "slow" below the tempo's 25th percentile while the optimizer used a fixed 100 BPM,
and the audit never checked the BPM-jump rule at all. That is a tool diagnosing against one
standard and repairing against another — so the count it reported could not be trusted.
Note on the "slow" threshold. It is absolute (100 BPM), not data-driven, deliberately: the optimizer evaluates the same sequence thousands of times while annealing, and a percentile threshold would drift as the permutation changes, so the cost would never settle. The cost is that a uniformly slow playlist flags every adjacent pair — that is real, not a sequencing failure, and the report says so.
Global narrative arc is optional. With no --arc, the optimizer only considers local
adjacency and does not impose an emotional or tempo trajectory. The host LLM should interpret a
free-form story from the brief and preserve its order through blocks; use a named preset only when
the user explicitly asks for that kind of curve.
| Axis | Target |
|---|---|
valence, energy, loudness | the chosen archetype, only when --arc is explicit |
tempo | inverted U — fast in the middle, only when --arc is explicit |
Six optional shapes: rags-to-riches, tragedy, man-in-a-hole, icarus, cinderella,
oedipus. The target curve and the shape the audit classifies come from the same table in
playlist_core.ARCHETYPES. The audit may still report which preset a playlist resembles; that
diagnosis does not make the preset a recommendation or an ordering target.
When a preset is explicitly requested, valence / energy / loudness follow its emotional
trajectory and tempo uses a separate inverted-U target. With no --arc, neither global target is
applied.
An arc cost is a distance from an explicitly requested feature curve, not a measure of whether a playlist is good. Coverage limits what it describes, and a lower cost cannot override the brief or prove that one sequence sounds better.
The optimizer preserves your grouping (movements / eras / moods) and only reorders within groups, so thematic structure survives the loudness tuning. Drop the grouping and it reorders freely — measurably "smoother", at the cost of your narrative.
Example from a real run: cost 104.46 → 19.41 with grouping preserved, → 1.28 ungrouped.
Accept-Encoding. Decoding the
raw bytes as UTF-8 yields garbage that looks like an empty body. (This is fixed in am_playlist.py.)DELETE only works on amp-api.music.apple.com. The documented host api.music.apple.com
returns 401 for playlist and library-song deletion.x-apple-client-version to amp-api — it turns into a 500.[70,160)
before comparing, or "two slow tracks adjacent" over-reports by ~4×."Title - Artist" straight into catalog search — you get live/remastered takes.
Search on the space-separated form and score versions afterwards.More in docs/apple-music-api-notes.md and
skill/reference.md.
| Doc | Contents |
|---|---|
SETUP.en.md | Credentials: what tokens exist, how to get each one, security notes, troubleshooting |
docs/client-setup.md | Client-specific MCP, Docker, Cordis/DSH, generic harness, and Harness Platform setup |
docs/publishing.md | Maintainer-only PyPI and official MCP Registry publishing checklist |
docs/apple-music-api-notes.md | Token model, endpoint contracts, measured API behaviour, eval of 7 automation approaches |
docs/how-to-build-a-good-playlist.md | Curation methodology: adjacency physics, arc data, six narrative shapes, the ISO principle |
docs/playlist-curation-survey.md | Survey of published curation guidance (platform rules, DJ methods, academic findings) |
docs/evaluation-signals.md | LLM-native curation: direct candidate comparison, catalog grounding, readable constraints, and why scalar theme scores stay out of the critical path |
docs/natural-language-curation-evidence.md | Primary research on broad music intent, reference-vs-request errors, intent hallucination, scrutable language profiles, and the reusable validation protocol |
examples/ | Reproducible multilingual curation evaluations: six difficult briefs, a baseline protocol, auditable evidence, privacy rules, and no golden track lists |
docs/algorithm-review.md | Review boundary for early heuristics: which decisions belong to the LLM and which deterministic algorithms should retain |
docs/platform-adapters.md | The platform-adapter boundary: what is platform-neutral, what an adapter must provide, and what breaks on a service that exposes no ISRC |
llms.txt | Concise agent-readable map of the curation boundary and the most useful project documents |
skill/ | Agent skill: workflow + the accumulated gotcha list |
CHANGELOG.md | Release history, including behaviour changes between versions |
preset/ | Cordis agent preset template that mounts the MCP server |
For the complete Simplified Chinese guide, see README_ZH_CN.md.
MIT — see LICENSE.
Unofficial community tooling. Not affiliated with or endorsed by Apple. Uses your own Apple Music account for personal use; follow Apple's terms of service.