Accompanist (Magenta RealTime 2)

Render chords, MIDI takes and free improvisations as audio with the experimental Accompanist module.

The Accompanist uses Magenta RealTime 2 small to turn pitch conditioning into 48 kHz stereo music. It is an experimental, browser-only, host-managed module. Inference runs in a worker outside the realtime audio callback.

Availability

Accompanist is a dedicated module in the Maru catalog. Its finite-take workflow is tested in the local website host with the local model bundle. Public online generation is unavailable until the model-delivery service, hosted files and production browser workflow are verified. Shipping the updated website controls alone does not enable generation.

Maru VST3, Audio Unit, CLAP and native standalone execution are unsupported: the native host does not implement the required model capability. Native MLX integration remains research, with no promised plugin release. Studio source editing likewise needs a host that supplies maru/accompanist for execution. The repository has command-line study/export scripts, but no supported Maru generation CLI matching this module's contract. Exported WAVs can be used in DAWs independently of native inference.

Load weights when needed

Browsing the catalog does not download weights. Activating Accompanist's audio host loads approximately 698 MB; generation remains idle in Clips mode until you request a take. Style changes use built-in tokens and need no extra model download. Persistent caching is not yet qualified, so do not assume a reload will work offline or avoid downloading again.

Direct downloads from Hugging Face are possible in principle, but are not enabled in this module. Google's official checkpoints need conversion for this browser runtime. The community ONNX export also needs a compatible embed graph and the validated Maru manifest; its current embed graph cannot load unchanged in our browser runtime. Public generation remains pending while compatible model delivery is prepared.

Choose a mode

Generation mode decides whether to make finite takes or generate continuously. Conditioning source decides which notes a finite take uses.

ChoiceStepsLimit
ClipsActivate, wait for loading, select a conditioning source and style, then Render loop or Render clip.This is the tested local workflow.
Held chordIn Clips, hold your chord and press Render. With no held notes, C4/E4/G4 is used.Takes a snapshot of the chord; it does not record subsequent chord changes.
MIDI takeIn Clips, set tempo/bars, Capture MIDI take, play or route MIDI, Finish MIDI take, then Render.Render length must cover the captured take.
FreeIn Clips, select Free, choose a style and press Render.No note conditioning.
StreamSelect Stream and play live MIDI. Return to Clips to stop generation.Experimental; current WASM generation is slower than realtime and may underrun. Captured MIDI is not replayed here.

Render loop uses bars and tempo to calculate a 4/4 duration, up to sixteen seconds. Render clip uses Length (2–16 seconds). After rendering, use Audition loop, Stop audition, Download WAV and Download recipe. Switch to Clips before auditioning a saved take. Cancel stops a render and keeps the previous completed take.

For chord changes or a solo line, capture MIDI instead of Held chord. Route MIDI-GPT's output into Accompanist's MIDI input to render generated notes as audio. Accompanist chooses the resulting ensemble; these controls do not guarantee isolated solo audio or exact harmonic adherence.

Render a take

Activation loads approximately 698 MB of model files. Clips is the default mode and leaves generation idle after loading. Choose a style, drums policy, note guidance, onset mode, pitch freedom, temperature, top-k and seed.

  1. Select Held chord, MIDI take or Free. Held chord uses current keyboard notes, or C4/E4/G4 if empty. Free supplies no MIDI notes.
  2. For a MIDI take, set tempo and bars, press Capture MIDI take and play or route notes into the module's MIDI input. Press Finish MIDI take to stop early; capture otherwise ends after the selected bar duration.
  3. Press Render loop for a 4/4 bar duration, or Render clip for seconds. Renders are limited to sixteen seconds, show generation/decoding progress, and can be cancelled while retaining the previous take.
  4. Audition loop plays through the host audio output. Download its WAV and JSON recipe to retain the audio, seed, style tokens and MIDI conditioning.

A fixed seed repeats a take within the same model/runtime/provider and recipe. Changing the seed explores another performance. This is an audio alternative or complement to MIDI-GPT: MIDI-GPT generates editable notes; MRT2 generates a stereo mixture and does not export MIDI or separated stems.

Variations and context controls

Keep the captured MIDI, style and other settings fixed, change Seed, then render again to explore another performance. Download each WAV and recipe to retain it; the module keeps one completed take at a time. An identical recipe repeats on the same tested model/runtime/provider, with no cross-provider byte-identity guarantee.

Auto-strum lets the model choose attacks for held pitches; Onsets carries your attack/sustain states. Note guidance and Pitch freedom influence how closely the output follows those notes. Drums offers Model decides, Off and On as model conditioning; it does not filter percussion out of decoded audio. Style, note and drums guidance each have a separate strength. Temperature and Top K change sampling variety. Frames Ahead and Decode Chunk adjust Stream buffering and decoding; they cannot make a slow model realtime. Watch frame throughput and underruns while experimenting.

Level changes monitoring volume. The WAV contains the generated take with peak attenuation where needed; monitoring level and connected rack effects are not baked into the export.

In Stream, MIDI Gate mutes output when no notes are held; inference can continue while muted.

Reset cancels a render, stops audition and clears captured/held notes and rolling model context. Prefill silence initializes rolling context with approximate silence. Finite takes reset context for each render, so prefill is mainly a Stream experiment. It does not supply an audio example. Arbitrary text prompts and audio-example conditioning are not exposed; choose from the committed style presets.

Timing and limits

Tempo sets duration and captured-note timing; the model receives no tempo token. Notes follow a forty-millisecond grid. Trimming to an exact bar length does not guarantee beat alignment or a seamless musical loop. Velocity, channel-specific timbres and expression are not modeled.

This fp16 bundle uses WASM because its WebGPU path has historical correctness failures. Browser generation is slower than realtime. Stream mode is optional, runs at any host sample rate (the model's 48 kHz output is converted to the device rate) and remains unqualified for sustained realtime playing. Its face reports measured frame throughput, queued audio and underruns; queued audio is only one contributor to total response delay. Offline audition uses the host's sample-rate conversion when the device runs at another rate.

Persistent weight caching is not qualified. Weight URLs come from the existing model-delivery service or the linked local development mirror. The shared stream bridge uses a ring on isolated pages and MessagePort transport otherwise.

Attribution

The Google model card identifies the weights as CC-BY-4.0 and the reference code as Apache-2.0. Attribution appears on the face and in the take recipe. The recipe's weight license describes the model; it does not assign a license to the generated recording.