Live Music Models

notes.

Live Music Models

Source: Lyria Team et al., “Live Music Models” (2025), arXiv:2508.04651. Full paper

Claim

Current music models can synthesize a continuous stream in real time while accepting changing text or audio controls. This is a distinct mechanism from a generative music engine that arranges predesigned loops or stems.

Systems

Magenta RealTime is an open-weights model intended for local or user-managed deployment. Lyria RealTime exposes a larger model through an API. Both generate audio tokens from short chunks and condition the output through a shared text–audio representation (pp. 1–3).

The system predicts two-second chunks from ten seconds of coarse audio history. Text and audio controls specify high-level qualities such as style, genre, instrumentation and mood, and the control can change between chunks (pp. 3–6).

The team evaluates audio quality, prompt adherence and transitions between prompts. It also uses internal play-testing and a small user study because interactive control cannot be reduced to offline audio metrics (pp. 4–8, appendix F).

Limits

This is a technical preprint from the team that built the models. It does not study functional listening, wellness effects, commercial adoption, artist labor, authorship, licensing or long-term listening. Its results establish a technical possibility, not a social outcome.

Presentation use

Use this paper to separate two claims. Endel already leaves some form open until playback through predesigned material and logic. A live music model can add real-time synthesis of the material itself. AI therefore extends the existing runtime form rather than creating functional listening or open form from nothing.

7 paragraphs257 words1,687 characters