Between Genres, Before Genre
A research dossier on genre prompts, model-specific musical spaces, and the social formation of genre.
Current research decision: an existing Udio study
Primary micro case
Sangheon Park and Claire Arthur’s “From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music” provides the micro analysis. The authors pair 200 real Udio prompts with their 32-second generated clips and collect 2,624 free descriptions from 148 listeners. Seventy participants belong to an English-speaking group and seventy-eight to a Korean-speaking group. The study already contains the generation, listening and coding work that an original case would require.
The bounded object is the movement from prompt language through generated audio to listener language. Park and Arthur classify the words in each stage as Genre, Mood and Emotion, Instrumentation, Music Theory, Timbre, Function, or Story and Narrative. This makes one part of Udio’s mediation observable: genre terms enter as compact instructions and return through the signs listeners hear in the result.
Genre appears in about 95 percent of the prompts and accounts for 38.8 percent of their coded vocabulary. It appears in about 61 percent of listener descriptions and accounts for 15.2 percent of descriptive vocabulary. Instrumentation becomes the largest descriptive category at 21.1 percent. Genre terms propagate more reliably than narrative terms, although listeners often describe a semantic neighbourhood or its musical signs. A prompt containing metal may return as rock, punk, electric guitar or screaming in a listener’s account.
These results permit a precise micro claim. Udio gives genre a privileged place in the instruction, and the generated audio often preserves enough associated signs for listeners to recover a related category. The category also changes during mediation. Prompt writers rely on genre names more heavily than listeners, who describe instruments and other audible properties. The system translates a cultural label into a bundle of musical cues rather than carrying one stable word through the process.
The study reports the aggregate relation across 200 cases. It does not isolate mixed-genre prompts, test whether listeners hear a new genre or reveal Udio’s internal representation. Its cultural comparison also remains exploratory because the two groups came through different recruitment channels and a large language model assisted part of the coding. These limits belong inside the micro analysis.
Supporting aesthetic case
Guilherme Coelho’s “Latent Music: Emergent Sonic Forms and Sonic Liminality in Text-to-Audio Systems” supplies the mixed-prompt case that Park and Arthur do not isolate. Coelho reports many hundreds of Udio generations and selects outputs that produced recurring experiences of “almostness.” One sequence uses john oswald plunderphonics experimental, deconstructed, classical music, stockhausen, no vocals and repeatedly extends the first thirty-two seconds with the same prompt. Other prompts combine glitch, plunderphonics, electroacoustic, experimental and classical descriptors.
Coelho describes signs that appear and recede without settling into one category. His term “genre slippage” is useful when an output moves among recognisable references over time. “Prompt assemblage” names the relation among descriptors whose effects cannot be separated into independent controls. These are Coelho’s reported listening judgments. The essay can analyse his account and examples without presenting them as our own listening observations.
His method sets a clear limit. Coelho selected the examples himself with prompts visible and did not report the retained proportion, blind listeners, independent coding or a repeatability estimate. His paper establishes a documented aesthetic encounter with mixed Udio prompts. It does not establish how frequent genre slippage is or whether listeners share the judgment.
Technical support from an open model
Marcel A. Vélez Vásquez, Charlotte Pouw, John Ashley Burgoyne and Willem Zuidema’s “Exploring the Inner Mechanisms of Large Generative Music Models” tests genre representations in three sizes of MusicGen. The authors construct 100 prompts for each of six genres and intervene on components across the model’s transformer layers. Instrument-to-instrument interventions work more effectively than genre-to-genre interventions. Their interventions reduce the original genre more reliably than they introduce the desired one, while other parts of the audio also change.
The study supports a narrow technical inference: genre behaves as a distributed combination of features in MusicGen rather than a modular switch like one instrument label. Its classifier-based evaluation contains no human listening study, and MusicGen cannot stand in for Udio’s proprietary architecture. The result still challenges the simple picture of two clean genre coordinates with an empty midpoint between them.
Four meanings of between
The published cases keep four relations separate:
- Lexical relation: a prompt places several genre names in one instruction.
- Computational relation: a model processes genre information through an internal representation.
- Sonic relation: listeners hear musical signs associated with several categories or hear them change over time.
- Social relation: musicians, listeners or institutions treat a recurrent difference as a category with shared rules.
Park and Arthur observe the movement from prompt to perception. Coelho documents one author’s experience of sonic instability. Vélez Vásquez and colleagues intervene inside an open model. None follows a new category through social adoption. Fabbri’s genre theory supplies that final level.
Current research question
How does Udio translate genre words into audible signs, and why does an interstitial output remain different from a new genre?
The question begins with Park and Arthur’s prompt-audio-description relation. Coelho adds the mixed-prompt aesthetic case. The cultural analysis then asks how a musical difference acquires rules, a public and institutional consequences.
Current thesis
Text-to-music systems turn genre from a descriptive and circulatory category into a generative control. Park and Arthur show that genre terms pass from Udio prompts into listener perception more reliably than narrative language, although listeners reconstruct them through instruments and related musical signs. Coelho’s mixed prompts produce selected cases of genre slippage. These systems can generate interstitial events, while genre formation depends on the later reproduction and social acceptance of musical rules.
Text-to-music platforms influence this later process because they define prompts, recommend tags, distribute outputs and classify what users publish. One generated clip remains a musical event. Repeated model behaviour may become a platform style. Fabbri’s account requires some social acceptance and reproduction before a difference functions as a genre rule.
Assignment fit
The assignment requires a move from a focused micro analysis to broader theoretical and cultural analysis. This route makes that movement inside one object instead of attaching a general discussion of AI after a close reading.
- Focused tool and phenomenon: Udio’s translation from real prompts into generated clips and listener descriptions.
- Micro analysis: Park and Arthur’s study design, taxonomy, quantitative results and limits, followed by Coelho’s documented mixed-prompt case.
- Theoretical bridge: the genre tag acts as a sign, a constraint, and an element of an epistemic instrument.
- Macro analysis: the conversion of musical culture into metadata, the platform’s power to make classifications productive, the cultural limits of an English-language prompt register and the social formation of genre.
- Critical reflection: the published cases establish prompt-perception relations and selected experiences of slippage. They do not reveal Udio’s latent topology or demonstrate the birth of a genre.
This route meets the brief without original generation or listening. It covers prompt-based composition, intersemiotic translation, distributed creativity and genre through a defined empirical case. A final length of approximately 1,400 to 1,600 words gives enough space for the micro evidence and the cultural analysis.
Title options
- Between Genres, Before Genre: Udio from Prompt to Perception
- Genre Labels as Musical Instructions in Udio
- Genre as a Generative Control: Prompting and Perception in Udio
Current preference: Between Genres, Before Genre: Udio from Prompt to Perception.
The movement from micro to macro
The genre label changes function
The decisive hinge is the change in what a genre word does. In criticism, record shops, catalogues, and streaming services, a genre label usually describes, sorts, or circulates music that already exists. In Udio, the same label enters before the sound and helps cause it. The model does not only classify a song as eurodance or black metal. It treats those terms as instructions for making one.
This change links the interface to a larger cultural history. The Session 4 course material describes presets and samples as sonic signs that carry expectations about genres, eras, moods, and techniques. A preset does more than store parameter values. It packages design decisions, historical memory, and a likely use. A genre prompt extends this logic. It turns a much larger cultural category into a reusable control that can request the rhythm, timbre, instrumentation, voice, form, and production of a whole song.
Thor Magnusson’s account of the digital instrument as an epistemic tool explains why this is culturally important. Software must formalise musical ideas in classifications, mappings, and permitted actions. The designer decides what the interface makes easy to think and do. Udio’s tag system presents genre as a set of named and combinable units. It makes eurodance + black metal legible as an operation even when no community has agreed on the relation between those histories. The interface is therefore a theory of music in usable form.
A chain of scales
The essay can move through seven connected scales:
- Prompt corpus: users translate intended music into genre terms and other verbal signs.
- Generation: Udio turns those instructions into audio through an undisclosed process.
- Perception: Park and Arthur’s listeners reconstruct the audio through genre, instruments, affect, theory and narrative.
- Aesthetic instability: Coelho hears several genre signs emerge and recede within selected Udio outputs.
- Model and corpus: the uneven relations reveal a system-specific operational concept of genre, while the proprietary representation remains unavailable.
- Platform: Udio supplies the vocabulary, interface and mechanisms through which users generate and publish music.
- Culture: musicians, listeners, critics and institutions can imitate, name, contest or ignore a recurrent difference. Socially accepted rules begin at this level.
Each scale answers a question raised by the previous one. An audible feature leads back to the instruction that requested it. Repeated variation leads to the model’s mediation. A stable model tendency raises the question of who can repeat and circulate it. Circulation raises the question of whether a category becomes socially consequential.
Classification becomes causation
Mads Krogh’s study of Spotify describes a shift from genres as qualitative musico-cultural worlds toward numerical relations in a continuous similarity space. Platform classifications already organise discovery and can create operational groupings around sound, mood, activity, or audience behaviour. Text-to-music adds another step: the platform’s categories do not only place existing works in relation. They condition future works.
The resulting feedback loop is a strong macro object:
historical music and discourse -> training data and labels -> prompt vocabulary and interface -> generated output -> user selection and publication -> new tags, listening data, and imitation -> later training and classification
The loop does not remove human agency. It redistributes it. Genre histories, often built through collective and contested practice, are compressed into prompt tokens. Developers make those tokens operational. Users assemble and select them. Listeners interpret the outputs. Platforms can then use the resulting activity to refine classifications and distribution. Georgina Born’s account of distributed and relayed creativity is useful here because no single actor explains the final object or its later cultural force.
The social test of a genre
Franco Fabbri supplies the boundary between the generated event and the macro claim. A genre is a set of real or possible musical events governed by socially accepted rules. These rules include formal and technical, semiotic, behavioural, social and ideological, and economic and juridical relations. A clip can combine the kick pattern and synthesiser vocabulary of eurodance with the guitar timbre and vocal behaviour associated with black metal. This covers only part of what either genre means.
An output at the intersection of two genres therefore need not establish a third. The stronger question is whether a recurrent difference becomes a model for further action. Do separate musicians imitate it? Do listeners recognise and name it? Do performances, visual signs, production techniques, values, and social identities gather around it? Do platforms or labels classify and promote it? Disagreement and institutional power are part of this process. Full consensus is unnecessary, but an isolated generation remains insufficient.
This distinction produces three possible findings from the micro case:
- An interstitial event: one clip carries signs from both prompt categories.
- A recurrent platform style: repeated generations produce a recognisable relation that can be requested again within one system.
- A genre-forming tendency: musicians and listeners begin to reproduce, name, circulate, and contest the relation beyond the original interface.
Park and Arthur support the first through many prompt-output-listener relations. Coelho supplies selected evidence of the second as an aesthetic possibility. Only longitudinal cultural evidence could support the third.
Deferred original prompt matrix
The source-based route makes this experiment unnecessary for Essay 4. Keep the design as a possible later project or appendix. Do not describe its expected observations in the submitted essay because no outputs have been generated or heard.
Case design
Use Udio Manual Mode as the primary system because the exact submitted tags remain visible. Generate instrumental clips first so that lyrical content does not overwhelm the genre relation. If the first set makes vocal behaviour essential to the comparison, add one matched vocal set and report it separately.
Run these six conditions:
| Condition | Exact prompt | Purpose |
|---|---|---|
| A | eurodance | First anchor |
| B | black metal | Second anchor |
| C | eurodance, black metal | Direct mixed prompt |
| D | black metal, eurodance | Prompt-order test |
| E | eurodance song structure, black metal instrumentation and production | Separation of formal and timbral functions |
| F | black metal song structure, eurodance instrumentation and production | Inverse functional separation |
Generate at least two outputs for every condition and retain every result. Record the date, platform, model or version if shown, interface mode, duration, exact prompt, tags, output URL or local reference, and any automatic settings. Do not post-process the audio before analysis. Use one identical mixed prompt in Suno as a limited cross-platform control. The control does not rank the systems. It tests whether “between” is stable across models.
The reverse-order and function-splitting prompts are important. A simple mixed prompt can produce several mechanisms that sound similar on first hearing. The system may let one category dominate, alternate sections, attach the timbre of one genre to the form of another, or produce a compromise that weakens both. Separating structure from instrumentation tests whether the model treats genre as a modular bundle of attributes or as a more entangled relation.
No original audio has yet been generated for this dossier. The protocol defines the evidence that the final essay still needs. Sonic conclusions must remain blank until the outputs exist and are heard.
Listening codebook
For each clip, write timestamped observations under the same headings:
- Rhythm and groove: pulse, meter, kick pattern, tempo feel, syncopation, blast-beat behaviour, dance-floor regularity.
- Form: section lengths, repetition, buildup, drop, breakdown, transition, climax, continuity.
- Harmony and melody: tonal centre, harmonic rhythm, riff type, melodic contour, consonance, dissonance.
- Instrumentation: synthesiser roles, guitars, drums, bass, orchestral or electronic layers.
- Timbre and production: distortion, density, brightness, spatial depth, compression, lo-fi or polished surface.
- Voice: delivery, intelligibility, register, screaming, singing, processing, gender coding, language, if vocals are present.
- Genre signs and affect: which features trigger a category judgment and whether the judgment changes over time.
- Combination mechanism: dominance, alternation, surface graft, functional split, compromise, instability, or a recurrent grammar not captured by the anchors.
- User action: accept, reject, extend, remix, alter the prompt, or stop, with the reason for the choice.
The last item keeps the human contribution inside the analysis. A convincing result is rarely the direct consequence of a prompt alone. It is also the survivor of a selection process. Record failed, ordinary, and over-dominant outputs so that the essay does not turn one convenient anomaly into a general property of the system.
Evidential limits
This micro study can support close reading, interface analysis, and a comparison of repeatable tendencies inside a small prompt matrix. It cannot estimate how often mixed genres work across Udio, infer the training corpus, map the actual latent geometry, establish novelty against all prior music, or demonstrate the existence of a new genre. Two outputs per condition reveal possibilities and variation, not prevalence. The final prose must identify each observation as a property of a particular run unless several retained runs support the same relation.
Prompt-based research audit
Platform documentation: the interface teaches users to combine culture
Udio’s current prompting guide describes genre, mood, tempo, instruments, theme, and era as prompt components. Manual Mode uses exact tags, and the guide proposes eurodance + black metal as a combination. Udio’s Styles feature permits one or two reference styles and exposes a blending control. Suno’s music glossary likewise presents genre, style, tempo, song structure, instrumentation, and effects as controllable terms and recommends genre blending.
These documents establish interface facts and intended user practice. They do not disclose the systems’ architectures or show that each word controls an independent musical dimension. Their importance is cultural: both companies teach users to treat named histories as modular and mixable materials.
Casini et al.: large-scale prompting practice
Luca Casini and colleagues’ “Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio” analyses 101,953 publicly shared Suno and Udio songs created between May and October 2024. After filtering, the authors analyse 16,881 unique English prompts and 35,746 unique tags. Genres, qualifiers, and instruments dominate the tag vocabulary. Users alternate between comma-separated descriptors and prose, while metatags attempt to control song sections, instruments, dynamics, and production.
The study supports the claim that prompting is already a cultural practice organised around musical metadata. It also shows high vocabulary dispersion: 80.7 percent of tags occur once, and only 1,193 tags occur at least ten times. This tension matters. Platform prompting standardises the role of metadata while users continually invent local combinations.
The corpus contains only published outputs and therefore overrepresents results that users chose to expose. The authors’ maps are Uniform Manifold Approximation and Projection visualisations of researcher-created text embeddings. They are not maps of Suno’s or Udio’s internal musical representations. Use the paper for prompt practice and public circulation, not as evidence that the model’s latent space has the same clusters.
Park and Arthur: prompts, sounds, and listeners do not use the same categories
Sangheon Park and Claire Arthur’s “From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music” compares 200 real Udio prompts with 32-second outputs and 2,624 free descriptions from 148 listeners, including 70 English-speaking and 78 Korean-speaking participants. The study codes language into Genre, Mood and Emotion, Instrumentation, Music Theory, Timbre, Function, and Story or Narrative.
Genre appears in about 95 percent of prompts and accounts for 38.8 percent of prompt vocabulary. It appears in about 61 percent of listener descriptions and accounts for 15.2 percent of descriptive vocabulary. Instrumentation becomes the largest descriptive category at 21.1 percent. Genre terms propagate more reliably than narrative terms, but often through semantic neighbourhoods and audible markers rather than exact repetition: metal can be heard through rock, punk, electric guitar, or screaming.
This gives the essay rare evidence across prompt, output, and reception. It supports a careful claim that genre is a privileged conditioning language, while listeners reconstruct the result through several kinds of musical and cultural signs. It also prevents a simple one-word-to-one-sound model.
The cultural comparison is exploratory. Korean participants used more affective, narrative, and functional descriptions, while English-speaking participants used more genre and music-theory terms. Recruitment channels differed, the study used one Udio system, and large language models assisted the coding. The findings cannot prove a universal national difference. They can support the narrower point that the platform’s efficient prompt vocabulary and the listener’s descriptive vocabulary are culturally situated. An English-centred metadata system can make one way of describing music more operational than others.
Coelho: close vocabulary for unstable output
Guilherme Coelho’s “Latent Music: Emergent Sonic Forms and Sonic Liminality in Text-to-Audio Systems” reports many hundreds of Udio generations and interprets a selected subset. “Prompt assemblage” names the pull among descriptors. “Genre slippage” describes signs that migrate without settling into a stable synthesis. “Ontological suspension” names the listener’s uncertainty about the kind of object or source that is present.
These concepts can sharpen the micro description, especially if a clip changes its genre relation over time. The method remains phenomenological. The paper does not report the total retained share, negative cases, blind listeners, independent coding, or a repeatability estimate. It cannot establish how Udio internally represents a genre or how often the reported instability occurs.
Magnusson, Fabbri, Born, and Krogh: the macro spine
Magnusson explains why an interface classification is an embodied musical theory. Fabbri distinguishes a musical event from a socially accepted genre rule system. Born places the object in a distributed assemblage of people, archives, technologies, institutions, and later uses. Krogh shows that platform categories can organise culture and become productive rather than merely descriptive. Together, these sources permit a continuous argument from exact prompt syntax to musical form, then to cultural classification and power.
Course-material map
The class materials support the argument through a sequence rather than one detached concept:
- Mediation, Affordances, and Interface supplies the first analytical move. The prompt field makes some musical relations visible, easy, and repeatable while concealing the technical process that produces them. Session 3’s compositional, behavioural, listening, aesthetic, and temporal affordances can structure the interface paragraph.
- Session 4, Presets, Sampling, and Sonic Signs supplies the historical bridge. Presets, samples, names, and categories store technical design and cultural memory. The genre prompt extends this representational practice from sound selection to whole-song generation.
- Sessions 10 and 11 supply the platform scale. Metadata, functional classification, recommendation, and circulation form feedback loops in which platforms shape production and listening.
- AI Mediation, Delegation, and Emergence supplies the account of current musicianship. Session 12 treats prompting as intersemiotic mediation and repeated listening as cartographic work through delegation, direction, and selection.
- Session 13 supplies the dispositif and semiotic-labour frame. Model, interface, company, discourse, economics, regulation, user practice, and listening act together. The course theory slides make the sequence explicit on pages 32, 37, and 43: metadata becomes a creative constraint, prompting shifts labour from making sound toward making meaning, and musicianship includes designing conditions through language and selection.
The course chain is therefore: sign and interface -> affordance and selection -> model tendency -> platform infrastructure -> cultural rule and power. It provides the essay’s movement from micro to macro. Original papers carry the load-bearing academic claims where available.
Rough essay outline: Udio from prompt to perception
Target length: approximately 1,450 words before references. Park and Arthur provide the main micro evidence. Coelho supplies one shorter mixed-prompt case, while Fabbri, Magnusson, Born and Krogh support the move to cultural analysis.
1. Opening: genre in the prompt, 150 to 180 words
Open with Park and Arthur’s asymmetry. Genre appears in about 95 percent of their Udio prompts and about 61 percent of the listener descriptions. The word enters as a dominant instruction and returns as one part of a broader account of the sound. Introduce the question of how Udio translates genre words into audible signs and why this process does not create a genre by itself. State the thesis.
2. Genre as interface knowledge, 190 to 230 words
Describe Udio’s prompt interface and the company’s treatment of genre, mood, era and instrumentation as compositional terms. The interface provides semantic conditioning without exposing an internal map. Use the course account of presets as cultural infrastructure and Magnusson’s epistemic instrument. Both frames show how software makes selected musical categories easy to think, combine and repeat. Define the lexical, computational, sonic and social relations before using the word “between.”
3. Park and Arthur’s prompt-audio-description study, 320 to 370 words
Explain the sample of 200 real prompts, the paired 32-second Udio clips, the 148 listeners and the 2,624 free descriptions. Present the seven-part taxonomy and the main proportions. Genre dominates the prompt vocabulary, but instrumentation dominates listener descriptions. Genre also propagates more reliably than narrative language through related terms and audible signs. Analyse this as intersemiotic mediation: Udio turns a compact category into sound, and listeners reconstruct the sound with a wider vocabulary. State the limits concerning one platform, recruitment differences and assisted coding.
4. Mixed prompts and genre slippage, 190 to 230 words
Park and Arthur do not isolate mixed genres. Use Coelho’s selected Udio cases for this smaller aesthetic question. Describe the dense prompts, repeated extension and his account of musical signs that emerge and recede. Attribute every listening judgment to Coelho. His “genre slippage” concerns unstable perception within an output, while “prompt assemblage” concerns the relation among descriptors. His selection method can document a possibility but cannot estimate frequency or listener agreement.
5. Genre as a distributed model concept, 150 to 190 words
Use the MusicGen intervention study as limited technical support. Its genre-to-genre interventions work less effectively than instrument substitutions and alter audio beyond the desired concept. The result suggests that one open model represents genre through several entangled features. It does not reveal Udio’s architecture. This finding supports the conceptual move away from genre as a single slider or clean point.
6. Musical event and genre rule, 230 to 270 words
Use Fabbri to distinguish an event at the intersection of existing genres from a socially accepted genre rule. His formal and technical rules keep the sound inside the analysis, while his semiotic, behavioural, social, ideological, economic and juridical rules show what the Udio studies leave untested. Separate an interstitial clip, a recurrent platform style and a genre-forming tendency. Recurrence, naming, imitation, circulation and contestation provide possible evidence of social acceptance.
7. Platform power and cultural limits, 250 to 300 words
Expand from the listener descriptions to the classification feedback loop. Krogh shows how platforms make categories consequential through retrieval, recommendation and promotion. Text-to-music makes them productive during composition. Park and Arthur’s exploratory language comparison raises a cultural problem because English-centred prompt systems can make one descriptive register more operational than others. Born distributes agency across prior musicians, archives, developers, users, listeners and platforms. The resulting question concerns who controls the conversion of collective genre histories into proprietary controls.
8. Conclusion: between genres, before genre, 120 to 150 words
Answer the research question with the measured result. Udio translates genre terms into audible signs that listeners often recover through related words, instruments and vocal behaviour. Coelho documents selected cases in which several references remain unstable. These findings establish mediation and interstitial events. They do not establish an objective midpoint or a new genre. AI changes the route through which a category may form because a platform can generate, label, publish and measure the same material. Genre formation still requires later social use.
Earlier research decision: published interpolation experiments
The concrete problem
In 2018, nineteen electronic-music practitioners listened to transitions between IDM, electro, and techno. They distinguished a prepared neural-interpolation condition from a prepared crossfade condition in 101 of 114 trials. Fifteen practitioners then built their own interpolations and rated their experience of the process and result highly for originality, surprise, usefulness, and value. The conditions were audibly distinguishable and the tool was valued under the study’s setup. Nobody was asked whether a midpoint belonged to both genres, neither genre, or a new genre Borghuis et al. 2018.
That missing question is the essay. The studies establish different things: condition discrimination and unblinded process-and-result ratings in Borghuis, classifier uncertainty in Tokui, and one author-listener’s selected examples of instability in Coelho. None tests human consensus that an output is genre-ambiguous, much less that a genre has come into existence. A model can produce a candidate difference before musicians, listeners, critics, platforms, and institutions decide what that difference is.
Assignment fit
Essay 4 asks for a critical analysis of AI as a mediating force in twenty-first-century music. The brief explicitly proposes latent-space exploration, stylistic interpolation, latent aesthetics, intersemiotic translation, distributed creativity, prompt-based composition, and cartographic latent-space musicianship. It requires a focused system, practice, output, or phenomenon; a clear thesis and conceptual frame; movement from micro analysis to macro reflection; critical evaluation; a conclusion; at least two sources; and at least 1,000 words.
This topic meets that contract without surveying “AI and music” in general:
- Focused phenomenon: AI-generated movement between existing genre categories.
- Micro level: the dataset, representation, interpolation path, interface, outputs, listening task, and selection procedure in three concrete studies.
- Macro level: how a musical event becomes a genre through rules, repetition, mediation, classification, circulation, and institutional power.
- Critical question: what each experiment actually measures, and what has been smuggled into the word genre before the experiment begins.
The assignment PDF gives 15/08/25, apparently a stale year. The live ISIS deadline previously checked for the course is 15 August 2026 at 23:59.
Final research question
What can AI experiments actually establish about the musical space between genres, and what still has to happen before an interstitial output becomes a genre?
Shorter version: Can AI generate a new genre, or only material from which one might later be made?
Working thesis
AI systems can generate coherent and useful musical events between established categories, but they do not create genres through interpolation alone. The apparent territory between genres is a model-specific construction shaped by a corpus, representation, metadata, objective, interface, and path through the model. Its outputs may be genuinely novel as musical events and unstable as classifications. Genre formation begins when a community accepts or codifies such a difference as a rule; recurrence, naming, circulation, and classification are possible evidence of that process rather than a mandatory checklist. AI systems and platforms participate in the process, sometimes forcefully, but no isolated output or uncertain classifier found in the literature establishes its result.
The deliberately narrow claim is stronger than either easy extreme. “The model only copies” cannot explain useful emergent transitions. “The model invented a genre” confuses a new event with the social history that would make the event generic.
Title options
- Between Genres, Before Genre: What AI Ambiguity Can and Cannot Invent
- A Map Made of Drums: AI Interpolation and the Social Life of Genre
- When the Classifier Has No Name: Latent Music Before Genre
- The Space Between Genres Is Not Empty
Current preference: Between Genres, Before Genre: What AI Ambiguity Can and Cannot Invent.
The argument in one page
Four levels that must not collapse into one another
- Representational space: a model encodes selected musical properties in a mathematical representation. This space is built from a corpus and a technical objective. It is not culture laid out objectively on a map.
- Generated event: interpolation or prompting produces a rhythm, note sequence, timbre, clip, or song. It may be coherent and unfamiliar; whether it differs from every training item requires a separate corpus comparison.
- Genre attribution: a person, classifier, vendor, critic, or platform assigns or withholds a category. Disagreement can reveal ambiguity, a poor representation, an unfamiliar event, or all three.
- Genre formation: some differences acquire socially accepted rules across formal and technical, semiotic, behavioural, social and ideological, and economic and juridical practice.
Most AI papers supply evidence at levels one and two. Tokui aims at level three by training with a genre-ambiguity objective, although the paper does not report generated posterior probabilities from a validated genre head. Coelho’s author-listener account studies an experience of level three through “genre slippage” and “almostness.” None of the located studies follows level four.
Claim chain
- “Between” is never neutral. It depends on which dimensions a model represents, which examples become endpoints, and which path connects them.
- Latent interpolation can nevertheless do more than change the amplitude of two endpoints. It can decode intermediate points into changing patterns, several of which Borghuis and colleagues observed outside their training points in a PCA projection.
- Practitioners can discriminate the prepared interpolation and crossfade conditions and value the interpolation process and result. Borghuis and colleagues provide the clearest direct evidence, although their design does not isolate the latent path as the sole audible cause.
- A generator can be rewarded for pushing a genre head toward equal probabilities. Tokui implements that objective and measures onset-distance divergence, but does not report posterior entropy for generated patterns or human genre perception.
- Coelho provides a careful vocabulary for hearing instability in contemporary text-to-audio output, but his selected Udio cases are phenomenological examples rather than a controlled demonstration.
- Fabbri prevents the leap from unstable sound to genre. A musical event may sit at the intersection of several genres; genre concerns socially accepted rules, not style alone.
- Born makes innovation temporal and mediated. Later practice can establish an object’s generative force; a present originality judgment cannot. Creative agency is distributed across people, technologies, archives, institutions, and time.
- Krogh complicates a romantic community-only account. Platforms can operationally produce categories from above, turning genre into numerical distance and using classification to enact new groupings.
- The defensible answer is therefore a genre-forming candidate. The shorthand pre-generic describes the current evidence, not an inherent property of the sound: AI can make candidate differences and alter the conditions under which genres form, while the literature stops before any community accepts them as rules.
Evidence audit: experiments closest to the question
1. Borghuis et al.: practitioners use a path between three genres
“Off the Beaten Track: Using Deep Learning to Interpolate Between Music Genres” is the closest experimental predecessor and should be the essay’s central case.
Design: One author, described as a professional musician, composed 1,782 four-bar drum patterns as representations of IDM (608), electro (690), and techno (484). Every pattern used six TR-808 voices, sixteenth-note quantisation, MIDI velocity, and an intended tempo of 129 BPM. A four-dimensional LSTM variational autoencoder encoded selected start and goal patterns. Spherical interpolation joined their latent codes, and the decoder produced a sequence of new patterns. The system ran as an Ableton Live and Max for Live instrument.
Human results: In the identification study, nineteen practitioners with two to thirty years of electronic-music experience compared neural interpolations with ordinary audio crossfades. They made thirteen errors across 114 trials, or 101 correct identifications. In a separate creative task, the prose reports fifteen practitioners choosing endpoints and making transitions themselves; the Figure 7 caption says n = 16, an internal inconsistency. Participants rated their experience of the process and result together. On seven-point Creative Product Analysis scales, mean ratings included originality 6.13, surprise 5.94, usefulness 6.13, value 6.50, organicness 5.82, and well-craftedness 6.06. A third study with thirty-eight participants rated a GAN drummer as competent at 4.24 out of 5 but lifelike at only 3.13.
What this establishes: Under the prepared conditions, practitioners reliably discriminated interpolation from crossfade and gave high unblinded ratings to their experience of making and hearing interpolations. This is stronger evidence than an author reporting that a midpoint sounds interesting, while remaining narrower than a clean causal test of the latent path.
What it does not establish: The listeners were shown the interpolation-versus-crossfade hypothesis visually, received a practice pair, and were asked for condition discrimination rather than genre classification. Interpolation endpoints were VAE reconstructions while crossfades were produced in Logic; the paper does not document matched reconstruction endpoints and rendering, so several audible causes remain entangled. The creative-task ratings had no competing instrument or control condition. The study reports no naming, imitation, release, circulation, or longitudinal follow-up. It reduces genre to drum pattern, brackets harmony, timbre, production, form, dance practice, and identity, and uses a small corpus written by one person. The “map” is one expert’s formalisation of three EDM rhythm vocabularies. It was made for the experiment; it was not discovered beneath musical culture. The paper is an arXiv preprint rather than a clearly published peer-reviewed article.
Argumentative use: This is the positive hinge. AI can produce musically useful material between defined endpoints. The essay does not need to deny that result in order to deny that a genre has already formed.
2. Tokui: a generator aims at genre ambiguity directly
Nao Tokui’s “Can GAN Originate New Electronic Dance Music Genres?” asks almost the exact essay question and supplies the cleanest negative lesson.
Design: Tokui selected 1,454 commercial Groove Monkee MIDI files across nine vendor-labelled categories: Breakbeats, drum and bass, Downtempo, Garage, House, Jungle, Old Skool, Techno, and Trance. Each example became a two-bar grid of nine drum classes by thirty-two sixteenth-note positions with velocity. Quantisation discarded microtiming. A preliminary standalone bidirectional LSTM classifier, similar to the later genre head, reached 93.09% validation accuracy. Creative-GAN then jointly trained a real/fake discriminator and a genre head while rewarding the generator for pushing that head toward approximately equal category probabilities. The paper does not report held-out accuracy for this jointly trained head or its posterior probabilities on generated patterns. Five hundred outputs from the converged model were evaluated using edit distance.
Result: Under Tokui’s summed nine-instrument onset edit-distance proxy, generated outputs sat farther from every training genre than ordinary within-corpus examples while remaining closer to the corpus than a random-onset baseline matched only on onset-count mean and standard deviation. The experiment demonstrates an ambiguity-training objective plus this distance divergence; it does not measure a validated classifier hesitating over the generated outputs.
Critical reading: Even measured classifier uncertainty would not identify a new aesthetic category. It could mean that the generator found a decision boundary, exploited a classifier weakness, or combined features omitted by the labels. Here the ambiguity is first a training target, while the reported evaluation is distance. The onset edit distance ignores velocity and says little about groove, quality, use, or historical significance. The in-training discriminator and random-onset baseline are weak realism checks, not listener evidence. Tokui explicitly leaves the question “are these rhythm patterns really good?” to a future listener study.
Argumentative use: The experiment shows how category ambiguity can be made into an optimisation target; it does not establish that a competent classifier or human listener actually withholds a label. It reaches toward the edge of classification and stops precisely where perceptual evidence and genre theory have to begin.
3. Coelho: contemporary text-to-audio “almostness”
Guilherme Coelho’s “Latent Music: Emergent Sonic Forms and Sonic Liminality in Text-to-Audio Systems” moves the question from symbolic drum patterns to Udio’s full-band audio.
Design: Coelho reports many hundreds of Udio v1 and v1.5 generations across prompts, continuations, input audio, remixes, and similarity settings, then selects the outputs that repeatedly produced “almostness.” One case starts with john oswald plunderphonics experimental, deconstructed, classical music, stockhausen, no vocals and extends the first thirty-two seconds repeatedly with the same prompt. A second case describes two single-generation objects produced from differently dense assemblages: one prompt contains more than twenty-five descriptors; the other uses glitch, plunderphonics, electroacoustic, experimental, deconstructed, indeterminacy, classical music.
Conceptual value: “Prompt assemblage” names the relational pull among descriptors. “Genre slippage,” “gradient identities,” “ontological suspension,” and “semiotic drift” describe outputs whose signs emerge and dissolve before settling. “Semiotic cartography” concerns practical navigation through prompt signs and taxonomies; Coelho’s “semiotic musicking” concerns the listener’s construction of meaning from the unstable result.
Methodological limit: The total sample, retained proportion, exclusion criteria, and negative cases are not reported. One author selects and interprets the examples with prompts visible. There are no blind listeners, independent coders, controls, acoustic measures, or estimates of repeatability. Udio’s coordinates, training data, and prompt expansion are undisclosed. Cartography is therefore an interface-level and phenomenological metaphor, not an observation of the system’s actual manifold.
Argumentative use: Coelho supplies the micro-aesthetic language for what instability can sound like. Borghuis and Tokui supply more formal experimental checks. None should be asked to do the other’s job.
4. Zheng, Xambó Sedó, and Bryan-Kinns: repeatability supports emerging technique
“Exploring Gestural Affordances in Audio Latent Space Navigation” suggests how latent navigation can support emerging instrumental technique. A RAVE model trained for two million steps on two hours of dry guitar plucks exposed an eight-dimensional representation through a two-dimensional tablet terrain. Six ninety-minute workshops recruited eighteen musicians; twelve usable datasets remained after opt-outs, lateness, and a crash. Participants located sonic material, repeated trajectories, made loops, and treated coordinates as an XY pad. Short-term repeatability helped them turn surprise into a usable strategy.
The study is small, short-term, and not about genre. Its analytical value is narrower: returning to a location or trajectory can support emerging technique within a workshop. It does not establish a mature repertoire or shared long-term practice, just as a surprising coordinate does not establish a genre.
Supporting experiment matrix
| Study | What was tested | Result worth keeping | What it cannot support |
|---|---|---|---|
| MusicVAE, Roberts et al. 2018 | Spherical interpolation among symbolic sequences; 1,024 endpoint pairs | Hierarchical latent midpoints remained plausible under a separate five-gram model while naive blends became less plausible | Its listening test assessed random samples, not interpolation or genre; smooth symbolic movement is not genre formation |
| NSynth, Engel et al. 2017 | Halfway interpolation between encoded monophonic instrument notes | Midpoints fused envelopes and harmonic structures rather than merely superposing waveforms | Author listening only; one note and timbral fusion do not establish a musical category |
| MusicLDM, Chen et al. 2023 | Beat-synchronous mixing during text-to-music training | Latent mixup reduced the proportion of generated clips above .95 CLAP similarity to a training neighbour from .047 to .020 while broadly preserving quality | “Novelty” means lower nearest-neighbour similarity; fifteen listeners rated quality, relevance, and musicality, not novelty or genre |
| Controllable Music Production, Levy et al. 2023 | Diffusion regeneration of 100 randomly paired song transitions | Mel-distance changed approximately linearly across a 2.5-second transition, as it also did for ordinary crossfade and VAE interpolation | The result does not make diffusion uniquely smooth, and smooth song transition does not demonstrate an interstitial genre identity |
| Combining Audio Control and Style Transfer, Demerlé et al. 2024 | Song structure from one source rendered toward a jazz, dub, lo-fi hip-hop, or rock target | The full model improved structural cover identification and target-genre classification over ablations and MusicGen | Success was defined as fidelity to an existing target genre; conflicting rhythms produced chaotic output and were engineered away |
| Stable Audio Open, Evans et al. 2024 | Open text-to-audio architecture and corpus construction | The autoencoder and diffusion transformer were trained from scratch on 486,492 Creative Commons recordings, about 7,300 hours, with conditioning from a pretrained T5-base encoder; metadata strings included tags and genres | It does not reveal Udio’s architecture and evaluates realism and prompt alignment rather than genre novelty |
The literature uses novelty for several different measurements: distance from training neighbours, classifier uncertainty, a plausible latent midpoint, a participant’s originality rating, or an author’s sense of almostness. These are not interchangeable. The essay’s analytical work is partly to stop them from quietly becoming one grand claim.
What no located experiment demonstrates
- Listeners converging on a new label for interpolated output.
- Musicians imitating a recurrent interstitial grammar across separate works.
- A classifier-ambiguous output becoming stable under a later classifier trained on community usage.
- Circulation through releases, playlists, clubs, criticism, fandom, or teaching.
- A longitudinal change from isolated event to accepted rules.
- A community attaching identity, value, conflict, or historical memory to the proposed category.
That absence does not prove AI could never participate in genre formation. It sets the boundary of the present evidence.
Theoretical framework
Coelho: a vocabulary for the micro event
Use Coelho to describe the output without overstating the mechanism. “Genre slippage” is more precise than “genre fusion” when signs migrate during one object rather than settle into a reproducible mixture. “Ontological suspension” names the listener’s inability to decide what kind of instrument, voice, or genre is present. “Prompt assemblage” prevents a simplistic account in which each word controls one independent slider.
Treat “latent music” as an aesthetic condition and analytical term, not as the new genre the essay is trying to prove. Treat “spectreme” as Coelho’s proposed concept for a distributed trace of prior music, not as an established technical entity. Use at most three of his neologisms in the final essay: prompt assemblage, genre slippage, and ontological suspension are sufficient.
Fabbri: the distinction between an event and a genre
Franco Fabbri defines a musical genre as “a set of musical events (real or possible) whose course is governed by a definite set of socially accepted rules” in “A Theory of Musical Genres: Two Applications”. His definition does three jobs here.
First, an event can sit in the intersection of several genres and belong to each. An interstitial output therefore does not force a new genre into existence. It may be an unusual event under several existing rule systems.
Second, style and form are only part of genre. Fabbri’s rules are formal and technical, semiotic, behavioural, social and ideological, and economic and juridical. A drum grid captures one thin slice of the object it names.
Third, new genres arise inside structured musical systems. A transgression becomes genre-forming when people use it as a model and codify it as a rule. “Success” is not proof of aesthetic value; it is a community’s codification of expectations. This supplies the missing temporal mechanism between one AI event and a genre.
Avoid turning Fabbri into a rigid checklist of a large scene, a settled name, and decades of history. He explicitly allows small or discredited communities, possible events, and aesthetic manifestos. The safer claim is that genre requires some socially consequential acceptance and reproduction of rules, however small or new the group may be.
Born: mediation, distributed creativity, and the test of time
Georgina Born’s “On Musical Mediation: Ontology, Technology and Creativity” supplies the macro frame. Music exists through assemblages of sonic, discursive, visual, artefactual, technological, social, and temporal mediations. This avoids both technological determinism and the fiction of a neutral tool. In the central case, the musical object is mediated by a musician’s genre examples, MIDI quantisation, an encoder, an interpolation rule, a decoder, Ableton Live, a chosen sound set, recruited practitioners, and the language of a research paper.
Born’s temporal account is decisive. Later practice can establish that an object generated further possibilities and expectations, but she warns that this relation neither exhausts nor guarantees innovation: some possibility-opening works are not innovative, and some innovative works fail to generate further “protentions.” The high originality score in Borghuis et al. describes a moment of reception. It cannot answer the historical question.
Born’s distributed creativity locates agency across the experiment’s human and technical assemblage more accurately than “the AI invented.” Her relayed creativity becomes relevant when generated material is reopened, recomposed, circulated, and transformed through sequential handoffs. The model makes a difference, but so do the musician who authored the corpus and the people who selected endpoints, chose a path, operated the interface, rendered the MIDI, listened, named, and reused the result.
Krogh: the platform can genre-fy from above
Mads Krogh’s “Rampant Abstraction as a Strategy of Singularization: Genre on Spotify” is the strongest counterargument to a simple story in which only an organic scene can form a genre. Within MIR-based classification, Spotify and The Echo Nest participate in a shift from qualitative musico-cultural worlds toward statistical correlations and numerical distances in a continuous, multidimensional similarity space. By 2022, Every Noise at Once represented Spotify’s underlying genre space with close to 6,000 entries, far more than the user-facing interface displayed. Content and contextual data can be used to identify, “seed,” and promote categories that were not conventional scene genres, including functional groupings organised around activity and mood.
Classification here is productive. A platform does not merely find a genre; playlists, interfaces, recommendation, and marketing can make a grouping consequential. This matters because future text-to-audio platforms could generate, label, cluster, recommend, and monetise an interstitial category before a local scene agrees on what it is.
The essay should therefore distinguish two kinds of classification:
- Operational platform category: a grouping organises retrieval, curation, recommendation, advertising, or promotion, whether or not it constitutes a musico-cultural genre.
- Musico-cultural genre: a rule system becomes meaningful across production, reception, identity, discourse, history, and institutions.
The two can feed one another. A platform label can help assemble a public, while musico-cultural trends, listener data, and scraped language feed back into platform classification. The distinction is analytical, not a claim that platform categories are fake. It prevents the essay from granting a corporation the final word merely because its database has a row.
Supporting frames: use briefly
- Jennifer Lena and Richard Peterson define genres as systems of orientations, expectations, and conventions connecting industry, performers, critics, and fans in “Classification as Culture”. They separately observe that many lone musical experimentalists go unheralded. Every genre in their explicitly nonrepresentative, nonprobability sample of sixty twentieth-century US commercial genres passed through a scene phase. The authors exclude many art and nonprofit musics and warn that the pattern may partly reflect retrospective historical storytelling, so scene remains supporting evidence rather than a universal law.
- Nikhil Singh, Manaswi Mishra, and Tod Machover’s “AI for Musical Discovery” draws on Boden’s combinational, exploratory, and transformational creativity. The authors judge current systems readily able to support the first two and say it is presently difficult to see an autonomous route to the third, while arguing that human–AI collaboration may alter a musical conceptual space. This is a position, not an experimental result. Their larger idea of discovery includes learning, fresh perspective, and community rather than output novelty alone. Use this once to classify the experiments, not as a second definition of genre.
- Eliot Bates’s “Actor-Network Theory and Organology” treats agency as the capacity to make a difference rather than the possession of intention. It can help depunctualise “the AI” into corpus, model, interface, prompt, user, listener, and institution. Born already performs most of this work and deals better with history, so Bates is optional.
Alternative published-case route
Case decision
Use Borghuis et al. as the focused system and practitioner experiment, Tokui as the designed ambiguity countertest, and Coelho as the contemporary full-audio phenomenological extension. This is one coherent case family rather than three unrelated examples: each makes “between genres” operational in a different way.
- Borghuis: between as a path joining selected genre endpoints.
- Tokui: between as uncertainty across a trained genre classifier.
- Coelho: between as a listener’s unstable encounter with full audio generated from words.
The three operationalisations do not converge automatically. Their disagreement is analytically productive. It shows that latent-space cartography is never one map.
No new Udio output was generated for these notes, and no sonic observation has been invented. Published experiment auditing is currently the empirical method. An original text-to-audio pilot remains possible, but the final essay already has a defensible micro case without spending platform credits or selecting a convenient strange output after the fact. The strongest remaining micro step is to listen closely to one published Coelho case object, add one timestamped observation, and keep it subordinate to the Borghuis experiment rather than opening a fourth case.
Analytical questions to ask of every experiment
- What does the study call a genre, and who supplied the labels?
- Which musical dimensions enter the representation, and which disappear?
- What makes two endpoints near, far, or opposed?
- Is the path a straight line, a spherical interpolation, classifier uncertainty, or a prompt relation?
- What evidence establishes coherence, novelty, ambiguity, or usefulness?
- Are listeners blind to the prompt and hypothesis?
- Are failed and ordinary outputs reported alongside selected successes?
- Does the study test a sound, a tool, a category, or a social practice?
- What later event would falsify the claim that a genre is forming?
Alternative detailed essay outline: interpolation studies
Target: 1,470–1,780 words excluding references. The micro analysis remains the largest part, but the theoretical turn is substantial enough to satisfy the assignment rather than appearing as a sociological disclaimer at the end. Borghuis remains the primary case; Tokui is the countertest and Coelho the short full-audio extension.
1. Opening: 101 correct answers and no name — 140–170 words
Open with the Borghuis identification result. Nineteen practitioners hear 114 paired transitions and identify the prepared neural-interpolation condition rather than the prepared crossfade condition in 101 cases. The creative-task participants later rate their experience of the process and result as original, surprising, useful, and valuable. Then state the omission: nobody is asked what the midpoint is.
Move directly to the research question and thesis. Do not begin with a general claim that AI is transforming music or that culture has exhausted every genre. The concrete paradox is enough: the study produces distinguishable conditions between genre endpoints, yet it supplies no evidence that the generated difference has become a category.
Possible opening:
Nineteen electronic-music practitioners heard six pairs of transitions. In one condition, two drum patterns were crossfaded. In the other, a neural network generated the intermediate patterns. The practitioners chose correctly 101 times out of 114. The experiment never asked them what the midpoint was.
End the introduction with the four-level distinction in compressed form: representation, event, attribution, formation.
2. The map is constructed — 170–210 words
Explain latent space in plain terms. An encoder places patterns into a compressed representation; an interpolation rule chooses intermediate coordinates; a decoder turns those points back into patterns. The user does not travel through culture itself. The user travels through what this corpus, architecture, objective, and representation have preserved.
Use the Borghuis dataset to make this material. The model knows 1,782 four-bar patterns written by one musician, played with six TR-808 voices at 129 BPM. The representation does not explicitly encode or test club history, fashion, venues, race, class, record distribution, or the argument over what “IDM” means, although rhythmic patterns can still carry historical and social traces. Its topology is musically consequential and remains an engineered proposition about three genres.
Introduce Coelho’s “semiotic cartography” as a qualified practice: musicians infer a changing map through repeated encounters. Keep the technical and phenomenological meanings of latent space distinct.
Transition: once the modest meaning of the map is fixed, the experiment’s positive result becomes easier to value without turning it into mythology.
3. What latent interpolation actually adds — 230–270 words
Describe the spherical interpolation and Ableton interface. A crossfade overlays two fixed endpoints while changing their amplitude. The VAE decodes each intermediate point into a changing drum pattern. The identification study shows that the prepared conditions were distinguishable in a musical context; because reconstruction and rendering were not fully matched, it does not isolate the latent path as the sole audible cause.
Then analyse the creative-task ratings. Originality and surprise matter, but usefulness and value are more interesting. They indicate that practitioners regarded their combined process-and-result experience as material for work, not only as a novelty demonstration. This is where Zheng’s workshops can enter in one sentence: even within a short workshop, returning to latent locations and trajectories can turn surprise into an emerging technique.
Add the methodological brake in the same paragraph rather than saving every limitation for later. The creative task has no control instrument; the samples are small and locally recruited; the endpoint corpus was authored by one person; the paper never asks for a genre judgment. The result supports a useful genre-forming candidate. Calling it pre-generic describes the limit of the evidence, not the sound’s destiny.
Transition: hearing a path is still different from hearing a new place.
4. When the classifier cannot decide — 190–230 words
Introduce Tokui’s stronger intervention. The Creative-GAN is explicitly rewarded when its jointly trained genre head distributes probability evenly across nine labels. The paper reports no held-out accuracy or generated posteriors for that head. Under a summed onset edit-distance proxy that ignores velocity, its 500 outputs are farther from every genre set than ordinary examples and closer to corpus examples than a random-onset baseline matched only on onset-count statistics.
Make the analytical distinction explicit: the experiment optimises an uncertainty objective under a particular genre head. It does not demonstrate that this target was reached under a validated classifier, much less discover an objectively empty region between genres. The vendor labels, quantised drum grid, learned features, and ambiguity loss construct the proposed boundary. A different representation or taxonomy would redraw it.
Use Tokui’s own unanswered question about whether the rhythms are good. The absence of listener evaluation blocks a claim about musical value; the absence of uptake blocks a claim about genre. An ambiguity objective may still be artistically useful, just as feedback and distortion were useful before anyone settled their meanings. It proposes a way to seek candidate difference, not category birth.
Optional one-sentence support: MusicLDM’s “novelty” metric likewise means reduced nearest-neighbour similarity, illustrating how quickly a technical proxy can outgrow its definition in discussion.
5. How instability sounds — 160–200 words
Move from symbolic drums to Coelho’s Udio cases. Briefly describe the repeated extensions and dense cross-genre prompts. Use “genre slippage” for identities that emerge and dissolve and “ontological suspension” for the listener’s inability to settle on an instrument, voice, or stylistic object.
Then expose the method. Coelho selected a subset from many hundreds of generations because it produced the experience his theory names. That is legitimate phenomenological concept formation and weak prevalence evidence. The examples show that one trained author-listener can find full-audio outputs worth describing as interstitial. They do not tell us how often this occurs, whether blind listeners agree, or whether the same prompt returns to the same tendency.
This paragraph prevents two mistakes: dismissing close listening because it is subjective, and treating selected listening as an experiment with a known hit rate.
Transition: at this point the essay has established three kinds of between-ness. None yet answers what a genre is.
6. A musical event is not yet a genre — 220–260 words
Introduce Fabbri’s definition and make it do the central conceptual work. A musical event can occupy the intersection of several genres. An AI output that sounds partly like techno, IDM, and something else may therefore remain an event governed by overlapping existing rules.
Explain that Fabbri’s rule system exceeds acoustic form: technical, semiotic, behavioural, social and ideological, and economic and juridical rules participate. The experiments capture formal and technical traces while deliberately excluding most of the rest.
Then give the diachronic mechanism. New genres start within existing systems. A transgression may become a model when a community accepts or codifies it as a rule. This is where the provisional term pre-generic earns its keep. It does not mean socially empty, aesthetically inferior, or condemned to remain outside genre. It names evidence that stops before codification.
Add Lena and Peterson carefully: they observe that many lone experimenters go unheralded, while scenes connect performers, audiences, critics, and industries. Every genre in their nonrepresentative sample passed through a scene phase; the study cannot establish that trajectory as a universal frequency or law. Virtual, tiny, platform-driven, and manifesto-based communities remain possible.
7. Mediation, time, and platform power — 240–290 words
Use Born to connect the micro event to a historical assemblage. The interpolated pattern is jointly conditioned by corpus author, genre taxonomy, representation, network, path, interface, sound set, operator, and listener. This is distributed creativity rather than a contest to locate one sovereign author. Relayed creativity enters if the output is subsequently reopened, recomposed, circulated, and transformed.
Use Born’s temporal argument to ask whether the object later generates possibilities and expectations, while retaining her warning that generativity and innovation are not identical. A high originality rating records present judgment; it cannot establish historical generativity or a future lineage. The essay can then infer, rather than attribute to Born, that cheaper variation may produce more abandoned novelties as easily as more genres.
Then bring in Krogh as the counterweight. Platforms can make classifications consequential through recommendation, playlisting, metadata, advertising, and promotion. A company does not have to wait for an organic scene before seeding an operational category. This means AI may participate in genre formation beyond sound generation: it could cluster outputs, supply labels, recommend examples, teach producers what is rewarded, and assemble listeners around the category.
Answer the counterargument without retreating. Platform enactment is part of genre’s social process, not proof that the sonic midpoint generated the genre alone. The macro question shifts from “did the model invent it?” to “which actors made this difference repeatable, visible, and valuable, and under whose category?”
8. Conclusion: before genre — 120–150 words
Return to the evidence. Borghuis’s practitioners discriminated prepared conditions and valued their process-and-result experience. Tokui optimised an ambiguity loss and measured onset-distance divergence. Coelho’s author-listener described selected full-audio identities flickering. These are real and different results.
State the answer plainly: AI can generate events from which genres might be made. It can also change the speed, scale, and institutional route through which categories are tested and imposed. The studies located and examined here do not show an AI-originated genre because they do not demonstrate a community accepting or codifying the proposed difference as a rule. Naming, repetition, imitation, circulation, and conflict would be useful evidence, not cumulative requirements.
End with a testable future condition rather than a slogan: the first persuasive demonstration will not be the strangest isolated output. It will be a documented sequence in which musicians and listeners return to a generative tendency, develop techniques and expectations around it, and make the category matter beyond the system that proposed it.
Counterarguments and answers
“Every genre is an arbitrary label. The model’s classifier is enough.”
Labels do real work, but different labels have different reach. Tokui’s classifier partitions a vendor-labelled drum corpus. Spotify’s classifications can organise discovery, recommendation, advertising, and promotion. A genre account should ask who uses the label, what it changes, and which musical and social practices it stabilises. The argument does not require an authentic folk origin; it requires consequential enactment.
“A single work can found a genre.”
Fabbri describes two distinct routes. Innovations in a successful event may later be used as a model and become rules; an aesthetic manifesto can codify rules for possible events in advance. For an isolated generated event with no such programme, “genre-founding” remains a hypothesis about later acceptance.
“AI only interpolates, so the output cannot be new.”
This confuses origin with identity. Borghuis’s decoder produces intermediate patterns through decoding rather than audio crossfade, and the authors observe several apparently new points in a PCA projection. They do not run an exhaustive nearest-neighbour or memorisation test. NSynth and MusicVAE provide related evidence at other musical scales. A new event can emerge from inherited relations; the stronger critique concerns the claim made for that event.
“Humans have always made hybrid genres, so AI changes nothing.”
Hybridity is old. The change lies in the mediational environment: high-dimensional representations, stochastic generation, promptable taxonomies, cheap iteration, automated classification, and platform-scale circulation can alter which differences become available and how quickly they are tested. The essay does not need the false claim that AI invented musical mixture.
“If listeners cannot classify it, it must be new.”
Uncertainty has several causes: unfamiliarity, ambiguity, poor labels, missing features, a weak classifier, or a damaged output. Newness requires additional evidence. “Neither” is not yet the name of a genre.
“Social genre theory ignores the sound.”
Fabbri’s formal and technical rules keep sonic form inside the account while showing that it is insufficient on its own. The experiments become more valuable when their precise sonic scope is stated rather than inflated.
Extended listener study after the primary case
The primary prompt matrix supplies the essay’s close case. A later study could test its emerging claims with more outputs and independent listeners. Preregister it before listening so that “almostness” is not selected after hundreds of attempts.
Minimal design
- Choose two genre anchors with recognisable but non-identical rhythm, timbre, form, and production conventions.
- Conditions: anchor A; anchor B; A plus B; A plus B plus one destabilising term; and one dense Coelho-style prompt.
- Generate at least eight outputs per condition with fixed platform, model/version, date, duration, settings, and seed where exposed.
- Preserve every output, exact prompt, automatic prompt expansion, extension, remix, and rejection.
- Randomise filenames before listening. Assign prompt-hidden and prompt-visible conditions to independent listener groups, or use disjoint, counterbalanced stimulus sets; do not have every listener hear the same output blind and then labelled in a fixed order.
- Preregister the codebook and agreement statistic. Have at least one second listener code a substantial preregistered subset independently, then report categorical agreement and scale reliability.
Listening codebook
- Primary attribution: A, B, both, neither, or unclear.
- Stability: stable category, stable hybrid, or category changes over time.
- Coherence, novelty, and usefulness on separate seven-point scales.
- Timestamped genre slippage, timbral liminality, structural rupture, and obvious generation defects.
- Recurrence: whether the same tendency appears across outputs rather than in one selected hit.
Stronger later study
Compare endpoint tracks, ordinary crossfades, latent interpolations, and text-prompt hybrids under blind listening. Ask participants for A/B/both/neither judgments, free labels, quality, novelty, and willingness to reuse. Test prompt visibility between independent groups or with counterbalanced disjoint stimuli. Then give musicians the system over several weeks and observe whether stable names, gestures, production techniques, and shared examples develop. Only the longitudinal stage reaches genre formation rather than immediate classification.
Claim discipline
- Do not claim that culture has tried every genre. Describe catalog saturation or the feeling produced by exhaustive platform taxonomies.
- Do not call latent space “every possible song.” It is a finite model’s learned representation of selected data under a particular objective.
- Do not say the user sees or directly navigates Udio’s coordinates. Prompt cartography is inferred through an opaque interface.
- Do not treat Casini et al.’s maps of prompt and tag embeddings as maps of Suno’s or Udio’s internal representations.
- Do not infer that two genre tokens are endpoints in a continuous musical geometry. They are instructions whose internal processing is undisclosed.
- Do not claim that one feature was caused by one prompt word after hearing a single clip. Use matched conditions and state the level of repetition.
- Do not claim that Udio’s example of
eurodance + black metalproduces a hybrid, an interstitial style, or a genre. The company documents the possibility of asking, not the sonic result. - Do not generalise Park and Arthur’s Udio findings to all text-to-music systems. Treat the Korean and English comparison as exploratory rather than a causal account of national listening cultures.
- Do not present Park and Arthur’s listener descriptions or Coelho’s listening judgments as our own close listening. Name the observer and study each time the source of the sonic claim could be unclear.
- Do not say Park and Arthur tested new genre formation or mixed-genre prompting. Their study tests the relation between prompting and description across a varied Udio sample.
- Do not infer Udio’s internal representation from the MusicGen intervention study. Use it only as evidence that genre behaved as an entangled concept in one open model.
- Do not equate mathematical midpoint, sonic hybrid, classifier uncertainty, listener surprise, and genre novelty.
- Do not report MusicLDM’s nearest-neighbour ratio as perceived creativity.
- Do not report MusicVAE’s listening test as evidence for its interpolations; listeners rated random samples.
- Do not say Coelho proved how frequent genre slippage is. His reported subset is purpose-selected, so it cannot establish prevalence.
- Do not say Borghuis represented EDM genre culture as a whole. The corpus contains one author’s quantised TR-808 patterns.
- Do not say Tokui’s output is good or humanly ambiguous. Those questions were not tested.
- Do not require a large, physical, long-lived scene. Require some consequential acceptance and reproduction of rules.
- Do not reduce platform categories to fake genres. Classification can organise attention and value and thereby help produce what it names.
- Do not anthropomorphise the system. Attribute agency as a specific difference made by corpus, model, interface, user, listener, or institution.
Source hierarchy
Core spine
- Udio, “Prompt Like a Master” and “Introducing Styles”. Primary interface evidence for tags, Manual Mode, genre combinations, references, and blending.
- Sangheon Park and Claire Arthur, “From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music” (2026 preprint). Prompt-output-listener evidence and exploratory cultural comparison.
- Luca Casini et al., “Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio” (2025). Large published-output corpus and prompting practice.
- Thor Magnusson, “Of Epistemic Tools: Musical Instruments as Cognitive Extensions” (2009), pp. 168–176. Software classifications, concealed musical theories, and interface responsibility.
- Franco Fabbri, “A Theory of Musical Genres: Two Applications” (1982). Event, rule, intersection, community, and codification.
- Georgina Born, “On Musical Mediation: Ontology, Technology and Creativity” (2005), pp. 7–36. Mediational assemblage, distributed and relayed creativity, temporal innovation.
- Mads Krogh, “Rampant Abstraction as a Strategy of Singularization: Genre on Spotify” (first published 2023; issue 2025), pp. 89–107. Continuous genre space and productive platform classification.
- Guilherme Coelho, “Latent Music: Emergent Sonic Forms and Sonic Liminality in Text-to-Audio Systems” (2026), pp. 122–129. Contemporary Udio phenomenology and vocabulary for unstable output.
Supporting sources
- Jennifer C. Lena and Richard A. Peterson, “Classification as Culture: Types and Trajectories of Music Genres”, American Sociological Review 73.5 (2008): 697–718. Genre systems, scenes, conventions, and trajectories; nonrepresentative commercial-US sample.
- Tijn Borghuis et al., “Off the Beaten Track: Using Deep Learning to Interpolate Between Music Genres” (2018). Practitioner evidence for symbolic rhythm interpolation under a prepared experiment.
- Nao Tokui, “Can GAN Originate New Electronic Dance Music Genres?” (2020). Category-ambiguity objective and onset-distance countertest.
- Marcel A. Vélez Vásquez, Charlotte Pouw, John Ashley Burgoyne and Willem Zuidema, “Exploring the Inner Mechanisms of Large Generative Music Models”, Proceedings of the 25th ISMIR Conference (2024), pp. 791–798. Genre and instrument interventions in three MusicGen sizes; classifier-based evaluation without a listener study.
- Nikhil Singh, Manaswi Mishra, and Tod Machover, “AI for Musical Discovery” (2024). Exploratory versus transformational creativity; position paper rather than experiment.
- Shuoyang Jasper Zheng, Anna Xambó Sedó, and Nick Bryan-Kinns, “Exploring Gestural Affordances in Audio Latent Space Navigation” (2025). Repeatable latent navigation as developing instrumental technique.
- Zach Evans et al., “Stable Audio Open” (2024). Technical grounding for text conditioning, learned audio representation, corpus, and metadata; architecture analogy, not evidence about Udio.
Background and triangulation
- Adam Roberts et al., “A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music” (2018).
- Jesse Engel et al., “Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders” (2017).
- Ke Chen et al., “MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies” (2023).
- Mark Levy et al., “Controllable Music Production with Diffusion Models and Guidance Gradients” (2023).
- Nils Demerlé et al., “Combining Audio Control and Style Transfer Using Latent Diffusion” (2024).
- Eliot Bates, “Actor-Network Theory and Organology” (2018), pp. 41–51.
- Guilherme Coelho, “The Artist Is Present: Evidence of Artists Residing and Spawning in Text-to-Audio AI” (2025). Useful evidence that apparent blank territory retains archive traces; selected results are consistent with artist-conditioned signals but do not prove that particular recordings were present in training.
Source and assignment corrections
- The course link labelled Stable Audio Open points to arXiv
2403.08997, an unrelated aerial RGB-thermal dataset. The correct paper is arXiv:2407.14358. - Coelho’s The Artist Is Present bibliography gives
2404.14677for Stable Audio Open; that identifier is also unrelated. Use2407.14358. - The course Bates link is a five-page 2017 conference paper. The Journal of the American Musical Instrument Society PDF is the expanded published version, volume 44 (2018), pp. 41–51; use that version if Bates remains.
- The linked Fabbri edition states that it was delivered at the First International Conference on Popular Music Studies in 1981 and first printed in Popular Music Perspectives in 1982, pp. 52–81. Cite this edition as 1982.
- Krogh was first published online in 2023 and appears in volume 19.1 dated 2025. Cite consistently according to the required style.
- Borghuis et al. is an arXiv preprint. Do not imply peer review.
Final writing test
The final draft is ready to write when every paragraph answers one of five questions: What entered the Udio prompt? What did the source study observe? What does its method establish? What does genre theory permit us to infer? Who or what could make the result socially consequential? Remove any paragraph that only describes latent space as vast, uncanny, or full of possibility.