Barcelona. Escola Superior de Música de Catalunya (ESMUC)
7-10 Sep 2026
CAMBIATA (Counterpoint And MusicXML By Instructing AI Through APIs): Generating Species Counterpoint with Large Language Models
Pierre Basso  1, 2@  , Liam Pond  1, 2@  
1 : Schulich School of Music [Montréal]
2 : Centre for Interdisciplinary Research in Music Media and Technology  (CIRMMT)

The present study investigates the capabilities of large language models (LLMs) to generate species counterpoint in accordance with Johann Joseph Fux's Gradus ad Parnassum (1725). Earlier approaches, from Schottstaedt (1984) and Lewin (1983) to the FuxCP project at UCLouvain (Wafflard 2023; Lamotte 2024), pursue an algorithmic path: the rules are encoded directly as conditions or search strategies. With LLMs, by contrast, the rules are conveyed in natural language and interpreted by the model. The procedure thereby comes closer to Fux's dialogic pedagogy.

For CAMBIATA (Counterpoint and MusicXML By Instructing AI Through APIs), we have written five teaching texts, one per species, that transpose Fux's approach into a form of language that is efficient for LLMs. Each guide formulates the rules of consonance and dissonance, types of motion, and cadential formulas in continuous prose, and ends with a self-verification step performed by the model. They are submitted, with one example drawn from Fux, to three LLMs (Claude Opus 4.6, GPT-5.3, Gemini 3 Pro); the cantus firmi too are taken from the Gradus.

A corpus of 450 generated solutions is examined both statistically by a custom analysis system and through close reading of selected scores. Three tendencies emerge: the third species is the most demanding, followed by the fifth; counterpoint above the cantus firmus contains fewer errors than counterpoint below, producing a clear asymmetry; surprisingly, 31 % of generations are musically identical to another run. The models contrast sharply: Gemini is the strongest performer (3.28 errors per solution on average) with exemplary dissonance treatment, ChatGPT the intermediate contender (3.97) with recurrent cadential errors, while Claude shows the highest average of mistakes (8.94) with systematic weakness in dissonance handling. In their most successful instances, the solutions are barely distinguishable from a human-written one. The paper evaluates the inferential capacities of large language models within a canonical object of computational music theory, and transposes Fux's dialogic pedagogy into a time when LLMs are becoming a defining tool for research and musical practice.



  • Poster
Loading... Loading...