A Practical Guide to Generative Music AI for Developers hero

Creative Coding

A Practical Guide to Generative Music AI for Developers

A hands-on, code-first journey through modern music AI — from audio features and VAEs to transformers and diffusion models — building working generative systems in Python. Includes guest sessions with Jordan Rudess and Christian Steinmetz.

Level

Advanced

Duration

36h 9m of video content

Format

Self-paced video

Watch a preview

An Introduction to AI Music (50')

Course overview

Over 12 sessions this course takes developers from the fundamentals of machine learning for audio through to the architectures powering today's generative music tools. You'll set up a Python environment, work with MIDI and spectrogram representations, and get hands-on with RAVE, EnCodec, Anticipatory Music Transformers, MusicGen and Stable Audio. Later sessions cover real-time inference (TorchScript, quantization, ONNX) and the commercial landscape, before you scope, build and present your own final project. Designed for developers comfortable with Python who want to understand and build music-AI systems rather than just use them.

Course content

Session 1: An Introduction to AI Music

2 videos, 1 resource, 2 lessons

+
  • Session 1 Recording
    Checking access...
  • Main Topics: AI Music Case Studies, Course Roadmap
    • Case study of the foundational tools behind the rise of Generative AI in music.
    • Exploration of the main themes of research and development in AI Music: Audio vs Symbolic, Language Models vs Diffusion, Composition vs Live Performance, General Public vs Artists
    • Outline of course objectives, final project, and portfolio goals.
  • An Introduction to AI Music (50')
    Checking access...
  • Session 1 PDF handout
  • Open Discussion
    • “What do you expect to get out of this course? What professional outcomes are you aiming for?”

Session 2: Setting up your environment

6 videos, 3 lessons

+
  • Session 2 Recording
    Checking access...
  • Main Topics: Environment Setup
    • Installing Python 3.12 and required libraries. Working with multiple environments on a single machine (pyenv, anaconda, venv).
    • Cloning the class repository from GitHub
  • GitHub Repository

    The GitHub repository can be accessed at lancelotblanchard/ai_music_course.

  • Setting up your environment - Cloning the class repository (10')
    Checking access...
  • Hands On
    • Using librosa and numpy to load audio, visualize waveforms, and extract simple features.
    • Using mido and fluidsynth to load and play MIDI data.
  • Setting up your environment - Hands On 1.1: Loading, visualizing, playing audio (10')
    Checking access...
  • Setting up your environment - Hands On 1.2: Extracting Audio Features, RMS and ZCR (13')
    Checking access...
  • Setting up your environment - Hands On 1.2: Extracting Audio Features, Spectrograms (13')
    Checking access...
  • Setting up your environment - Hands On 2: Manipulating MIDI Data (20')
    Checking access...

Session 3: Core Machine Learning Concepts for Music and Audio

6 videos, 2 lessons

+
  • Session 3 Recording
    Checking access...
  • Main Topics: Audio vs Symbolic Music, Basics of Generative AI, Data Acquisition and Ethics
    • Go over significant differences between audio and symbolic music.
    • Cover the statistical basics of generative modeling in artificial intelligence.
    • Understand how Autoencoders are used in musical representation and generation.
    • Learn how to legally source and curate music datasets (copyright, licensing).
    • Industry Relevance: Ensuring compliance in data collection. 
  • Statistical Basics of Generative Modeling in Artificial Intelligence (10')
    Checking access...
  • Variational Autoencoders (16')
    Checking access...
  • Hands On
    • Downloading a subset of Lakh MIDI dataset and Free Music Archive and analyzing examples.
    • Using and training RAVE models for new music generation.
  • Hands On 1.1: Lakh MIDI Dataset (12')
    Checking access...
  • Hands On 1.2: Free Music Archive (6')
    Checking access...
  • Hands On 2: Using RAVE
    Checking access...

Session 4: Real-Life Collaborations between Artists and Engineers with Guest Speaker Jordan Rudess

2 videos, 2 lessons

+
  • Session 4 Recording
    Checking access...
  • Main Topics: Human-Computer Interaction, Iterative Design, Continuous Deployment
    • Explore Human-Computer Interaction (HCI) principles behind the development of tools with user feedback integration.
    • Analyze continuous deployment strategies to ensure seamless updates and integration with artists’ workflows.
    • Case study of Jordan Rudess’s project with the MIT Media Lab: ‘Jordan and the jam_bot’.


  • Human-Computer Interaction & User-Centered Design
    Checking access...
  • Open Discussion with Jordan Rudess
    • “What are the most exciting applications of Generative AI in music creation?”
    • “Which challenges need addressing to optimize Generative AI’s role in music making?”
    • “How can we foster effective collaborations between artists and engineers in music technology advancements?”

Session 5: Representation Learning for Music

4 videos, 2 lessons

+
  • Session 5 Recording
    Checking access...
  • Deep Dive into MIDI & Spectrograms
    • Overview of the different representations used for music (e.g., tokenization: MusicGen for audio and Music Transformer for MIDI, spectrograms, chromagrams, latent representations).
    • Study of the trade-offs of each representation and the best situations to use them.
  • Comparing Musical Representations & Encodec Deep Dive (14')
    Checking access...
  • Understanding RVQ in Encodec (12')
    Checking access...
  • Hands On
    • Computing spectrograms from audio files
    • Creating sequences of tokens for MIDI files.
    • Using HiFi-GAN to synthesize spectrograms.
    • Using Encodec or DAC to compress and resynthesize audio.
  • Hands On: Encodec (29')
    Checking access...

Session 6: Autoregressive Music Generation

5 videos, 2 lessons

+
  • Session 6 Recording
    Checking access...
  • Main Topics: Autoregressive modeling, the Transformer architecture, HuggingFace Hub
    • Overview of Autoregressive Distribution Estimation and the Attention mechanism.
    • Walkthrough of the Transformer architecture.
    • Case study of Anticipatory Music Transformer for Symbolic Music Generation.


  • The Transformer architecture (15')
    Checking access...
  • Understanding Anticipatory Music Transformers (13')
    Checking access...
  • Hands On
    • Run a Music Transformer to generate new MIDI data.
  • Hands On: Using AMT to generate MIDI data (Part 1) (18')
    Checking access...
  • Hands On: Using AMT to generate MIDI data (Part 2) (17')
    Checking access...

Session 7: Autoregressive Music Generation (Part 2)

3 videos, 2 lessons

+
  • Session 7 Recording
    Checking access...
  • Main Topics: MusicGen & Audio Generation with Transformers
    • Case study of the different sampling and token modeling methodologies (Residual Vector Quantization and codebook patterns.
    • Case study of the MusicGen architecture.
  • Understanding MusicGen (8')
    Checking access...
  • Hands On
    • Using MusicGen to generate audio autoregressively.
  • Hands On: Using MusicGen to generate audio (38')
    Checking access...

Session 8: Diffusion Models for Music Generation

8 videos, 2 lessons

+
  • Session 8 Recording
    Checking access...
  • Main Topics: Diffusion Models, Latent Diffusion Models
    • Brief overview of the theory behind Denoising Diffusion Probabilistic Models (stochastic differential equations, score-matching)
    • Overview of the modalities of diffusion models (Text Conditioning, Classifier-Free Guidance, Inference-Time Optimization).
  • Intro to Diffusion Models Part 1 (11')
    Checking access...
  • Intro to Diffusion Models Part 2 (14')
    Checking access...
  • Conditioning & Classifier-Free Guidance (10')
    Checking access...
  • The UNet Architecture (6')
    Checking access...
  • Inference-Time Optimization: DITTO (6')
    Checking access...
  • Hands On
    • Using stable-audio-tools to generate spectrograms.
    • Experimenting with different initial noises, infilling strategies, conditioning methods.
    • Using Latent Diffusion to generate new latents for our RAVE models from Session 3.
  • Hands On: Using Stable Audio Part 1 (15')
    Checking access...
  • Hands On: Using Stable Audio Part 2 (18')
    Checking access...

Session 9: Real-Time Generative AI & Commercial Applications of Generative AI in Music [Guest: Christian Steinmetz]

4 videos, 2 lessons

+
  • Session 9 Recording
    Checking access...
  • Introduction to TorchScript (18')
    Checking access...
  • 8-bit Linear Quantization (9')
    Checking access...
  • ONNX and Graph Optimizations (15')
    Checking access...
  • Main Topics: Landscape of companies in AI and Music, Available Commercial Products
    • Overview of the current landscape of commercial solutions making use of Generative AI for music and audio.
    • Main differences between academia and industry in the world of AI and music.
    • Tips on developing a career in Generative AI and music (AI engineering, data science, product design)
  • Demo
    • Try out Suno’s commercial product for music generation.

Session 10: Final Project Planning

6 videos, 2 lessons

+
  • Session 10 Recording 1
    Checking access...
  • Session 10 Recording 2
    Checking access...
  • Session 10 Recording 3
    Checking access...
  • Main Topics: Setting up a project specification, timeline, and scope
    • Setting up a project specification, timeline, and scope—skills you can directly apply in music tech roles.
    • Group or individual project proposals aligned with personal career objectives.
  • Training Part 1 (8')
    Checking access...
  • Training Part 2 (13')
    Checking access...
  • Training Part 3 (16')
    Checking access...
  • Peer Review & Feedback
    • Students pitch final project plans, receive structured critique for improvement.
    • Emphasis on real-world communication styles used in professional music and tech teams.

Session 11: Final Project Lab

1 video, 2 lessons

+
  • Session 11
    Checking access...
  • Lab Session: Guided Coding & Troubleshooting
    • Students focus on building out their final AI-driven music projects: from automated composition, stem separation and AI mixing to real-time applications.
    • Instructor mentorship on optimization and code review—mimicking real-world R&D or dev team environments.
  • Milestone Check-Ins
    • Verification that each project meets requirements: code structure, documentation, and creative/music quality.

Session 12: Project Showcase & Next Steps

1 video, 2 lessons

+
  • Session 12 Recording
    Checking access...
  • Final Presentations
    • Students present completed projects, highlighting technical and creative achievements.
    • Demonstration of functionalities and professional-grade code base.


  • Next Steps
    • Discussion of relevant job roles, interview readiness, and resources for continued learning.

Instructors

Lancelot Blanchard

Lancelot Blanchard

Instructor

Lancelot Blanchard is a musician, engineer, and AI researcher at the MIT Media Lab's Responsive Environments group. His research focuses on generative AI systems for human-AI co-created live musical performances, pioneering the concept of Symbiotic Virtuosity in AI-human music collaboration. He has collaborated with Grammy-winning keyboardist Jordan Rudess and brings eight years of software engineering experience to his teaching.

Ready to join?

One-off enrolment. Includes all course materials.