Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset

Adam Roberts; Andriy Stasyuk; Cheng-Zhi Anna Huang; Curtis Hawthorne; Douglas Eck; Erich Elsen; Ian Simon; Jesse Engel; Sander Dieleman

arxiv: 1810.12247 · v5 · pith:23OHCS44new · submitted 2018-10-29 · 💻 cs.SD · cs.LG· eess.AS· stat.ML

Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset

Curtis Hawthorne , Andriy Stasyuk , Adam Roberts , Ian Simon , Cheng-Zhi Anna Huang , Sander Dieleman , Erich Elsen , Jesse Engel

show 1 more author

Douglas Eck

This is my paper

classification 💻 cs.SD cs.LGeess.ASstat.ML

keywords audiodatasetmusicmusicalmaestromodelingmodelsnetworks

0 comments

read the original abstract

Generating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales. Fortunately, most music is also highly structured and can be represented as discrete note events played on musical instruments. Herein, we show that by using notes as an intermediate representation, we can train a suite of models capable of transcribing, composing, and synthesizing audio waveforms with coherent musical structure on timescales spanning six orders of magnitude (~0.1 ms to ~100 s), a process we call Wave2Midi2Wave. This large advance in the state of the art is enabled by our release of the new MAESTRO (MIDI and Audio Edited for Synchronous TRacks and Organization) dataset, composed of over 172 hours of virtuosic piano performances captured with fine alignment (~3 ms) between note labels and audio waveforms. The networks and the dataset together present a promising approach toward creating new expressive and interpretable neural models of music.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing
cs.CV 2026-07 unverdicted novelty 7.0

DisciplineGen-1M is a million-scale multidisciplinary dataset for text-to-image generation and editing, paired with a discipline-informed model that improves results on discipline-specific benchmarks.
SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning
cs.LG 2026-06 unverdicted novelty 7.0

HRM adapters via Hankel reduced-order modeling outperform LoRA on long-context tasks in Mistral-7B when used as SSM residual modules with FFT-based parallel scan.
ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence
cs.SD 2026-04 unverdicted novelty 7.0

ONOTE is a multi-format benchmark that applies a deterministic pipeline to expose a disconnect between perceptual accuracy and music-theoretic comprehension in leading omnimodal AI models.
Latent Fourier Transform
cs.SD 2026-04 unverdicted novelty 7.0

LatentFT uses latent-space Fourier transforms and frequency masking in diffusion autoencoders to enable timescale-specific manipulation of musical structure in generative models.
Self-Supervised Test-Time Tuning for Packet Loss Concealment
eess.AS 2026-07 unverdicted novelty 6.0

TTT-PLC adapts existing PLC models at test time via self-supervised synthetic masking of received audio packets, improving concealment on the same lossy signal in both file and streaming settings.
PJ-RoPE: A Fourier-Jet-Affine Position Space for Relative Attention
cs.LG 2026-06 unverdicted novelty 6.0

PJ-RoPE organizes relative-position mechanisms as a learnable Fourier-Jet-Affine space derived from lag-shift dynamics, extending RoPE and ALiBi with explicit jets and sector selection.
Rubato: Transcribing Piano Music with Timestamps
cs.SD 2026-05 unverdicted novelty 6.0

Rubato model with InterMo representation outperforms cascade methods in generating timestamped piano sheet music from audio, even when cascades receive ground-truth MIDI.
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
cs.SD 2026-05 unverdicted novelty 6.0

Introduces the first large-scale Persian music dataset and shows fine-tuned MusicGen produces compositions more aligned with Persian stylistic conventions via tag-based evaluation.
Music Transcription with (Almost) No Supervision
cs.SD 2026-05 unverdicted novelty 5.0

Cycle-consistent translation enables competitive music transcription performance with mostly unpaired audio and scores plus minimal paired supervision.
A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
eess.AS 2026-05 unverdicted novelty 2.0

A structured survey of audio bandwidth extension that organizes the transition from deterministic discriminative DNNs to generative approaches including GANs, diffusion models, and flow-based methods.