Daejeon, July 27–31
AI music generation has developed rapidly, and aligning its output with human preferences has received increasing attention. Yet across these stages, questions of whose preferences are collected, which dimensions are emphasized, and how preferences are encoded in training remain largely underexplored. We develop a four-stage framework and apply it to recent post-training systems, evaluation benchmarks, and metric audits, tracing how human preference shapes the system at each stage. Across this pipeline, we identify a gap between rhetoric and practice: systems described as “human-aligned” increasingly rely on automated proxies and post-hoc validation rather than direct human preference data.
This analysis examines four stages of human involvement in AI music generation post-training.
Stage 1: Inherited pretraining priors
Post-training systems begin from pretrained foundation models such as LeVo (Lei et al. 2025), MelodyFlow (Le Lan et al. 2024), or MusicLM (Agostinelli et al. 2023), whose training data already reflect choices about which genres, languages, and traditions are represented. These priors are inherited, and they propagate into any preference alignment performed at later stages.
Stage 2: Aesthetic dimensions and rating rubrics
Rubric design decides which aspects of musical experience are measurable. It translates complex listening judgments into discrete categories and shapes which musical dimensions are encoded into the system.
Stage 3: Preference data and post-training optimization
Human preference data is collected through listening studies, then used to fine-tune the foundation model. The collected preferences serve as the optimization objective for the fine-tuning process.
Stage 4: Model evaluation
After post-training, the generated music is evaluated on selected dimensions, revealing which aspects developers consider meaningful. Whether listeners function as humans in the loop or as post-hoc evaluators varies across systems.
We compare recent works across three layers of AI music human-preference alignment: post-training systems (MusicRL (Cideron et al. 2024), Hallucination-Free Music (Zhang et al. 2025), MR-FlowDPO (Ziv et al. 2026)), benchmarks that define and measure it (Audiobox-Aesthetics (Tjandra et al. 2025), MusicEval (Liu et al. 2025), SongEval (Yao et al. 2025)), and an audit (Grötschla et al. 2025) that tests whether deployed metrics actually represent human preference.
For each work, we examine data collection procedures, proxy reward models, and evaluation rubrics to identify where human preference shapes the model and where automated proxies substitute for it. This approach treats AI music pipelines as cultural artifacts whose design decisions influence how musical aesthetics become encoded.
Across these systems, four patterns emerge.
Stages 2 and 3 are increasingly compressed into a single algorithmic step: the proxy becomes the rubric.
Rubric design (deciding which dimensions of music to measure) and preference data collection (providing the optimization objective) are no longer clearly separate. Instead, developers select off-the-shelf proxy models that do both at once: define what counts as preferable and serve as the optimization objective. For example, Hallucination-Free Music uses Phoneme Error Rate as the proxy for hallucination, while MR-FlowDPO uses Contrastive Language-Audio Pretraining (CLAP), Audiobox-Aesthetics, and HuBERT-likelihood to represent alignment, quality, and musicality. MusicRL is the main counterexample, collecting human preferences directly through large-scale listener studies. Later systems treat this approach as a baseline to surpass rather than a standard.
Even as proxies substitute for human preference data, this substitution is not consistently validated.
MR-FlowDPO collects 32K human preference pairs and reports that proxy-based pairing achieves comparable post-training performance to human annotation at the same scale. Hallucination-Free Music does not make this comparison. Grötschla et al. show that several widely deployed metrics correlate poorly with human preference, raising questions about the reliability of proxy-based optimization objectives.
Evaluation dimensions agree in name more than in measurement.
Terms such as text-audio alignment, audio quality, and musicality recur across these systems, but their measurements vary. Grötschla et al. show that “audio quality” may refer to five different Fréchet Audio Distance (FAD) variants, while “text-audio alignment” may be measured by seven different CLAP-family checkpoints, each yielding different metric behavior on the same outputs. Subjective benchmarks are similarly fragmented: Audiobox-Aesthetics has four axes, SongEval five, MusicEval two.
Stage 4 is where humans are most clearly involved, but in proxy-based systems they function as post-hoc evaluators rather than as humans in the loop.
Listening tests are conducted after training is complete and cannot influence the optimization objective; listeners only evaluate whether the output sounds acceptable. Combined with the Stage 2 and 3 collapse, this creates a gap between rhetoric and practice: systems can be described as “human-aligned” even when human feedback shapes neither the rubric nor the optimization objective, but only the post-hoc verdict.