Demucs vs Moises Stem Quality: A Reproducible Listening Test
Compare Demucs and Moises stem separation without guessing which model a service uses: level-match the outputs, inspect bleed and transients, and record a repeatable result.
Two separated stems can sound similar without coming from the same model. Model architecture, training data, inference settings, overlap, shifts, post-processing, and export encoding can all change the result. Moises does not publicly document that its separation service runs Demucs, so a useful Demucs vs Moises comparison must not start with that assumption.
Direct answer: use one checksum-recorded source, request the closest available stem layouts, preserve original outputs, align and level-match them, and write timestamped notes for target retention, bleed, transients, sustained tone, and musical continuity. Choose the result that serves the passage and workflow. Do not declare a universal winner or infer hidden model identity from sound.
The defensible question is narrower: which output better preserves the instrument you need for this song and this practice task? You can answer that with a controlled listening test.
What the vendors actually disclose
Demucs v4 is open source. Its repository identifies the default htdemucs family as a hybrid waveform/spectrogram architecture and documents model selection, shifts, overlap, segment length, output stems, and WAV or MP3 export options. The original Facebook Research repository also states that it is no longer actively maintained and is archived/read-only. Record the exact source, revision, environment, command, and model rather than treating “Demucs” as one current service.
Moises documents a hosted product with multiple separation choices and says its AI models are trained on licensed material. Its public documentation does not identify the separation service as Demucs. Treat its internal architecture and training recipe as undisclosed unless Moises publishes a specific statement. Similar sound is not evidence of shared weights.
That distinction matters. A comparison that says “same model, same quality” cannot explain a result and cannot be reproduced. A comparison that preserves the input, settings, outputs, and observations can.
Build an apples-to-apples test
Use two or three short excerpts that expose different failure modes rather than one easy song:
- A dense chorus where vocals, guitars, and cymbals overlap.
- A bass passage that shares attacks with the kick drum.
- A quiet intro with reverb tails or room sound.
Start from the same lossless source file. Request the closest matching stem layout from each tool. Keep every original export; do not normalize, denoise, or convert one result before the other.
For listening, align the files to the same sample start and level-match them. Louder usually sounds “better” in a quick comparison even when it contains more bleed. If an export has a different sample rate or codec, record that fact rather than hiding it in a conversion.
Listen for evidence, not a single quality score
Solo each target stem and then audition it against the residual mix. Check four boundaries:
- Target retention: Are quiet notes, consonants, ghost notes, and reverb tails still present?
- Bleed: How much unrelated vocal, cymbal, guitar, or kick energy remains in the stem?
- Transient damage: Do drum attacks smear, double, or disappear?
- Musical continuity: Does the part pump or change tone when another instrument enters?
Write observations against timestamps. “Bass note at 00:42 loses its attack” is useful. “Version A sounds more AI” is not. Repeat the comparison at normal playback speed and at the slower speed you actually use for practice; artifacts that are harmless in a mix can become distracting inside a loop.
Why SDR numbers do not settle this comparison
SDR is meaningful only when systems are evaluated against the same reference stems, dataset split, stem definition, alignment, and metric implementation. A paper’s Demucs score and a service’s marketing number are not automatically comparable. Without matched reference audio, report a structured listening result instead of inventing a dB difference.
A compact test record can include:
| Field | Record |
|---|---|
| Source | filename, duration, sample rate, checksum |
| Requested stems | exact separation mode in each tool |
| Local model | model name, revision, command, shifts/overlap |
| Service run | date, selected quality/mode, exported format |
| Comparison | level adjustment and alignment method |
| Findings | timestamped retention, bleed, transient, and continuity notes |
This record lets you rerun the test after either implementation changes.
Choose the workflow after the sound check
Demucs is a strong fit when you want a documented local model, batchable commands, and direct control over inference and output. A hosted service is convenient when you want its account, mobile, library, and cloud workflow. These are operational differences; neither proves a universal quality winner.
Session Craft takes a third path for musicians who want local separation inside the practice session. It bundles a revision-pinned HTDemucs-derived ONNX asset with its source provenance and checksum, then keeps stems beside speed, pitch, A/B loops, chord context, and the saved project. The Community edition is free, processing stays on the desktop, and the useful result is not merely a folder of stems: it is a repeatable practice workspace.
For the broader local-versus-cloud decision, continue with the desktop Moises alternative workflow or the stem-separation practice guide.
Match stem layouts before judging
The common Demucs default produces Vocals, Drums, Bass, and Other. Its documented experimental six-source model adds Guitar and Piano-like output. Session Craft’s current local project exposes Vocals, Drums, Bass, Guitar, Keyboard, and Other. Moises separation choices vary by current tier and requested mode.
If one result has four roles and another has six:
- compare equivalent combined material;
- document how Guitar/Keyboard are recombined;
- avoid calling one “cleaner” merely because content moved into another category;
- keep Other in every reconstruction check.
More output stems do not guarantee better source preservation.
Prepare the source
Use audio you are authorized to process. Preserve the original and record:
- filename and checksum;
- container and codec;
- sample rate, channel count, and duration;
- exact time ranges;
- any prior mastering or conversion;
- intended practice task.
Do not compare one lossless source with a separately downloaded lossy version.
Choose difficult excerpts
| Excerpt | Failure mode exposed |
|---|---|
| Sparse vocal | Consonant and ambience retention |
| Dense chorus | Cross-role bleed and pumping |
| Drum fill | Transient smear and timing |
| Kick plus bass | Low-frequency assignment |
| Guitar plus keyboard | Six-role ambiguity |
| Long reverb tail | Shared effect handling |
The benchmark should represent the user’s repertoire, not only a separator-friendly demo.
Align before listening
Exports can begin at different sample positions or contain codec delay. Align an identifiable transient. If the files drift, investigate sample-rate or duration differences before judging.
Level-match conservatively. Do not process away differences with denoising, EQ, limiting, or normalization applied to one output only. Keep raw exports.
Score by task, not one number
Use a small rubric:
| Criterion | 1 | 3 | 5 |
|---|---|---|---|
| Target retention | Important material missing | Mostly usable | Critical material preserved |
| Bleed | Prevents task | Noticeable but manageable | Minimal for task |
| Transients | Timing misleading | Some softening | Attacks remain useful |
| Sustained tone | Strong warble/pumping | Audible but usable | Stable for task |
| Context reconstruction | Form damaged | Mostly coherent | Musical context preserved |
The scores are listening notes, not scientific SDR measurements. Include timestamps and a written reason.
Test at practice settings
A stem may sound acceptable at normal speed and poor at the slowdown used for transcription. Test:
- normal speed;
- required slower speed;
- solo target;
- target plus anchor;
- full reconstructed context.
The winning raw stem may not produce the winning practice loop.
Reproducibility record
For local Demucs:
repository/commit
environment
model
command
shifts/overlap/segment
output format
For Moises:
date
platform
account tier
selected separation
quality option
export format
For Session Craft:
application version
platform
bundled model provenance/checksum
six-stem completion
project and source path
Redact private account data and never publish tokens or signed URLs.
Operational comparison
After audio quality, compare:
- local or upload processing;
- compute and wait dependency;
- account and tier;
- cross-device access;
- project recall;
- A/B, speed, and pitch workflow;
- export entitlement;
- source and stem recovery.
An output can sound best but violate source custody. Another can be slightly less isolated but complete the practice task locally and repeatably.
Comparison QA checklist
- Same authorized source and excerpts.
- Exact stem modes recorded.
- Original exports preserved.
- Alignment verified.
- Levels matched.
- Four-versus-six mapping documented.
- Timestamped retention and bleed notes.
- Transients and sustain checked.
- Normal and practice speeds tested.
- Workflow and privacy evaluated separately.
- No hidden-model claim.
- No incompatible SDR comparison.
Frequently asked questions
Does Moises use Demucs?
Do not claim that unless Moises publishes specific current evidence. Similar output is not proof of model identity.
Is Demucs actively maintained?
The original Facebook Research repository says it is no longer actively maintained and is archived. Exact forks, packages, or wrappers need their own maintenance check.
Is six-stem output better than four-stem output?
Not automatically. Additional categories can help isolate Guitar or Keyboard but can also increase misassignment. Judge the target task.
Can I compare with SDR?
Only with matched ground-truth stems, definitions, alignment, dataset, and metric implementation. For a user’s commercial song without ground truth, use structured listening evidence.
What does Session Craft add?
It bundles a pinned derived local model asset and keeps six stems with speed, pitch, A/B loops, and project state. Community includes that core workflow; Professional adds chord detection and export.
Which result should I choose?
Choose the output and workflow that preserve the required role and context for the exact song while satisfying custody and recovery requirements.
Final recommendation
A defensible Demucs-versus-Moises comparison preserves inputs, settings, outputs, and timestamped observations. It separates audio quality from architecture and refuses to turn one passage into a universal ranking.
Publish the limitations with the result
A useful comparison states the song type, source quality, tested passages, output layouts, alignment method, level-matching method, listening system, test date, and product or model versions. It also states what was not tested. Without that context, “better quality” cannot be reproduced.
Keep observations role-specific: vocal retention, bass attack, drum transient, guitar or keyboard misassignment, shared reverb, and full-mix continuity. A result can win one role and lose another. If layouts differ, compare the closest musical task rather than pretending every stem maps one-to-one.
Session Craft should be evaluated as its current local six-stem practice workflow, not presented as the upstream Demucs command-line project or as proof of Moises internals. Product behavior, inference asset provenance, practice controls, and project recovery are separate pieces of evidence.
<!-- multilingual-related-reading:start -->Related guides
Continue with the same-language pages below. They cover adjacent stages without changing the canonical owner of this topic:
<!-- multilingual-related-reading:end -->