Results

Every method per stage, every full-pipeline combination, and the patterns behind the numbers.

No pipeline combinations for this field mapping yet.

#
Loading results…

No results for this view yet.

Patterns the leaderboard tables don't surface, computed live from the same results. The recurring lesson: a method's standing depends heavily on what you score, and on the pipeline it sits in, not just on a single number.

Loading results…

The isolated ranking misleads about real pipelines

Dipole inversion ranked on the clean ground-truth field (left) versus chained after real background removal (right), by NRMSE, best at top. Green rises when composed (robust to upstream error); red falls. TV-family methods top the clean ranking, then collapse; several networks climb.

Accurate but fragile, or reliable but modest

Each dipole method by its mean NRMSE across background-removal partners (x) versus how much that error varies between partners (y). Bottom-left is the sweet spot: accurate and consistent. Iterative methods lead on average but scatter wildly with a poor upstream stage; networks are dull but dependable.

Error compounds through the pipeline

The best achievable xSIM (blue) and the median (green box = interquartile range) at three pipeline depths: dipole inversion alone from the ground-truth local field, then with real background removal added (from the ground-truth total field), then a full pipeline from raw phase. Each real upstream stage roughly halves the median and caps the best — the ceiling falls from 0.78 → 0.68 → 0.39. Upstream error is never recovered downstream, so a method's isolated score badly overstates its end-to-end result.

Pick your dipole independently of background removal

For each dipole method, its rank among dipoles (dot = mean, bar = min–max) across the eight strong background-removal methods. Most bars are tight: the dipole ranking barely shifts with the upstream stage (mean pairwise Spearman ρ ≈ 0.87), so the stages are near-separable — wh-qsm and hd-qsm lead everywhere, modl-qsm/qsmgan/nltv trail everywhere. A handful in the middle (xqsm, l1-qsm) swing widely — those are the genuinely upstream-sensitive methods. Coloured by family.

The best background-removal method depends on what you measure

Each background-removal method's rank (1 = best, green good → red poor) under five metrics, rows sorted by xSIM. The similarity metrics — xSIM, NRMSE, HFEN — are nearly interchangeable (their columns match). But calcification and iron linearity reshuffle the ranking: iron linearity is essentially uncorrelated with similarity (ρ ≈ −0.02), and the deep-learning bfrnet is worst-tier on similarity yet best on iron quantification. A single headline number hides which sources a method actually recovers.

The best method depends on what you measure

Isolated dipole methods across five metrics, each normalised so higher is better. Grey lines are all methods; the coloured lines are the winner of a given axis, showing they lose elsewhere: the global-NRMSE / xSIM leader is among the worst at calcification, and the fast direct methods that under-quantify iron.

Where the error lives

Isolated dipole error by region and artefact (rows sorted by overall NRMSE). Green is good, red is poor, coloured within each column. Venous blood is universally hardest; the calcification columns reshuffle the ranking entirely.

Which metrics are redundant

Rank agreement (Spearman ρ) between metrics over the isolated dipole methods, sign-aligned so +1 means they reward the same methods. NRMSE, detrended NRMSE and correlation are near-interchangeable; DGM linearity, calcification and runtime are their own axes, so a single headline number can't stand in for them.

Accuracy vs compute

Isolated dipole inversion: structural similarity (xSIM, higher is better) against wall-clock runtime on a log scale. Green points are Pareto-optimal (nothing is both faster and more accurate); grey points are dominated. Some methods cost minutes for no accuracy gain.

Build your own comparison

Plot any two metrics against each other for any isolated stage. Points are methods, coloured by family; hover to identify, click to open a method. Runtime uses a log scale.

In silico · QSM Challenge 2.0 (2019)

The main dataset is the QSM Reconstruction Challenge 2.0 in-silico head phantom, forward-simulated to a BIDS multi-echo acquisition with qsm-forward. Its per-voxel susceptibility is known exactly, so QSM-CI scores each pipeline stage on its own (field mapping, background removal, dipole inversion) rather than only the final map. It backs the Full pipelines, Stages and Findings views.

Why a synthetic ground truth

QSM is an ill-posed dipole deconvolution, so different algorithms produce visibly different maps from the same phase. Judging them needs a true susceptibility map, which no in vivo acquisition provides, so every in vivo reference is a surrogate that biases the comparison: a COSMOS reconstruction (Liu et al., 2009) averages anisotropic susceptibility across orientations, and simple piecewise-constant phantoms (de Rochefort, 2008; Wharton, 2010) flatter the TV/TGV methods that reconstruct piecewise-constant maps by construction (Knoll, 2011). A ground truth is never neutral; its texture decides which assumptions get rewarded.

Challenge 2.0 (Marques et al., 2021) answered this with an in-silico head phantom: a realistic digital brain from a 0.64 mm 7 T acquisition, with an interhemispheric calcification, whose susceptibility is known exactly and from which a data-consistent gradient-echo acquisition is forward-simulated. It deliberately carries physiological texture so it does not reward piecewise-smooth solutions the way earlier phantoms did.

What QSM-CI scores, and what it leaves out

Because the total field, local field and tissue susceptibility are all known, QSM-CI scores each stage independently and reruns methods as they appear. The simulation carries realistic background fields from air/tissue interfaces and wrapped multi-echo phase, so field mapping and background removal genuinely have work to do, not just dipole inversion.

It stays idealised in known ways: susceptibility is an isotropic scalar (no white-matter anisotropy), there is no flow or motion, microstructure enters only through the χ-separation compartments, noise is complex Gaussian, and the phantom comes from a single subject. A QSM-CI ranking is a strong controlled test, not a guarantee of in vivo performance.

Reproduce the phantom

qsm-forward simulates the acquisition from the Challenge 2.0 tissue-property maps. Request those from the Donders repository (data.ru.nl/…/DSC_3015069.02_542, DOI 10.34973/m20r-jt17), then generate:

pip install qsm-forward
# 7 T, 4 echoes, 1 mm, fixed seed: matches the scored phantom
qsm-forward head ./qsm-challenge-2.0 bids_qsm --save-field

--save-field emits the total field (fieldmap) and local field (fieldmap-local) that the isolated background-removal and dipole stages are scored against. Acquisition parameters (--B0, --TEs, --voxel-size, --peak-snr, --random-seed) default to the challenge phantom.

Citations

  • Head phantom: Marques J.P., Meineke J., Milovic C., et al. “QSM reconstruction challenge 2.0: A realistic in silico head phantom for MRI data simulation and evaluation of susceptibility mapping procedures.” Magn Reson Med 2021;86(1):526–542. doi:10.1002/mrm.28716. Data: doi:10.34973/m20r-jt17.
  • Challenge 2.0 design & results: QSM Challenge 2.0 Organization Committee. “QSM reconstruction challenge 2.0: Design and report of results.” Magn Reson Med 2021;86(3):1241–1255. doi:10.1002/mrm.28754.
  • COSMOS reference: Liu T., Spincemaille P., de Rochefort L., et al. “Calculation of susceptibility through multiple orientation sampling (COSMOS).” Magn Reson Med 2009;61(1):196–204. doi:10.1002/mrm.21828.
  • Piecewise-constant inversion: de Rochefort L., et al. Magn Reson Med 2008;60(4):1003–1009. doi:10.1002/mrm.21710. Wharton S., et al. Magn Reson Med 2010;63(5):1292–1304. doi:10.1002/mrm.22334.
  • TGV regularisation: Knoll F., Bredies K., Pock T., Stollberger R. “Second order total generalized variation (TGV) for MRI.” Magn Reson Med 2011;65(2):480–491. doi:10.1002/mrm.22595.
  • Forward simulation: qsm-forward (Stewart A., et al.).
# Algorithm

No chi-separation results yet.

Myelin (χ−) is harder to recover than iron (χ+)

Structural similarity (xSIM) of the recovered paramagnetic χ+ (iron) source versus the diamagnetic χ− (myelin / calcium) source. Every method drops from left to right, so the diamagnetic map is systematically the harder of the two.

Cross-contamination between the two sources

How much each source map bleeds into the other (regression slope, 0 = clean): iron leaking into the myelin map (x) versus myelin leaking into iron (y). The origin is perfect separation; distance from it is contamination that xSIM can miss.

χ-separation (2026)

Susceptibility source separation splits the net susceptibility χ into a paramagnetic part χ⁺ (iron) and a diamagnetic part χ⁻ (myelin, calcium). On this dataset each method is fed the ground-truth local field, R2′ and χtotal, and its recovered χ⁺ and χ⁻ maps are scored separately. It extends the in-silico Challenge 2.0 phantom, so the susceptibility ground truth is the same one behind the In silico dataset.

The phantom

The χ-separation phantom is generated with qsm-forward, which ports and extends the neuropoly Susceptibility-Separation-Phantom (Ridani et al., 2026), itself an extension of the Challenge 2.0 head phantom. It adds three things the plain QSM phantom lacks: the χ⁺/χ⁻ source maps (the scoring targets), an R2′ map, and a source-separation-aware signal model. The χ⁺/χ⁻ split and relaxation model follow the reference phantom; qsm-forward additionally offers a matched spin-echo acquisition and a multi-compartment signal model.

The forward model

Relaxation: R2, R2*, and R2′

The gradient-echo magnitude decays at R2* = R2 + R2′. R2 is the irreversible rate; R2′ is the reversible rate from static field inhomogeneity around susceptibility sources, and it is the channel source separation draws on. The phantom simulates R2 independently from per-tissue literature T2 values and adds it to R2′ to reach R2*, with no fixed ratio between them (the common κ shortcut, R2* ≈ 1.9·R2′ from Dimov et al., 2022, is avoided because it makes recovering R2′ = R2* − R2 degenerate). The R2′ map is shipped directly as a provided input; qsm-forward can instead simulate a matched spin echo (--save-se) that lets a method derive R2′ from paired GRE/SE data (Stoll, 2025), a capability this dataset does not currently ship.

The relaxivity Dr

Relaxivity converts a susceptibility source into the reversible relaxation it produces. χ-separation uses a single kernel shared by both source types:

R2′ = D_r · ( |χ⁺| + |χ⁻| ),   D_r = 137 Hz/ppm

One kernel for both, because in the static-dephasing regime (Yablonskiy & Haacke, 1994) reversible relaxation depends on the magnitude of the field perturbation, not its sign. Dr = 137 Hz/ppm is the value Shin et al. (2021) measured. A split (Dr⁺ ≠ Dr⁻) is left off by default, since the two relaxivities are not recoverable from one QSM and one R2′ map; it stays available as an opt-in (--dr-neg) for sensitivity studies.

The multi-compartment magnitude

Field-domain methods need only R2′ and the field. Signal-domain separators (e.g. DECOMPOSE) fit the multi-echo complex signal per voxel, where a plain mono-exponential decay carries nothing to separate. With --chisep-multicompartment the voxel magnitude becomes the modulus of a sum of compartments:

S(TE) = | C₊·e^(−(R2 + D_r|χ⁺| + i·ω·χ⁺)·TE)
        + C₋·e^(−(R2 + D_r|χ⁻| + i·ω·χ⁻)·TE)
        +  C₀·e^(−R2·TE) |,     ω = (2/3)·γ·B₀

Each compartment is a static-dephasing exponential with its own decay and off-resonance, so the pools beat against each other and give the non-mono-exponential magnitude those methods rely on. The rates reuse the same Dr kernel as R2′, so the effective R2* is unchanged and field-domain methods see identical inputs; only the shape of the decay is enriched. It defaults to off.

Reproduce the phantom

From the same Challenge 2.0 maps, add the source-separation outputs:

qsm-forward head ./qsm-challenge-2.0 bids_chisep \
  --save-field --save-chi-pos --save-chi-neg --save-r2prime \
  --chisep-signal --chisep-multicompartment

--save-chi-pos/--save-chi-neg write the scoring targets, --save-r2prime the provided R2′ input, and the --chisep-* flags select the source-separation signal model. qsm-forward writes a standard BIDS tree; point qsm-ci run straight at the NIfTIs.

Citations

  • Separation phantom: Ridani S., De Leener B., Alonso-Ortiz E. “A realistic in-silico brain phantom for quantifying susceptibility anisotropy-induced error in susceptibility separation.” bioRxiv 2026. doi:10.64898/2026.04.07.716972.
  • χ-separation model & Dr: Shin H.G., Lee J., Yun Y.H., et al. “χ-separation: Magnetic susceptibility source separation toward iron and myelin mapping in the brain.” NeuroImage 2021;240:118371. doi:10.1016/j.neuroimage.2021.118371.
  • Static-dephasing regime: Yablonskiy D.A., Haacke E.M. “Theory of NMR signal behavior in magnetically inhomogeneous tissues: the static dephasing regime.” Magn Reson Med 1994;32(6):749–763. doi:10.1002/mrm.1910320610.
  • DECOMPOSE: Chen J., et al. “Decompose QSM to sub-voxel diamagnetic and paramagnetic components based on gradient-echo MRI data.” NeuroImage 2021. doi:10.1016/j.neuroimage.2021.118735.
  • SUSEP-Net: Li, Gao, Sun, et al. “SUSEP-Net: simulation-supervised contrastive learning for susceptibility source separation via a sub-voxel QSM network.” arXiv:2506.13293, 2025. doi:10.48550/arXiv.2506.13293.
  • WaveSep: Fang Z., Shin H.G., van Zijl P., Li X., Sulam J. “WaveSep: A Flexible Wavelet-Based Approach for Source Separation in Susceptibility Imaging.” MLCN, MICCAI 2023. doi:10.1007/978-3-031-44858-4_6.
  • κ approximation: Dimov A.V., et al. “Magnetic susceptibility source separation solely from gradient echo data: histological validation.” Tomography 2022;8(3):1544–1551. doi:10.3390/tomography8030127.
  • χ-sep spin-echo model: Stoll P. “Development of a Deep Learning Framework for Iron and Myelin Mapping from Quantitative Susceptibility Maps.” MSc thesis, ETH Zurich, 2025.
# Algorithm vs COSMOS vs STI χ33 Runtime

No in-vivo results yet.

COSMOS vs STI χ33 — the harder reference

Each method's structural similarity against the two references. Every point sits below the diagonal — STI χ33 is a systematically harder target than COSMOS (roughly +20 NRMSE across the board) — yet the two references rank methods almost identically. Coloured by family; hover for values.

Do NRMSE and HFEN agree?

The headline error metric (NRMSE) against the fine-detail metric (HFEN, the 2016 challenge's signature), both vs COSMOS. A high Spearman ρ means they reward the same methods; where a point strays off the trend, a method wins on bulk error but loses fine structure, or vice-versa.

Runtime vs accuracy

In-vivo structural similarity (xSIM vs COSMOS) against wall-clock runtime on a log scale. Green points are Pareto-optimal (nothing is both faster and more accurate); grey points are dominated.

In vivo · 2016 Reconstruction Challenge

This dataset is real acquired brain data from the 2016 QSM Reconstruction Challenge: a single-orientation 3 T scan of one volunteer. Each method inverts the ground-truth local tissue field to a susceptibility map, scored against two reference reconstructions rather than a true susceptibility (none exists in vivo). Only the dipole-inversion stage is scored, so the ranking reads as agreement with a reference, not accuracy against ground truth.

The challenge and its references

The 2016 Reconstruction Challenge (Langkammer et al., 2018) took the χ33 component of a susceptibility-tensor (STI) reconstruction as its reference, hoping to sidestep the orientation bias of COSMOS. Twenty-seven submissions were scored on RMSE, HFEN, SSIM and regional error. The lesson was as much about metrics as algorithms: tuned to minimise those errors, the winning maps were over-smoothed and lost fine structure, and near-identical scores hid visibly different maps. A follow-up (Milovic et al., 2020) then showed the χ33 reference itself did not match the measured phase once anisotropy, microstructure and noise were accounted for.

QSM-CI scores each method against both available references: the multi-orientation COSMOS reconstruction and the STI χ33 map. Source: the 2016 challenge (qsm.neuroimaging.at); the local tissue field and both reference χ maps are in ppm.

Citations

  • 2016 Challenge report: Langkammer C., Schweser F., Shmueli K., et al. “Quantitative susceptibility mapping: Report from the 2016 reconstruction challenge.” Magn Reson Med 2018;79(3):1661–1673. doi:10.1002/mrm.26830.
  • 2016 Challenge lessons: Milovic C., Tejos C., Acosta-Cabronero J., et al. “The 2016 QSM Challenge: Lessons learned and considerations for a future challenge design.” Magn Reson Med 2020;84(3):1624–1637. doi:10.1002/mrm.28185.
  • COSMOS reference: Liu T., Spincemaille P., de Rochefort L., Kressler B., Wang Y. “Calculation of susceptibility through multiple orientation sampling (COSMOS).” Magn Reson Med 2009;61(1):196–204. doi:10.1002/mrm.21828.

Dipole inversion across datasets

The same isolated dipole-inversion methods, scored on the in-silico phantom (2019, true χ) and the in-vivo challenge (2016, COSMOS reference). Does a ranking on the idealised phantom carry over to real acquired data?

Phantom rank does not predict in-vivo rank

Each method's xSIM on the in-silico phantom (x) against its xSIM in vivo (y), coloured by family. The two are essentially uncorrelated: classical iterative and direct methods that win on the clean phantom collapse in vivo, while deep-learning methods hold up. Hover for values; click to open a method.

How methods move between datasets

Each method's rank on the in-silico phantom (left) linked to its rank in vivo (right), by xSIM. Lines coloured by family: deep-learning methods tend to climb left-to-right; classical methods fall.

Domain robustness by family

Mean xSIM per method family on each dataset. Deep-learning methods barely drop from phantom to in-vivo; classical iterative and direct methods lose far more — tuned to the idealised phantom, they generalise worse to real 3 T data.