Research results
Seven open research models across X-ray, CT, MRI and mammogram, trained on public data, for education and research.
Research results on public test data. They say nothing about your own scan.
How to read the charts
- AUROC: 0.5 is a coin toss and 1.0 is perfect.
- Dice: how much the outline overlaps the dataset's own outlines; 1.0 is a perfect match.
- Kappa: 0 is chance agreement and 1.0 is full agreement.
- Dots: a filled dot is a held-out test, a hollow dot is development. The line through a dot is its 95% interval.
X-ray
Chest X-ray research model
A team of ten networks trained to give 15 scores for one frontal chest X-ray, one for each of the 14 NIH ChestX-ray14 label types and one for its abnormal label.
Research results on public test data. They say nothing about your own scan.
Results
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Agreement with ChestX-ray14's labels, mean over its 14 label types | AUROC | 0.818 | 0.811 to 0.823 | NIH's official test list: 25,596 pictures of 2,797 patients | held-out test |
| Agreement with ChestX-ray14's abnormal label | AUROC | 0.736 | 0.724 to 0.748 | NIH's official test list: 25,596 pictures | held-out test |
Exact values: Agreement with ChestX-ray14's labels, mean over its 14 label types: AUROC 0.8175 (95% interval 0.811 to 0.8231); Agreement with ChestX-ray14's abnormal label: AUROC 0.7363 (95% interval 0.7241 to 0.7477).
Agreement with ChestX-ray14's labels, per label type
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Atelectasis | AUROC | 0.775 | 0.762 to 0.788 | NIH's official test list, 25,596 pictures; 3,279 pictures | held-out test |
| Cardiomegaly | AUROC | 0.902 | 0.886 to 0.914 | NIH's official test list, 25,596 pictures; 1,069 pictures | held-out test |
| Effusion | AUROC | 0.839 | 0.828 to 0.849 | NIH's official test list, 25,596 pictures; 4,658 pictures | held-out test |
| Infiltration | AUROC | 0.707 | 0.696 to 0.718 | NIH's official test list, 25,596 pictures; 6,112 pictures | held-out test |
| Mass | AUROC | 0.824 | 0.803 to 0.842 | NIH's official test list, 25,596 pictures; 1,748 pictures | held-out test |
| Nodule | AUROC | 0.754 | 0.733 to 0.774 | NIH's official test list, 25,596 pictures; 1,623 pictures | held-out test |
| Pneumonia | AUROC | 0.727 | 0.702 to 0.750 | NIH's official test list, 25,596 pictures; 555 pictures | held-out test |
| Pneumothorax | AUROC | 0.861 | 0.845 to 0.877 | NIH's official test list, 25,596 pictures; 2,665 pictures | held-out test |
| Consolidation | AUROC | 0.761 | 0.746 to 0.775 | NIH's official test list, 25,596 pictures; 1,815 pictures | held-out test |
| Edema | AUROC | 0.854 | 0.838 to 0.870 | NIH's official test list, 25,596 pictures; 925 pictures | held-out test |
| Emphysema | AUROC | 0.896 | 0.880 to 0.910 | NIH's official test list, 25,596 pictures; 1,093 pictures | held-out test |
| Fibrosis | AUROC | 0.838 | 0.815 to 0.861 | NIH's official test list, 25,596 pictures; 435 pictures | held-out test |
| Pleural thickening | AUROC | 0.784 | 0.766 to 0.804 | NIH's official test list, 25,596 pictures; 1,143 pictures | held-out test |
| Hernia | AUROC | 0.923 | 0.878 to 0.957 | NIH's official test list, 25,596 pictures; 86 pictures | held-out test |
Test set: NIH's official test list, 25,596 pictures.
Where it falls short
- Agreement is lower for some label types: infiltration scores 0.7071, pneumonia 0.7267 and nodule 0.7544.
- Rare label types have wide intervals. Hernia has only 86 test pictures, and its interval runs from 0.8776 to 0.9574.
- All the labelled pictures come from one hospital, so nothing shows how it behaves on other scanners, other patient groups or children.
Trained and tested on
- NIH ChestX-ray14. Licence: NIH asks for credit; use is unrestricted
- TAIX-Ray (pre-training pictures only). Licence: CC BY 4.0
- COVID-19-NY-SBU (pre-training pictures only). Licence: CC BY 4.0
- Tuberculosis chest X-ray sets, Kiran and Jabeen (pre-training pictures only). Licence: CC BY 4.0
- BDCXR-3257 (pre-training pictures only). Licence: CC BY 4.0
Weights
Not released yet.
CT
CT organ outliner
A two-step network trained to label each point of a 3D CT scan as one of 117 organs and bones, such as the liver, spleen, aorta, ribs and vertebrae, or as nothing.
Research results on public test data. They say nothing about your own scan.
Results
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Mean over 115 structures | Dice | 0.867 | 0.854 to 0.878 | Held-out test: 89 CT scans | held-out test |
| Mean over 115 structures, development scans | Dice | 0.816 | 0.812 to 0.819 | 1,139 development scans, each read by models that never saw it | development |
Exact values: Mean over 115 structures: Dice 0.8668 (95% interval 0.8538 to 0.8781); Mean over 115 structures, development scans: Dice 0.8156 (95% interval 0.8116 to 0.8194).
Mean Dice by group of scans
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| All | Dice | 0.867 | 0.854 to 0.878 | Held-out test: 89 CT scans (a scan can sit in more than one group); 89 scans | held-out test |
| Abdomen scans | Dice | 0.880 | 0.868 to 0.891 | Held-out test: 89 CT scans (a scan can sit in more than one group); 47 scans | held-out test |
| Thorax scans | Dice | 0.888 | 0.876 to 0.899 | Held-out test: 89 CT scans (a scan can sit in more than one group); 32 scans | held-out test |
| Head or neck scans | Dice | 0.877 | 0.855 to 0.895 | Held-out test: 89 CT scans (a scan can sit in more than one group); 19 scans | held-out test |
| Siemens scanners | Dice | 0.871 | 0.853 to 0.886 | Held-out test: 89 CT scans (a scan can sit in more than one group); 58 scans | held-out test |
| Other scanner makers | Dice | 0.864 | 0.845 to 0.881 | Held-out test: 89 CT scans (a scan can sit in more than one group); 31 scans | held-out test |
| The dataset's main institute | Dice | 0.857 | 0.836 to 0.874 | Held-out test: 89 CT scans (a scan can sit in more than one group); 55 scans | held-out test |
| All other institutes | Dice | 0.887 | 0.873 to 0.898 | Held-out test: 89 CT scans (a scan can sit in more than one group); 34 scans | held-out test |
| Scans with marked pathology | Dice | 0.866 | 0.851 to 0.879 | Held-out test: 89 CT scans (a scan can sit in more than one group); 73 scans | held-out test |
Test set: Held-out test: 89 CT scans (a scan can sit in more than one group).
Where it falls short
- Both kidney cysts score 0.0000 on the test scans. Vertebra C6, the gallbladder, the portal and splenic veins and the adrenal glands are the next weakest.
- It outlines a structure that is not in the scan about 4% of the time on the test scans and about 13% of the time on the development scans.
- Most test scans come from one institute and one scanner maker, and the score is lower for that institute (0.8573) than for the others (0.8866).
Trained and tested on
- TotalSegmentator dataset v2.0.1 (1,228 CT scans with masks of 117 structures). Licence: CC BY 4.0
Weights
Not released yet.
MRI
Disc grading research model (MRI)
Three networks trained to give each lower-spine disc on a sagittal MRI a Pfirrmann grade from 1 to 5 and a score for seven wear label types.
Research results on public test data. They say nothing about your own scan.
Results
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Agreement with SPIDER's Pfirrmann grades, held-out test | kappa | 0.773 | 0.664 to 0.853 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
| Share of discs given SPIDER's exact grade, development patients | percent | 59.6 | 55.5 to 63.4 | SPIDER development patients, out of fold: 168 patients, 1,004 discs | development |
Exact values: Agreement with SPIDER's Pfirrmann grades, held-out test: kappa 0.773 (95% interval 0.664 to 0.853); Share of discs given SPIDER's exact grade, development patients: percent 59.6 (95% interval 55.5 to 63.4).
Agreement with SPIDER's labels, seven label types, earlier version 1
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Herniation (per disc) | AUROC | 0.818 | 0.710 to 0.912 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
| Narrowing | AUROC | 0.956 | 0.930 to 0.980 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
| Bulging | AUROC | 0.874 | 0.824 to 0.919 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
| Slipped vertebra | AUROC | 0.887 | 0.789 to 0.964 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
| Upper endplate change | AUROC | 0.817 | 0.734 to 0.893 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
| Lower endplate change | AUROC | 0.850 | 0.784 to 0.908 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
| Any Modic change | AUROC | 0.832 | 0.743 to 0.911 | SPIDER held-out test: 44 patients, 264 discs, opened once | held-out test |
Test set: SPIDER held-out test: 44 patients, 264 discs, opened once.
Where it falls short
- The three current versions have not been scored on held-out test patients; their results are development results.
- Herniation is the weak spot: for the dataset's herniation label per patient, the three versions together agree on 53% (38% to 70%) of patients with one and 74% of patients without one.
- A new hospital can cost it most of its skill: version 1 scored kappa 0.42 on a collection it had never seen, against 0.78 for the version trained on that collection's development discs.
Trained and tested on
- SPIDER. Licence: CC BY 4.0
- LSMA-PQR v2. Licence: CC BY 4.0
- Lumbar Spine MRI Dataset v2. Licence: CC BY 4.0
- Multi-Disorder Annotations for Lumbar Spine Mid-Sagittal Images v1. Licence: CC BY 4.0
Weights
Not released yet for versions 133, 134 and 140. Version 1 already has a download.
X-ray
Fracture research model (X-ray)
Two sets of five small networks trained to match a dataset's fracture labels on one wrist, hand, leg, hip or shoulder X-ray.
Research results on public test data. They say nothing about your own scan.
Results
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Agreement with GRAZPEDWRI-DX's fracture labels, children's wrists | AUROC | 0.979 | 0.974 to 0.984 | GRAZPEDWRI-DX held-out test: 4,067 pictures of 1,176 patients with a certain label | held-out test |
| Agreement with FracAtlas's fracture labels, mixed bones | AUROC | 0.918 | 0.888 to 0.945 | FracAtlas held-out test: 798 pictures | held-out test |
| Agreement with GRAZPEDWRI-DX's fracture labels, first visits only | AUROC | 0.969 | 0.960 to 0.978 | GRAZPEDWRI-DX held-out test, first visits: 2,172 pictures of 990 patients | held-out test |
Exact values: Agreement with GRAZPEDWRI-DX's fracture labels, children's wrists: AUROC 0.979 (95% interval 0.974 to 0.984); Agreement with FracAtlas's fracture labels, mixed bones: AUROC 0.918 (95% interval 0.888 to 0.945); Agreement with GRAZPEDWRI-DX's fracture labels, first visits only: AUROC 0.969 (95% interval 0.96 to 0.978).
Agreement with FracAtlas's fracture labels, by body part
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Leg | AUROC | 0.922 | 0.877 to 0.961 | FracAtlas held-out test (hand not given with an interval in the card); 445 pictures | held-out test |
| Hip | AUROC | 0.933 | 0.856 to 0.988 | FracAtlas held-out test (hand not given with an interval in the card); 67 pictures | held-out test |
| Shoulder | AUROC | 0.692 | 0.494 to 0.863 | FracAtlas held-out test (hand not given with an interval in the card); 65 pictures | held-out test |
Test set: FracAtlas held-out test (hand not given with an interval in the card).
Where it falls short
- Shoulders are weak: on 65 FracAtlas test pictures the AUROC is 0.692 (0.494 to 0.863), from very few pictures, so the true figure is uncertain.
- Hands raise many false alarms: on 299 FracAtlas hand pictures it agrees with the no-fracture label on 0.645 of them, so roughly a third disagree.
- It learned mostly from children's wrists (16,157 of 19,383 pictures) and started from a network trained on chest X-rays, not bones.
Trained and tested on
- GRAZPEDWRI-DX v2. Licence: CC BY 4.0
- FracAtlas v7. Licence: CC BY 4.0
- NIH ChestX-ray14 (pre-training only, through the chest X-ray research model). Licence: NIH asks for credit; use is unrestricted
Weights
Not released yet.
CT
Lung nodule research model (CT)
A chain of 3D networks trained to list places in a chest CT scan that look like small round spots (nodules), plus one network trained on LIDC-IDRI's malignancy ratings of nodules.
Research results on public test data. They say nothing about your own scan.
Results
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Share of LUNA16's marked nodules matched, averaged over seven false-alarm allowances | percent | 88.0 | 86.0 to 89.6 | LUNA16: 888 thin-slice LIDC-IDRI scans, five folds, each read by a model that never trained on it | development |
| Agreement with LIDC-IDRI's malignancy ratings, per patient | AUROC | 0.717 | 0.599 to 0.827 | 117 LIDC patients (86 cancers, 31 benign), out of fold | development |
| Agreement with LIDC-IDRI's malignancy ratings, per nodule | AUROC | 0.949 | 0.930 to 0.965 | 749 LIDC nodules, out of fold | development |
Exact values: Share of LUNA16's marked nodules matched, averaged over seven false-alarm allowances: percent 88.0 (95% interval 86.0 to 89.6); Agreement with LIDC-IDRI's malignancy ratings, per patient: AUROC 0.717 (95% interval 0.599 to 0.827); Agreement with LIDC-IDRI's malignancy ratings, per nodule: AUROC 0.9485 (95% interval 0.93 to 0.9652).
Where it falls short
- All the nodule-matching numbers are development scores from the same 888 scans used to choose between designs; there is no untouched test.
- Per patient, the largest nodule's diameter alone scores 0.755, higher than the model's 0.717.
- At one false alarm in eight scans it matches 70.5% of LUNA16's marked nodules under the strict reading.
Trained and tested on
- LIDC-IDRI. Licence: CC BY 3.0
- LUNA16 (answer key and scoring script). Licence: CC BY 4.0
- LIDC-annot-NLST501. Licence: CC BY 4.0
- NLST-Sybil (label corrections only). Licence: CC BY 4.0
- SPIE-AAPM Lung CT Challenge (LUNGx), practice scans checked only. Licence: CC BY 3.0
Weights
Not released yet.
Mammogram
Mammogram research model
Five versions of a network trained to match a dataset's benign/malignant labels for a lesion already marked on a mammogram.
Research results on public test data. They say nothing about your own scan.
Results
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Agreement with CMMD's benign/malignant labels, per picture | AUROC | 0.793 | 0.774 to 0.811 | CMMD: 3,744 pictures of 1,775 patients, five folds split by patient | development |
| Agreement with CMMD's benign/malignant labels, per breast | AUROC | 0.815 | 0.794 to 0.835 | CMMD: 1,872 breasts, five folds split by patient | development |
| Agreement with CBIS-DDSM's benign/malignant labels, masses, per lesion and view | AUROC | 0.829 | 0.775 to 0.873 | CBIS-DDSM official test list: 378 masses, opened once | held-out test |
Exact values: Agreement with CMMD's benign/malignant labels, per picture: AUROC 0.793 (95% interval 0.774 to 0.811); Agreement with CMMD's benign/malignant labels, per breast: AUROC 0.8147 (95% interval 0.7945 to 0.8346); Agreement with CBIS-DDSM's benign/malignant labels, masses, per lesion and view: AUROC 0.829 (95% interval 0.775 to 0.873).
Where it falls short
- Masses are the weak spot: they score 0.7371 against 0.8424 for calcifications, and masses are 2,298 of the 3,744 pictures.
- Where no lesion outline exists the score is close to a coin toss: version 38 scores 0.5726 on those 1,002 pictures.
- Every CMMD number comes from CMMD itself. There is no held-out test and no second hospital, so the figure may be a little high for new pictures.
Trained and tested on
- CMMD. Licence: CC BY 4.0
- TOMPEI-CMMD (lesion outlines, training only). Licence: CC BY 4.0
- CBIS-DDSM. Licence: CC BY 3.0
- OMAMA-DB. Licence: MIT
- NIH ChestX-ray14 (start of the training chain). Licence: NIH asks for credit; use is unrestricted
Weights
Not released yet.
X-ray
Lumbar spine X-ray corner reader
Small networks trained to place the corners of the lower-back vertebrae L1 to L5 on a spine X-ray, one for the side view and one for the front view.
Research results on public test data. They say nothing about your own scan.
Results
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Share of BUU-LSPINE's marked vertebrae matched, side view (IoU 0.5) | percent | 97.9 | 97.3 to 98.4 | All 2,000 BUU-LSPINE patients, out of fold | development |
| Share of BUU-LSPINE's marked vertebrae matched, front view (IoU 0.5) | percent | 95.8 | 95.0 to 96.6 | All 2,000 BUU-LSPINE patients, out of fold | development |
| Share of the second set's marked corners within 5 pixels, side view | percent | 64.3 | 61.7 to 66.9 | Outside set: 590 side-view pictures from the Spondylolisthesis Vertebral Landmark set, no network trained on it, scored once | held-out test |
Exact values: Share of BUU-LSPINE's marked vertebrae matched, side view (IoU 0.5): percent 97.9 (95% interval 97.3 to 98.4); Share of BUU-LSPINE's marked vertebrae matched, front view (IoU 0.5): percent 95.8 (95% interval 95.0 to 96.6); Share of the second set's marked corners within 5 pixels, side view: percent 64.3 (95% interval 61.7 to 66.9).
Share of BUU-LSPINE's marked vertebrae matched, development patients
Show the numbers
| Result | Metric | Score | 95% interval | Test set | Kind |
|---|---|---|---|---|---|
| Side view, IoU 0.5 | percent | 97.9 | 97.3 to 98.4 | All 2,000 BUU-LSPINE patients, out of fold; 2,000 patients | development |
| Side view, stricter reading (0.50 to 0.95) | percent | 89.4 | 88.8 to 89.9 | All 2,000 BUU-LSPINE patients, out of fold; 2,000 patients | development |
| Front view, IoU 0.5 | percent | 95.8 | 95.0 to 96.6 | All 2,000 BUU-LSPINE patients, out of fold; 2,000 patients | development |
| Front view, stricter reading (0.50 to 0.95) | percent | 81.2 | 80.4 to 81.9 | All 2,000 BUU-LSPINE patients, out of fold; 2,000 patients | development |
Test set: All 2,000 BUU-LSPINE patients, out of fold.
Where it falls short
- On the second set it places 64.3% of corners within 5 pixels; the pictures of the Burapha patients score 71.6% and those of the Honduran patients score 18.3%.
- Everything it learned from comes from one dataset collected in Thailand.
- The average miss is bigger than the typical miss: on the side view 11.78 pixels against a median patient of 7.10.
Trained and tested on
- BUU-LSPINE, 2,000-patient release (Burapha University, Thailand). Licence: Burapha University End-User License Agreement, accepted before download
- Spondylolisthesis Vertebral Landmark Dataset v1 (testing only). Licence: CC BY 4.0
Weights
Not released yet.