rAidiology

Research results

Seven open research models across X-ray, CT, MRI and mammogram, trained on public data, for education and research.

Research results on public test data. They say nothing about your own scan.

How to read the charts

  • AUROC: 0.5 is a coin toss and 1.0 is perfect.
  • Dice: how much the outline overlaps the dataset's own outlines; 1.0 is a perfect match.
  • Kappa: 0 is chance agreement and 1.0 is full agreement.
  • Dots: a filled dot is a held-out test, a hollow dot is development. The line through a dot is its 95% interval.
Chest X-ray, a public dataset picture used as an example of this scan type.
Chest X-ray. NIH ChestX-ray14, No restrictions (NIH), attribution requested. Source

X-ray

Chest X-ray research model

A team of ten networks trained to give 15 scores for one frontal chest X-ray, one for each of the 14 NIH ChestX-ray14 label types and one for its abnormal label.

Research results on public test data. They say nothing about your own scan.

Results

AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with ChestX-ray14's labels, mean over its 14 label types: 0.818 (95% interval 0.811 to 0.823); Agreement with ChestX-ray14's abnormal label: 0.736 (95% interval 0.724 to 0.748). Research results on public test data.AUROC, 0.5 to 1Held-out testAgreement with ChestX-ray14'slabels, mean over its 14 labeltypes0.818 (0.811 to 0.823)Agreement with ChestX-ray14'sabnormal label0.736 (0.724 to 0.748)0.50.751coin toss (0.5)AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with ChestX-ray14's labels, mean over its 14 label types: 0.818 (95% interval 0.811 to 0.823); Agreement with ChestX-ray14's abnormal label: 0.736 (95% interval 0.724 to 0.748). Research results on public test data.AUROC, 0.5 to 1Held-out testAgreement with ChestX-ray14's labels, meanover its 14 labeltypes0.818 (0.811 to 0.823)Agreement with ChestX-ray14's abnormal label0.736 (0.724 to 0.748)0.50.751coin toss (0.5)
Show the numbers
ResultMetricScore95% intervalTest setKind
Agreement with ChestX-ray14's labels, mean over its 14 label typesAUROC0.8180.811 to 0.823NIH's official test list: 25,596 pictures of 2,797 patientsheld-out test
Agreement with ChestX-ray14's abnormal labelAUROC0.7360.724 to 0.748NIH's official test list: 25,596 picturesheld-out test

Exact values: Agreement with ChestX-ray14's labels, mean over its 14 label types: AUROC 0.8175 (95% interval 0.811 to 0.8231); Agreement with ChestX-ray14's abnormal label: AUROC 0.7363 (95% interval 0.7241 to 0.7477).

Agreement with ChestX-ray14's labels, per label type

Agreement with ChestX-ray14's labels, per label typeDot and interval chart. Filled dot: held-out test. Hollow dot: development. Atelectasis: 0.775 (95% interval 0.762 to 0.788); Cardiomegaly: 0.902 (95% interval 0.886 to 0.914); Effusion: 0.839 (95% interval 0.828 to 0.849); Infiltration: 0.707 (95% interval 0.696 to 0.718); Mass: 0.824 (95% interval 0.803 to 0.842); Nodule: 0.754 (95% interval 0.733 to 0.774); Pneumonia: 0.727 (95% interval 0.702 to 0.750); Pneumothorax: 0.861 (95% interval 0.845 to 0.877); Consolidation: 0.761 (95% interval 0.746 to 0.775); Edema: 0.854 (95% interval 0.838 to 0.870); Emphysema: 0.896 (95% interval 0.880 to 0.910); Fibrosis: 0.838 (95% interval 0.815 to 0.861); Pleural thickening: 0.784 (95% interval 0.766 to 0.804); Hernia: 0.923 (95% interval 0.878 to 0.957). Research results on public test data.Agreement with ChestX-ray14's labels, per label typeHeld-out testAtelectasis0.775 (0.762 to 0.788)Cardiomegaly0.902 (0.886 to 0.914)Effusion0.839 (0.828 to 0.849)Infiltration0.707 (0.696 to 0.718)Mass0.824 (0.803 to 0.842)Nodule0.754 (0.733 to 0.774)Pneumonia0.727 (0.702 to 0.750)Pneumothorax0.861 (0.845 to 0.877)Consolidation0.761 (0.746 to 0.775)Edema0.854 (0.838 to 0.870)Emphysema0.896 (0.880 to 0.910)Fibrosis0.838 (0.815 to 0.861)Pleural thickening0.784 (0.766 to 0.804)Hernia0.923 (0.878 to 0.957)0.50.751coin toss (0.5)Agreement with ChestX-ray14's labels, per label typeDot and interval chart. Filled dot: held-out test. Hollow dot: development. Atelectasis: 0.775 (95% interval 0.762 to 0.788); Cardiomegaly: 0.902 (95% interval 0.886 to 0.914); Effusion: 0.839 (95% interval 0.828 to 0.849); Infiltration: 0.707 (95% interval 0.696 to 0.718); Mass: 0.824 (95% interval 0.803 to 0.842); Nodule: 0.754 (95% interval 0.733 to 0.774); Pneumonia: 0.727 (95% interval 0.702 to 0.750); Pneumothorax: 0.861 (95% interval 0.845 to 0.877); Consolidation: 0.761 (95% interval 0.746 to 0.775); Edema: 0.854 (95% interval 0.838 to 0.870); Emphysema: 0.896 (95% interval 0.880 to 0.910); Fibrosis: 0.838 (95% interval 0.815 to 0.861); Pleural thickening: 0.784 (95% interval 0.766 to 0.804); Hernia: 0.923 (95% interval 0.878 to 0.957). Research results on public test data.Agreement with ChestX-ray14's labels, per label typeHeld-out testAtelectasis0.775 (0.762 to 0.788)Cardiomegaly0.902 (0.886 to 0.914)Effusion0.839 (0.828 to 0.849)Infiltration0.707 (0.696 to 0.718)Mass0.824 (0.803 to 0.842)Nodule0.754 (0.733 to 0.774)Pneumonia0.727 (0.702 to 0.750)Pneumothorax0.861 (0.845 to 0.877)Consolidation0.761 (0.746 to 0.775)Edema0.854 (0.838 to 0.870)Emphysema0.896 (0.880 to 0.910)Fibrosis0.838 (0.815 to 0.861)Pleural thickening0.784 (0.766 to 0.804)Hernia0.923 (0.878 to 0.957)0.50.751coin toss (0.5)
Show the numbers
ResultMetricScore95% intervalTest setKind
AtelectasisAUROC0.7750.762 to 0.788NIH's official test list, 25,596 pictures; 3,279 picturesheld-out test
CardiomegalyAUROC0.9020.886 to 0.914NIH's official test list, 25,596 pictures; 1,069 picturesheld-out test
EffusionAUROC0.8390.828 to 0.849NIH's official test list, 25,596 pictures; 4,658 picturesheld-out test
InfiltrationAUROC0.7070.696 to 0.718NIH's official test list, 25,596 pictures; 6,112 picturesheld-out test
MassAUROC0.8240.803 to 0.842NIH's official test list, 25,596 pictures; 1,748 picturesheld-out test
NoduleAUROC0.7540.733 to 0.774NIH's official test list, 25,596 pictures; 1,623 picturesheld-out test
PneumoniaAUROC0.7270.702 to 0.750NIH's official test list, 25,596 pictures; 555 picturesheld-out test
PneumothoraxAUROC0.8610.845 to 0.877NIH's official test list, 25,596 pictures; 2,665 picturesheld-out test
ConsolidationAUROC0.7610.746 to 0.775NIH's official test list, 25,596 pictures; 1,815 picturesheld-out test
EdemaAUROC0.8540.838 to 0.870NIH's official test list, 25,596 pictures; 925 picturesheld-out test
EmphysemaAUROC0.8960.880 to 0.910NIH's official test list, 25,596 pictures; 1,093 picturesheld-out test
FibrosisAUROC0.8380.815 to 0.861NIH's official test list, 25,596 pictures; 435 picturesheld-out test
Pleural thickeningAUROC0.7840.766 to 0.804NIH's official test list, 25,596 pictures; 1,143 picturesheld-out test
HerniaAUROC0.9230.878 to 0.957NIH's official test list, 25,596 pictures; 86 picturesheld-out test

Test set: NIH's official test list, 25,596 pictures.

Where it falls short

  • Agreement is lower for some label types: infiltration scores 0.7071, pneumonia 0.7267 and nodule 0.7544.
  • Rare label types have wide intervals. Hernia has only 86 test pictures, and its interval runs from 0.8776 to 0.9574.
  • All the labelled pictures come from one hospital, so nothing shows how it behaves on other scanners, other patient groups or children.

Trained and tested on

Weights

Not released yet.

Read the model card on GitHub

CT, organs and bones coloured, a public dataset picture used as an example of this scan type.
CT, organs and bones coloured. TotalSegmentator dataset v2.0.1, CC BY 4.0. Source

CT

CT organ outliner

A two-step network trained to label each point of a 3D CT scan as one of 117 organs and bones, such as the liver, spleen, aorta, ribs and vertebrae, or as nothing.

Research results on public test data. They say nothing about your own scan.

Results

Dice overlap, 0 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Mean over 115 structures: 0.867 (95% interval 0.854 to 0.878); Mean over 115 structures, development scans: 0.816 (95% interval 0.812 to 0.819). Research results on public test data.Dice overlap, 0 to 1Held-out testMean over 115 structures0.867 (0.854 to 0.878)DevelopmentMean over 115 structures,development scans0.816 (0.812 to 0.819)00.51Dice overlap, 0 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Mean over 115 structures: 0.867 (95% interval 0.854 to 0.878); Mean over 115 structures, development scans: 0.816 (95% interval 0.812 to 0.819). Research results on public test data.Dice overlap, 0 to 1Held-out testMean over 115structures0.867 (0.854 to 0.878)DevelopmentMean over 115structures,development scans0.816 (0.812 to 0.819)00.51
Show the numbers
ResultMetricScore95% intervalTest setKind
Mean over 115 structuresDice0.8670.854 to 0.878Held-out test: 89 CT scansheld-out test
Mean over 115 structures, development scansDice0.8160.812 to 0.8191,139 development scans, each read by models that never saw itdevelopment

Exact values: Mean over 115 structures: Dice 0.8668 (95% interval 0.8538 to 0.8781); Mean over 115 structures, development scans: Dice 0.8156 (95% interval 0.8116 to 0.8194).

Mean Dice by group of scans

Mean Dice by group of scansDot and interval chart. Filled dot: held-out test. Hollow dot: development. All: 0.867 (95% interval 0.854 to 0.878); Abdomen scans: 0.880 (95% interval 0.868 to 0.891); Thorax scans: 0.888 (95% interval 0.876 to 0.899); Head or neck scans: 0.877 (95% interval 0.855 to 0.895); Siemens scanners: 0.871 (95% interval 0.853 to 0.886); Other scanner makers: 0.864 (95% interval 0.845 to 0.881); The dataset's main institute: 0.857 (95% interval 0.836 to 0.874); All other institutes: 0.887 (95% interval 0.873 to 0.898); Scans with marked pathology: 0.866 (95% interval 0.851 to 0.879). Research results on public test data.Mean Dice by group of scansHeld-out testAll0.867 (0.854 to 0.878)Abdomen scans0.880 (0.868 to 0.891)Thorax scans0.888 (0.876 to 0.899)Head or neck scans0.877 (0.855 to 0.895)Siemens scanners0.871 (0.853 to 0.886)Other scanner makers0.864 (0.845 to 0.881)The dataset's main institute0.857 (0.836 to 0.874)All other institutes0.887 (0.873 to 0.898)Scans with marked pathology0.866 (0.851 to 0.879)00.51Mean Dice by group of scansDot and interval chart. Filled dot: held-out test. Hollow dot: development. All: 0.867 (95% interval 0.854 to 0.878); Abdomen scans: 0.880 (95% interval 0.868 to 0.891); Thorax scans: 0.888 (95% interval 0.876 to 0.899); Head or neck scans: 0.877 (95% interval 0.855 to 0.895); Siemens scanners: 0.871 (95% interval 0.853 to 0.886); Other scanner makers: 0.864 (95% interval 0.845 to 0.881); The dataset's main institute: 0.857 (95% interval 0.836 to 0.874); All other institutes: 0.887 (95% interval 0.873 to 0.898); Scans with marked pathology: 0.866 (95% interval 0.851 to 0.879). Research results on public test data.Mean Dice by group of scansHeld-out testAll0.867 (0.854 to 0.878)Abdomen scans0.880 (0.868 to 0.891)Thorax scans0.888 (0.876 to 0.899)Head or neck scans0.877 (0.855 to 0.895)Siemens scanners0.871 (0.853 to 0.886)Other scanner makers0.864 (0.845 to 0.881)The dataset's maininstitute0.857 (0.836 to 0.874)All other institutes0.887 (0.873 to 0.898)Scans with markedpathology0.866 (0.851 to 0.879)00.51
Show the numbers
ResultMetricScore95% intervalTest setKind
AllDice0.8670.854 to 0.878Held-out test: 89 CT scans (a scan can sit in more than one group); 89 scansheld-out test
Abdomen scansDice0.8800.868 to 0.891Held-out test: 89 CT scans (a scan can sit in more than one group); 47 scansheld-out test
Thorax scansDice0.8880.876 to 0.899Held-out test: 89 CT scans (a scan can sit in more than one group); 32 scansheld-out test
Head or neck scansDice0.8770.855 to 0.895Held-out test: 89 CT scans (a scan can sit in more than one group); 19 scansheld-out test
Siemens scannersDice0.8710.853 to 0.886Held-out test: 89 CT scans (a scan can sit in more than one group); 58 scansheld-out test
Other scanner makersDice0.8640.845 to 0.881Held-out test: 89 CT scans (a scan can sit in more than one group); 31 scansheld-out test
The dataset's main instituteDice0.8570.836 to 0.874Held-out test: 89 CT scans (a scan can sit in more than one group); 55 scansheld-out test
All other institutesDice0.8870.873 to 0.898Held-out test: 89 CT scans (a scan can sit in more than one group); 34 scansheld-out test
Scans with marked pathologyDice0.8660.851 to 0.879Held-out test: 89 CT scans (a scan can sit in more than one group); 73 scansheld-out test

Test set: Held-out test: 89 CT scans (a scan can sit in more than one group).

Where it falls short

  • Both kidney cysts score 0.0000 on the test scans. Vertebra C6, the gallbladder, the portal and splenic veins and the adrenal glands are the next weakest.
  • It outlines a structure that is not in the scan about 4% of the time on the test scans and about 13% of the time on the development scans.
  • Most test scans come from one institute and one scanner maker, and the score is lower for that institute (0.8573) than for the others (0.8866).

Trained and tested on

Weights

Not released yet.

Read the model card on GitHub

MRI of the lumbar spine, vertebrae and discs coloured, a public dataset picture used as an example of this scan type.
MRI of the lumbar spine, vertebrae and discs coloured. SPIDER lumbar spine MRI, CC BY 4.0. Source

MRI

Disc grading research model (MRI)

Three networks trained to give each lower-spine disc on a sagittal MRI a Pfirrmann grade from 1 to 5 and a score for seven wear label types.

Research results on public test data. They say nothing about your own scan.

Results

Kappa agreement, 0 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with SPIDER's Pfirrmann grades, held-out test: 0.773 (95% interval 0.664 to 0.853). Research results on public test data.Kappa agreement, 0 to 1Held-out testAgreement with SPIDER's Pfirrmanngrades, held-out test0.773 (0.664 to 0.853)00.510 = chance agreementKappa agreement, 0 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with SPIDER's Pfirrmann grades, held-out test: 0.773 (95% interval 0.664 to 0.853). Research results on public test data.Kappa agreement, 0 to 1Held-out testAgreement withSPIDER's Pfirrmanngrades, held-out test0.773 (0.664 to 0.853)00.510 = chance agreement
Percent, 0 to 100Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Share of discs given SPIDER's exact grade, development patients: 59.6 (95% interval 55.5 to 63.4). Research results on public test data.Percent, 0 to 100DevelopmentShare of discs given SPIDER'sexact grade, development patients59.6% (55.5 to 63.4%)050100Percent, 0 to 100Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Share of discs given SPIDER's exact grade, development patients: 59.6 (95% interval 55.5 to 63.4). Research results on public test data.Percent, 0 to 100DevelopmentShare of discs givenSPIDER's exact grade,development patients59.6% (55.5 to 63.4%)050100
Show the numbers
ResultMetricScore95% intervalTest setKind
Agreement with SPIDER's Pfirrmann grades, held-out testkappa0.7730.664 to 0.853SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test
Share of discs given SPIDER's exact grade, development patientspercent59.655.5 to 63.4SPIDER development patients, out of fold: 168 patients, 1,004 discsdevelopment

Exact values: Agreement with SPIDER's Pfirrmann grades, held-out test: kappa 0.773 (95% interval 0.664 to 0.853); Share of discs given SPIDER's exact grade, development patients: percent 59.6 (95% interval 55.5 to 63.4).

Agreement with SPIDER's labels, seven label types, earlier version 1

Agreement with SPIDER's labels, seven label types, earlier version 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Herniation (per disc): 0.818 (95% interval 0.710 to 0.912); Narrowing: 0.956 (95% interval 0.930 to 0.980); Bulging: 0.874 (95% interval 0.824 to 0.919); Slipped vertebra: 0.887 (95% interval 0.789 to 0.964); Upper endplate change: 0.817 (95% interval 0.734 to 0.893); Lower endplate change: 0.850 (95% interval 0.784 to 0.908); Any Modic change: 0.832 (95% interval 0.743 to 0.911). Research results on public test data.Agreement with SPIDER's labels, seven label types, earlier version 1Held-out testHerniation (per disc)0.818 (0.710 to 0.912)Narrowing0.956 (0.930 to 0.980)Bulging0.874 (0.824 to 0.919)Slipped vertebra0.887 (0.789 to 0.964)Upper endplate change0.817 (0.734 to 0.893)Lower endplate change0.850 (0.784 to 0.908)Any Modic change0.832 (0.743 to 0.911)0.50.751coin toss (0.5)Agreement with SPIDER's labels, seven label types, earlier version 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Herniation (per disc): 0.818 (95% interval 0.710 to 0.912); Narrowing: 0.956 (95% interval 0.930 to 0.980); Bulging: 0.874 (95% interval 0.824 to 0.919); Slipped vertebra: 0.887 (95% interval 0.789 to 0.964); Upper endplate change: 0.817 (95% interval 0.734 to 0.893); Lower endplate change: 0.850 (95% interval 0.784 to 0.908); Any Modic change: 0.832 (95% interval 0.743 to 0.911). Research results on public test data.Agreement with SPIDER's labels, seven label types, earlier version 1Held-out testHerniation (per disc)0.818 (0.710 to 0.912)Narrowing0.956 (0.930 to 0.980)Bulging0.874 (0.824 to 0.919)Slipped vertebra0.887 (0.789 to 0.964)Upper endplate change0.817 (0.734 to 0.893)Lower endplate change0.850 (0.784 to 0.908)Any Modic change0.832 (0.743 to 0.911)0.50.751coin toss (0.5)
Show the numbers
ResultMetricScore95% intervalTest setKind
Herniation (per disc)AUROC0.8180.710 to 0.912SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test
NarrowingAUROC0.9560.930 to 0.980SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test
BulgingAUROC0.8740.824 to 0.919SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test
Slipped vertebraAUROC0.8870.789 to 0.964SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test
Upper endplate changeAUROC0.8170.734 to 0.893SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test
Lower endplate changeAUROC0.8500.784 to 0.908SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test
Any Modic changeAUROC0.8320.743 to 0.911SPIDER held-out test: 44 patients, 264 discs, opened onceheld-out test

Test set: SPIDER held-out test: 44 patients, 264 discs, opened once.

Where it falls short

  • The three current versions have not been scored on held-out test patients; their results are development results.
  • Herniation is the weak spot: for the dataset's herniation label per patient, the three versions together agree on 53% (38% to 70%) of patients with one and 74% of patients without one.
  • A new hospital can cost it most of its skill: version 1 scored kappa 0.42 on a collection it had never seen, against 0.78 for the version trained on that collection's development discs.

Trained and tested on

Weights

Not released yet for versions 133, 134 and 140. Version 1 already has a download.

Read the model card on GitHub

X-ray of a child's wrist, a public dataset picture used as an example of this scan type.
X-ray of a child's wrist. GRAZPEDWRI-DX v2, CC BY 4.0. Source

X-ray

Fracture research model (X-ray)

Two sets of five small networks trained to match a dataset's fracture labels on one wrist, hand, leg, hip or shoulder X-ray.

Research results on public test data. They say nothing about your own scan.

Results

AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with GRAZPEDWRI-DX's fracture labels, children's wrists: 0.979 (95% interval 0.974 to 0.984); Agreement with FracAtlas's fracture labels, mixed bones: 0.918 (95% interval 0.888 to 0.945); Agreement with GRAZPEDWRI-DX's fracture labels, first visits only: 0.969 (95% interval 0.960 to 0.978). Research results on public test data.AUROC, 0.5 to 1Held-out testAgreement with GRAZPEDWRI-DX'sfracture labels, children'swrists0.979 (0.974 to 0.984)Agreement with FracAtlas'sfracture labels, mixed bones0.918 (0.888 to 0.945)Agreement with GRAZPEDWRI-DX'sfracture labels, first visitsonly0.969 (0.960 to 0.978)0.50.751coin toss (0.5)AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with GRAZPEDWRI-DX's fracture labels, children's wrists: 0.979 (95% interval 0.974 to 0.984); Agreement with FracAtlas's fracture labels, mixed bones: 0.918 (95% interval 0.888 to 0.945); Agreement with GRAZPEDWRI-DX's fracture labels, first visits only: 0.969 (95% interval 0.960 to 0.978). Research results on public test data.AUROC, 0.5 to 1Held-out testAgreement withGRAZPEDWRI-DX'sfracture labels,children's wrists0.979 (0.974 to 0.984)Agreement withFracAtlas's fracturelabels, mixed bones0.918 (0.888 to 0.945)Agreement withGRAZPEDWRI-DX'sfracture labels, firstvisits only0.969 (0.960 to 0.978)0.50.751coin toss (0.5)
Show the numbers
ResultMetricScore95% intervalTest setKind
Agreement with GRAZPEDWRI-DX's fracture labels, children's wristsAUROC0.9790.974 to 0.984GRAZPEDWRI-DX held-out test: 4,067 pictures of 1,176 patients with a certain labelheld-out test
Agreement with FracAtlas's fracture labels, mixed bonesAUROC0.9180.888 to 0.945FracAtlas held-out test: 798 picturesheld-out test
Agreement with GRAZPEDWRI-DX's fracture labels, first visits onlyAUROC0.9690.960 to 0.978GRAZPEDWRI-DX held-out test, first visits: 2,172 pictures of 990 patientsheld-out test

Exact values: Agreement with GRAZPEDWRI-DX's fracture labels, children's wrists: AUROC 0.979 (95% interval 0.974 to 0.984); Agreement with FracAtlas's fracture labels, mixed bones: AUROC 0.918 (95% interval 0.888 to 0.945); Agreement with GRAZPEDWRI-DX's fracture labels, first visits only: AUROC 0.969 (95% interval 0.96 to 0.978).

Agreement with FracAtlas's fracture labels, by body part

Agreement with FracAtlas's fracture labels, by body partDot and interval chart. Filled dot: held-out test. Hollow dot: development. Leg: 0.922 (95% interval 0.877 to 0.961); Hip: 0.933 (95% interval 0.856 to 0.988); Shoulder: 0.692 (95% interval 0.494 to 0.863). Research results on public test data.Agreement with FracAtlas's fracture labels, by body partHeld-out testLeg0.922 (0.877 to 0.961)Hip0.933 (0.856 to 0.988)Shoulder0.692 (0.494 to 0.863)0.50.751coin toss (0.5)Agreement with FracAtlas's fracture labels, by body partDot and interval chart. Filled dot: held-out test. Hollow dot: development. Leg: 0.922 (95% interval 0.877 to 0.961); Hip: 0.933 (95% interval 0.856 to 0.988); Shoulder: 0.692 (95% interval 0.494 to 0.863). Research results on public test data.Agreement with FracAtlas's fracture labels, by body partHeld-out testLeg0.922 (0.877 to 0.961)Hip0.933 (0.856 to 0.988)Shoulder0.692 (0.494 to 0.863)0.50.751coin toss (0.5)
Show the numbers
ResultMetricScore95% intervalTest setKind
LegAUROC0.9220.877 to 0.961FracAtlas held-out test (hand not given with an interval in the card); 445 picturesheld-out test
HipAUROC0.9330.856 to 0.988FracAtlas held-out test (hand not given with an interval in the card); 67 picturesheld-out test
ShoulderAUROC0.6920.494 to 0.863FracAtlas held-out test (hand not given with an interval in the card); 65 picturesheld-out test

Test set: FracAtlas held-out test (hand not given with an interval in the card).

Where it falls short

  • Shoulders are weak: on 65 FracAtlas test pictures the AUROC is 0.692 (0.494 to 0.863), from very few pictures, so the true figure is uncertain.
  • Hands raise many false alarms: on 299 FracAtlas hand pictures it agrees with the no-fracture label on 0.645 of them, so roughly a third disagree.
  • It learned mostly from children's wrists (16,157 of 19,383 pictures) and started from a network trained on chest X-rays, not bones.

Trained and tested on

Weights

Not released yet.

Read the model card on GitHub

CT of the chest, lung window, a public dataset picture used as an example of this scan type.
CT of the chest, lung window. LIDC-IDRI, CC BY 3.0. Source

CT

Lung nodule research model (CT)

A chain of 3D networks trained to list places in a chest CT scan that look like small round spots (nodules), plus one network trained on LIDC-IDRI's malignancy ratings of nodules.

Research results on public test data. They say nothing about your own scan.

Results

AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with LIDC-IDRI's malignancy ratings, per patient: 0.717 (95% interval 0.599 to 0.827); Agreement with LIDC-IDRI's malignancy ratings, per nodule: 0.949 (95% interval 0.930 to 0.965). Research results on public test data.AUROC, 0.5 to 1DevelopmentAgreement with LIDC-IDRI'smalignancy ratings, per patient0.717 (0.599 to 0.827)Agreement with LIDC-IDRI'smalignancy ratings, per nodule0.949 (0.930 to 0.965)0.50.751coin toss (0.5)AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with LIDC-IDRI's malignancy ratings, per patient: 0.717 (95% interval 0.599 to 0.827); Agreement with LIDC-IDRI's malignancy ratings, per nodule: 0.949 (95% interval 0.930 to 0.965). Research results on public test data.AUROC, 0.5 to 1DevelopmentAgreement with LIDC-IDRI's malignancyratings, per patient0.717 (0.599 to 0.827)Agreement with LIDC-IDRI's malignancyratings, per nodule0.949 (0.930 to 0.965)0.50.751coin toss (0.5)
Percent, 0 to 100Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Share of LUNA16's marked nodules matched, averaged over seven false-alarm allowances: 88.0 (95% interval 86.0 to 89.6). Research results on public test data.Percent, 0 to 100DevelopmentShare of LUNA16's marked nodulesmatched, averaged over sevenfalse-alarm allowances88.0% (86.0 to 89.6%)050100Percent, 0 to 100Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Share of LUNA16's marked nodules matched, averaged over seven false-alarm allowances: 88.0 (95% interval 86.0 to 89.6). Research results on public test data.Percent, 0 to 100DevelopmentShare of LUNA16'smarked nodulesmatched, averaged overseven false-alarmallowances88.0% (86.0 to 89.6%)050100
Show the numbers
ResultMetricScore95% intervalTest setKind
Share of LUNA16's marked nodules matched, averaged over seven false-alarm allowancespercent88.086.0 to 89.6LUNA16: 888 thin-slice LIDC-IDRI scans, five folds, each read by a model that never trained on itdevelopment
Agreement with LIDC-IDRI's malignancy ratings, per patientAUROC0.7170.599 to 0.827117 LIDC patients (86 cancers, 31 benign), out of folddevelopment
Agreement with LIDC-IDRI's malignancy ratings, per noduleAUROC0.9490.930 to 0.965749 LIDC nodules, out of folddevelopment

Exact values: Share of LUNA16's marked nodules matched, averaged over seven false-alarm allowances: percent 88.0 (95% interval 86.0 to 89.6); Agreement with LIDC-IDRI's malignancy ratings, per patient: AUROC 0.717 (95% interval 0.599 to 0.827); Agreement with LIDC-IDRI's malignancy ratings, per nodule: AUROC 0.9485 (95% interval 0.93 to 0.9652).

Where it falls short

  • All the nodule-matching numbers are development scores from the same 888 scans used to choose between designs; there is no untouched test.
  • Per patient, the largest nodule's diameter alone scores 0.755, higher than the model's 0.717.
  • At one false alarm in eight scans it matches 70.5% of LUNA16's marked nodules under the strict reading.

Trained and tested on

Weights

Not released yet.

Read the model card on GitHub

Mammogram, a public dataset picture used as an example of this scan type.
Mammogram. CMMD (Chinese Mammography Database), CC BY 4.0. Source

Mammogram

Mammogram research model

Five versions of a network trained to match a dataset's benign/malignant labels for a lesion already marked on a mammogram.

Research results on public test data. They say nothing about your own scan.

Results

AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with CMMD's benign/malignant labels, per picture: 0.793 (95% interval 0.774 to 0.811); Agreement with CMMD's benign/malignant labels, per breast: 0.815 (95% interval 0.794 to 0.835); Agreement with CBIS-DDSM's benign/malignant labels, masses, per lesion and view: 0.829 (95% interval 0.775 to 0.873). Research results on public test data.AUROC, 0.5 to 1Held-out testAgreement with CBIS-DDSM'sbenign/malignant labels, masses,per lesion and view0.829 (0.775 to 0.873)DevelopmentAgreement with CMMD'sbenign/malignant labels, perpicture0.793 (0.774 to 0.811)Agreement with CMMD'sbenign/malignant labels, perbreast0.815 (0.794 to 0.835)0.50.751coin toss (0.5)AUROC, 0.5 to 1Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Agreement with CMMD's benign/malignant labels, per picture: 0.793 (95% interval 0.774 to 0.811); Agreement with CMMD's benign/malignant labels, per breast: 0.815 (95% interval 0.794 to 0.835); Agreement with CBIS-DDSM's benign/malignant labels, masses, per lesion and view: 0.829 (95% interval 0.775 to 0.873). Research results on public test data.AUROC, 0.5 to 1Held-out testAgreement with CBIS-DDSM'sbenign/malignantlabels, masses, perlesion and view0.829 (0.775 to 0.873)DevelopmentAgreement with CMMD'sbenign/malignantlabels, per picture0.793 (0.774 to 0.811)Agreement with CMMD'sbenign/malignantlabels, per breast0.815 (0.794 to 0.835)0.50.751coin toss (0.5)
Show the numbers
ResultMetricScore95% intervalTest setKind
Agreement with CMMD's benign/malignant labels, per pictureAUROC0.7930.774 to 0.811CMMD: 3,744 pictures of 1,775 patients, five folds split by patientdevelopment
Agreement with CMMD's benign/malignant labels, per breastAUROC0.8150.794 to 0.835CMMD: 1,872 breasts, five folds split by patientdevelopment
Agreement with CBIS-DDSM's benign/malignant labels, masses, per lesion and viewAUROC0.8290.775 to 0.873CBIS-DDSM official test list: 378 masses, opened onceheld-out test

Exact values: Agreement with CMMD's benign/malignant labels, per picture: AUROC 0.793 (95% interval 0.774 to 0.811); Agreement with CMMD's benign/malignant labels, per breast: AUROC 0.8147 (95% interval 0.7945 to 0.8346); Agreement with CBIS-DDSM's benign/malignant labels, masses, per lesion and view: AUROC 0.829 (95% interval 0.775 to 0.873).

Where it falls short

  • Masses are the weak spot: they score 0.7371 against 0.8424 for calcifications, and masses are 2,298 of the 3,744 pictures.
  • Where no lesion outline exists the score is close to a coin toss: version 38 scores 0.5726 on those 1,002 pictures.
  • Every CMMD number comes from CMMD itself. There is no held-out test and no second hospital, so the figure may be a little high for new pictures.

Trained and tested on

Weights

Not released yet.

Read the model card on GitHub

Side X-ray of the lower spine, a public dataset picture used as an example of this scan type.
Side X-ray of the lower spine. Spondylolisthesis Vertebral Landmark Dataset v1 (Reyes, VSB-TU Ostrava), CC BY 4.0. Source

X-ray

Lumbar spine X-ray corner reader

Small networks trained to place the corners of the lower-back vertebrae L1 to L5 on a spine X-ray, one for the side view and one for the front view.

Research results on public test data. They say nothing about your own scan.

Results

Percent, 0 to 100Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Share of BUU-LSPINE's marked vertebrae matched, side view (IoU 0.5): 97.9 (95% interval 97.3 to 98.4); Share of BUU-LSPINE's marked vertebrae matched, front view (IoU 0.5): 95.8 (95% interval 95.0 to 96.6); Share of the second set's marked corners within 5 pixels, side view: 64.3 (95% interval 61.7 to 66.9). Research results on public test data.Percent, 0 to 100Held-out testShare of the second set's markedcorners within 5 pixels, sideview64.3% (61.7 to 66.9%)DevelopmentShare of BUU-LSPINE's markedvertebrae matched, side view (IoU0.5)97.9% (97.3 to 98.4%)Share of BUU-LSPINE's markedvertebrae matched, front view(IoU 0.5)95.8% (95.0 to 96.6%)050100Percent, 0 to 100Dot and interval chart. Filled dot: held-out test. Hollow dot: development. Share of BUU-LSPINE's marked vertebrae matched, side view (IoU 0.5): 97.9 (95% interval 97.3 to 98.4); Share of BUU-LSPINE's marked vertebrae matched, front view (IoU 0.5): 95.8 (95% interval 95.0 to 96.6); Share of the second set's marked corners within 5 pixels, side view: 64.3 (95% interval 61.7 to 66.9). Research results on public test data.Percent, 0 to 100Held-out testShare of the secondset's marked cornerswithin 5 pixels, sideview64.3% (61.7 to 66.9%)DevelopmentShare of BUU-LSPINE'smarked vertebraematched, side view(IoU 0.5)97.9% (97.3 to 98.4%)Share of BUU-LSPINE'smarked vertebraematched, front view(IoU 0.5)95.8% (95.0 to 96.6%)050100
Show the numbers
ResultMetricScore95% intervalTest setKind
Share of BUU-LSPINE's marked vertebrae matched, side view (IoU 0.5)percent97.997.3 to 98.4All 2,000 BUU-LSPINE patients, out of folddevelopment
Share of BUU-LSPINE's marked vertebrae matched, front view (IoU 0.5)percent95.895.0 to 96.6All 2,000 BUU-LSPINE patients, out of folddevelopment
Share of the second set's marked corners within 5 pixels, side viewpercent64.361.7 to 66.9Outside set: 590 side-view pictures from the Spondylolisthesis Vertebral Landmark set, no network trained on it, scored onceheld-out test

Exact values: Share of BUU-LSPINE's marked vertebrae matched, side view (IoU 0.5): percent 97.9 (95% interval 97.3 to 98.4); Share of BUU-LSPINE's marked vertebrae matched, front view (IoU 0.5): percent 95.8 (95% interval 95.0 to 96.6); Share of the second set's marked corners within 5 pixels, side view: percent 64.3 (95% interval 61.7 to 66.9).

Share of BUU-LSPINE's marked vertebrae matched, development patients

Share of BUU-LSPINE's marked vertebrae matched, development patientsDot and interval chart. Filled dot: held-out test. Hollow dot: development. Side view, IoU 0.5: 97.9 (95% interval 97.3 to 98.4); Side view, stricter reading (0.50 to 0.95): 89.4 (95% interval 88.8 to 89.9); Front view, IoU 0.5: 95.8 (95% interval 95.0 to 96.6); Front view, stricter reading (0.50 to 0.95): 81.2 (95% interval 80.4 to 81.9). Research results on public test data.Share of BUU-LSPINE's marked vertebrae matched, development patientsDevelopmentSide view, IoU 0.597.9% (97.3 to 98.4%)Side view, stricter reading (0.50to 0.95)89.4% (88.8 to 89.9%)Front view, IoU 0.595.8% (95.0 to 96.6%)Front view, stricter reading(0.50 to 0.95)81.2% (80.4 to 81.9%)050100Share of BUU-LSPINE's marked vertebrae matched, development patientsDot and interval chart. Filled dot: held-out test. Hollow dot: development. Side view, IoU 0.5: 97.9 (95% interval 97.3 to 98.4); Side view, stricter reading (0.50 to 0.95): 89.4 (95% interval 88.8 to 89.9); Front view, IoU 0.5: 95.8 (95% interval 95.0 to 96.6); Front view, stricter reading (0.50 to 0.95): 81.2 (95% interval 80.4 to 81.9). Research results on public test data.Share of BUU-LSPINE's marked vertebrae matched, development patientsDevelopmentSide view, IoU 0.597.9% (97.3 to 98.4%)Side view, stricterreading (0.50 to 0.95)89.4% (88.8 to 89.9%)Front view, IoU 0.595.8% (95.0 to 96.6%)Front view, stricterreading (0.50 to 0.95)81.2% (80.4 to 81.9%)050100
Show the numbers
ResultMetricScore95% intervalTest setKind
Side view, IoU 0.5percent97.997.3 to 98.4All 2,000 BUU-LSPINE patients, out of fold; 2,000 patientsdevelopment
Side view, stricter reading (0.50 to 0.95)percent89.488.8 to 89.9All 2,000 BUU-LSPINE patients, out of fold; 2,000 patientsdevelopment
Front view, IoU 0.5percent95.895.0 to 96.6All 2,000 BUU-LSPINE patients, out of fold; 2,000 patientsdevelopment
Front view, stricter reading (0.50 to 0.95)percent81.280.4 to 81.9All 2,000 BUU-LSPINE patients, out of fold; 2,000 patientsdevelopment

Test set: All 2,000 BUU-LSPINE patients, out of fold.

Where it falls short

  • On the second set it places 64.3% of corners within 5 pixels; the pictures of the Burapha patients score 71.6% and those of the Honduran patients score 18.3%.
  • Everything it learned from comes from one dataset collected in Thailand.
  • The average miss is bigger than the typical miss: on the side view 11.78 pixels against a median patient of 7.10.

Trained and tested on

Weights

Not released yet.

Read the model card on GitHub