New Study: BoneXpert bone age has a precision of 0.08 years

A new longitudinal study from Philadelphia finds that BoneXpert’s automated bone age has a precision (repeatability) of 0.08 years – far better than manual rating. The authors examine what this level of precision could mean when a child’s bone age is followed over time.

The first question you ask of any automated bone age method is: how accurate is it? For BoneXpert, the answer is that the deviation from the average rating of six radiologists is 0.45 years (the root mean square deviation). Accuracy is what matters the first time a child is assessed – to establish whether bone age is delayed or advanced, and to predict adult height.

But there is a second property of bone age method, at least as important and often overlooked: how precise is it? That is, if a hand is X-rayed several times, what is the standard deviation (SD) of the automated ratings? Precision is what matters once a child is followed longitudinally, because then the quantity of interest is the change in bone age from one visit to the next.

As an example of where precision matters, consider a girl with precocious puberty, who is on a course of bone age advancing at, say, 1.5 years per year. Hormone therapy is often initiated, to slow down this progression, and hence one wants to know whether that rate has changed. Detecting such a change depends entirely on how precisely each reading can be made – so here it is precision, not accuracy, that sets the limit.

Manual rating has notoriously poor precision. When 12 experts rated the same images, the SD was 0.63 years, so the SD of a bone age increment (the difference of two readings) is √2 × 0.63 = 0.89 years. At that level, manual reading cannot reliably resolve year-to-year changes.

The new study – a placebo-controlled trial following 90 children with repeated hand X-rays – reports that automated bone age is markedly more precise than manual reading: BoneXpert’s bone age precision is 0.08 years. The SD of an increment is therefore √2 × 0.08 = 0.113 years, and the smallest change that can be detected with confidence is 0.23 years. At this level, an increment of 0.5 years is clearly distinguishable from one of 1.5 years. The authors discuss what this could mean for following children and adolescents over the course of growth-modulating treatment and stress the importance of using the same X-ray equipment for the baseline and the follow-up images.

Where does the high precision come from? Two things:

A) BoneXpert assesses each of the 21 tubular bones (radius, ulna, metacarpals, and phalanges) independently and averages the 21 results, thereby reducing the noise.

B) The averaging is standardised, with equal 1/21 weight on each bone. Methods, that interpret the hand as a whole, are prone to changing the relative weights of the bones between visits, which harms precision, because bone age varies slightly across the bones.

Thus, BoneXpert is designed to yield high precision, while modern automated bone age methods based on deep learning focusses solely on accuracy and never report precision, but there is one study that allows an estimate: It found an SD between and left- and right-hand bone age of 0.29 years, indicating a precision of 0.20 years – the analogous left-right analysis using BoneXpert yields 0.11 years, indicating that BoneXpert is about twice as precise as the deep learning method.

Is your radiology AI sycophantic?

The word sycophant is Greek for “showing figs”—a gesture of calculated pleasing. In the world of AI, a sycophantic model is one that is frightened of silence. It forces an answer for every image, even when the data is inadequate, just to “please” the radiologist.

But in medicine, a guess is dangerous.

Large Language Models like Gemini 3.0 are finally learning the power of saying “I don’t know.” It’s time Radiology AI adopts the same epistemic humility.

Read more about why silence is a safety feature.

Read more “Is your radiology AI sycophantic?”

BoneXpert validation study in Turkey

Researchers from Koc University Hospital in Istanbul have published a study validating two bone age systems: BoneXpert and Vuno.

292 images were rated with BoneXpert and Vuno Med bone age and compared to a reference formed by two manual ratings.

The accuracy of the two systems was found to be the same, as illustrated in the below Bland Altman plot for girls

Accuracy is one aspect of an automated method – the ability to explain the result is another, and here the two systems are very different, as the authors illustrated by juxtaposing the outputs from BoneXpert and Vuno:

BoneXpert’s main bone age result, 7.54 y, is derived as an average over the 21 tubular bones with equal weights on the bones. In addition, BoneXpert reports a carpal bone age – the average over the 7 carpals.

These results are “explained” by showing the contour of each bone as well as its bone age. Occasionally, a bone is left out due to abnormal shape, but not in this example.

Vuno’s “explanation” is a heat map, depicting where the deep neural net output is most sensitive to changes in the image. In this example, it showed a high intensity in radius, ulna, PP2 and DP3. Since the most reliable bone age is derived by averaging over as many bones as possible, it is of concern that the method collected information from mainly these four bones. BoneXpert’s output suggested that PP2 has bone age 9.0 y, much larger than the average 7.5 y. BoneXpert assigns the same weight to all 21 bones, which tends to average out differences between the bones and this provides precision and standardisation. In a follow-up exam, the Vuno method could – at its own discretion – decide to focus on a different small subset of bones, and this would raise the question whether a change in bone age was due a biological change, or merely a result of emphasising a different subset of bones.

The full article pdf is freely available here

Bone age validation study from Leipzig

Radiologists at the University Hospital in Leipzig have published a comparison of three automated bone age methods: BoneXpert, BoneView form Gleamer and PANDA from Image BiopsyLab

The study collected images of 306 children covering the age range 1–18 years. This wide range allowed the study to reveal how the methods behaved at low and high bone ages.

A reference bone age (denoted  “Ground Truth”) was formed as the average of three human ratings.

The overall agreement between each AI and the reference was almost the same for the three methods, although BoneXpert still showed the best correlation

BoneXpert is the only method that covers the full bone age range 0-19 yr, and it is interesting to see how the other methods handled the low- and high-bone age ends.

BoneView rejected images if the chronological age of the DICOM file was below 3 y and also rejected images with a bone age 17 and above.

PANDA accepted all images, but gave large errors below 5 years as is clearly visible. Also, for females with reference bone age above 15 y, PANDA tended to give predicted bone age not much larger than 15 y – a kind of saturation effect.

The corresponding author is Dr Daniel Gräfe and the full text is available online

Google AI talks about BoneXpert AI

We might be crossing a line these days, as AI is not only taking over some of the radiologists’ work, but also taking over talking about it.

Listen to a podcast, which was generated autonomously by Google NotebookLM, from the pdf of the latest paper on the BoneXpert method for autonomous bone age determination.

  • The paper: Link
  • The podcast:
    0:00
    1:03

The podcast is surprisingly good at summarising the general aspects, but for an exact account of the accuracy, please consult the paper.

Swiss study compares two bone age algorithms

The accuracy of bone age determination by BoneXpert and Panda in 188 images was reported at ECR.

Federica Zanca from Leuven, together with four co-authors from Switzerland, presented a study comparing two bone age algorithms, BoneXpert from Visiana and Panda (based on deep learning) from ImageBiopsyLab. The study included 188 images taken under real-world conditions across 11 centres in Switzerland. The ground truth was provided by an exceptionally reliable manual rater. BoneXpert’s intended use includes autonomous use, while Panda’s is intended to only assist the radiologist. However, in this study, both algorithms were used without human interference.

The mean absolute deviation between the algorithm and the ground truth was 0.36 y for BoneXpert and 0.42 y for Panda, and this difference was significant with p=0.01.

In the Bland Altman plots below one can clearly see that the agreement is better with BoneXpert.

Figure 1: BoneXpert versus radiologist

 

Figure 2: Panda versus radiologist

The plots reveal that there are markedly fewer large deviations with BoneXpert.

So what is the clinical signficance of the different performance? The authors adressed this by defining the clinically acceptable limit of agreement to be ±1 year, and they found twice as many such significant deviations for Panda.

The table below summarises all the findings.

The poster is available though myESR (requiring log in)

Table: The deviation between the bone age algorithm and the radiologist

BoneXpert Panda
Mean Absolute Deviation 0.36 y 0.42 y
Root Mean Square Deviation 0.47 y 0.55 y
Number of deviations > 1 y 7 14
Largest deviation 1.2 y 1.9 y