Baby cry analyzer accuracy: what we measured, what we found
We answer the baby cry analyzer accuracy question with a measurement, not a percentage. The short answer: a cry carries who the baby is and how old they are; it does not carry why they are crying. This page explains what we measured, which numbers we found, where the high-percentage claims come from, and what Agubu does as a result.
Every figure here comes from our own measurement or from the named publication. The same page sits inside the app under Settings → How does this app decide? and will never be locked.
What we measured
We worked on donateacry-corpus, the only public, commercially usable dataset labelled by cause: 457 recordings from 221 babies. From each recording we extracted YAMNet audio embeddings and asked the same classifier three separate questions: which baby is this clip from, how old is the baby, and why is the baby crying.
We split the data by baby: one baby's recordings sit either in training or in testing, never both. That detail decides everything; with the same baby on both sides, the model learns the baby rather than the question, and the score inflates.
Our metric is balanced accuracy, not raw accuracy. The reason: 83.6% of the recordings carry the label “hungry”, so a model that learns nothing and always answers “hungry” already scores 83.6% on raw accuracy.
The results table
Same data, same embeddings, same classifier, same baby-wise cross-validation. Chance level in brackets:
- Identity — “which baby is this?” → 62.5% (chance 9.1%). Gain +53.4 points.
- Age — newborn / infant / older → 40.2% (chance 33.3%). Gain +6.9 points.
- Cause — three classes → 32.9% (chance 33.3%). Gain −0.4 points.
- Cause — hungry or not → 49.2% (chance 50.0%). Gain −0.8 points.
- Cause — pain or not → 51.0% (chance 50.0%). Gain +1.0 points.
How to read it
The pipeline works: it tells individual babies apart from their cries far above chance, and it can estimate age. The same pipeline, on the same data, finds nothing about the cause. The method is not the problem. The information is not in the sound.
This is not just our finding
A 2023 study by Lockhart-Bouron and colleagues in Communications Psychology examined 39,201 cry sequences recorded in the homes of 24 babies. Across the 676 clips with a known cause, neither adult listeners nor an algorithm trained on the parental action that stopped the cry could reliably classify the cause, while age and identity came through reliably. Source: Communications Psychology, 2023.
Our measurement is an independent replication of that result on our own data: a cry carries identity and age, but not cause.
Where “96% accuracy” claims come from
High-percentage claims are common in this category; we measured the mechanism too. Of the 221 babies in the dataset, 197 (89%) contributed only a single label, so recognising the baby amounts to knowing the label. Split the recordings randomly instead of by baby and the same baby appears in both training and testing, and the score rises on its own: on identical data the three-class task went from 32.9% to 35.5%, “hungry or not” from 49.1% to 59.3%, “pain or not” from 50.9% to 62.7%. That is 10 to 12 points on a simple model, and far more on a large one.
The second factor is raw accuracy: with 83.6% of recordings labelled “hungry”, even a model that always says “hungry” scores high. Combine a random split with raw accuracy and you get an impressive figure — one that measures nothing.
We tested a model on the market
We ran the model shipped inside one app in this category on the same data. The result fell far below the advertised figure, even though the test favoured the model, since it was probably trained on that very data. We do not name the app here; the point is to show a method, not a product.
Its behaviour matters more than its score. Given absolute silence, it named a cause with 100% confidence — as it did for all six synthetic inputs we tried. Across different seconds of one recording it changed its answer on more than half the clips, averaging 96.9% confidence while doing so. That is the numerical counterpart of the store reviews saying “it gives a result even in a silent room” and “it gave four different answers to the same cry”.
Cry detection works; cause detection does not
Two separate questions, and one genuinely has an answer. “Is this a cry?” Answerable: in our calibration real cries scored a median of 0.737, the hardest synthetic input 0.0095 — a gap of roughly 78 times. Our threshold is 0.02; it passes about 90% of real cries and blocks every synthetic input. “Why is this baby crying?” Not answerable; that is the table above.
So what Agubu does
The result is not an answer but a ranked set of possibilities. Hiding uncertainty does not remove it; it only hands you misplaced confidence. The full feature: baby cry analyzer; the shorter telling: Why a cry cannot be decoded.
- Detection: the microphone listens for 8 seconds; an on-device model confirms the sound is really a cry.
- Quality gate: silent, noisy or short recordings are rejected. No reason is invented for silence.
- Ranking from context: time since the last feed, time awake, the last diaper change, age and the hour of the day rank the likely reasons with probabilities; once enough examples exist, your feedback gains weight.
- Saying when it isn't sure: if the possibilities are close, the result says “I'm not sure” and lists the top three.
Sources
- Our own measurements — donateacry-corpus (457 recordings, 221 babies), YAMNet embeddings, baby-wise cross-validation; reproducible from the “How does it work?” page inside the app
- Lockhart-Bouron et al., Infant cries convey both stable and dynamic information about age and identity, Communications Psychology, 2023 — https://www.nature.com/articles/s44271-023-00022-z
- Our threshold calibration for AudioSet-based cry detection

Frequently asked questions
What is Agubu's accuracy rate?
We do not give a percentage, because the only honest percentage for predicting the cause from sound is chance level, and we show that in the table above. What we can say and have measured: cry detection passes about 90% of real cries and blocks every synthetic input; the ranking of reasons comes from context, not from sound.
Other apps claim “96% accuracy”. Are they lying?
Usually not lying — mismeasuring. The recordings are split randomly rather than by baby, and raw accuracy is used. On the same data those two choices raise the score on their own. Split by baby and look at balanced accuracy, and the figure drops to chance.
Can I reproduce this measurement myself?
Yes. The dataset is public (donateacry-corpus), the embeddings come from YAMNet, and the split is by baby. Anyone following the same steps reaches the same table; the method is also written up on the “How does it work?” page inside the app.
Try Agubu
Free on iPhone and Android. No account, no ads; your entries stay on your phone.
