A dog can be trained to pick which of seven urine samples in a line-up came from a person with bladder cancer, and to do it more often than chance allows. That was the result of the proof of principle study by Carolyn Willis and colleagues at Amersham Hospital, published in the BMJ in 2004: six dogs, a mean success rate of 41 per cent against the 14 per cent expected from guessing, with a confidence interval running from 23 to 58 per cent. Two decades of follow-up work has produced accuracy figures that in several studies sit alongside the tests currently used in clinics. It has not produced an account of what the dogs are responding to.

What exists is not one paper, and not a consensus either. It runs from single-dog pilots with a few dozen samples to a prostate study of more than nine hundred participants, and the reported accuracy ranges from barely better than chance to almost perfect. Both ends of that range are real.

What the accuracy numbers cover

Prostate and colorectal research has produced the strongest results. Jean-Nicolas Cornu’s group at Tenon Hospital in Paris trained a single Belgian Malinois over twenty-four months, then tested it double-blind on urine from 33 biopsy-confirmed prostate cancer patients and 33 men whose biopsies came back clear. The dog found the cancer sample in 30 of 33 runs, giving sensitivity and specificity of 91 per cent in European Urology in 2011. One of the three men it called wrongly was re-biopsied and found to have cancer after all.

A larger Italian study led by Gianluigi Taverna used two German Shepherds from the Italian Ministry of Defence veterinary centre, testing 362 prostate cancer patients against 540 controls. Writing in the Journal of Urology in 2015, they reported sensitivity of 100 per cent and specificity of 98.7 for the first dog, 98.6 and 97.6 for the second. On the colorectal side, a Kyushu University team led by Hideto Sonoda worked with one eight-year-old black Labrador and reported in Gut in 2011 a breath-sample sensitivity of 0.91 and specificity of 0.99 measured against colonoscopy, with watery stool samples performing slightly better again. Their dog held up on early-stage disease and was not thrown off by smoking or benign bowel conditions.

Set against those, the lung cancer study by Rainer Ehmann and colleagues in Stuttgart, published in the European Respiratory Journal in 2012 under the title “Canine scent detection in the diagnosis of lung cancer: revisiting a puzzling phenomenon”, used four dogs and 220 volunteers including a group with chronic obstructive pulmonary disease. Sensitivity landed at 71 per cent, specificity at 93.

Same species, same broad method, materially different result.

Why the compounds have not been pinned down

The assumption running through most of these papers is that a tumour and the tissue around it release volatile organic compounds, that some of those reach breath or urine, and that a trained nose is picking up a subset of them. Thorsten Walles, the senior author on the Stuttgart paper, put the frustration plainly in the European Respiratory Society’s announcement of the work: It is unfortunate that dogs cannot communicate the biochemistry of the scent of cancer.

Fifteen years on, the chemistry has not converged.

The clearest window into why comes from a 2021 pilot in PLOS ONE led by Claire Guest of Medical Detection Dogs, with collaborators at Johns Hopkins, MIT and the University of Texas at El Paso. The team ran the same urine samples three ways: past two trained dogs, through gas chromatography-mass spectrometry, and through 16S sequencing of the urinary microbiota. The dogs, a Labrador called Florin and a wire-haired Hungarian Vizsla called Midas, each correctly indicated five of seven Gleason 9 samples, giving 71.4 per cent sensitivity with specificity between 70 and 76 per cent. Gas chromatography turned up more than 1,157 distinct compounds in the sample headspace. The handful that differed significantly between cancer and control were not the ones earlier urinary volatilome studies had flagged, and one of them was trimethyl silanol, a breakdown product of silicones rather than anything a prostate makes.

The case for a mixture rather than a marker

The Guest paper’s authors argue something more interesting than a failed search. Their position is that work across prostate, colorectal, liver and lung cancer has repeatedly found selection rules that separate positives from negatives inside a given dataset, without any of those rule sets generalising into a template that holds elsewhere. On their reading, what a dog responds to may be scent character rather than molecular composition, an emergent property of a mixture in which no single compound does the work.

This matters commercially as much as scientifically. A dog is not a scalable diagnostic, and every group in the field is ultimately aiming at a device. A device needs a target, and if the target is a pattern in perceptual space rather than a named compound at a measurable concentration, the engineering problem changes shape. Several groups have since shifted to training machine-learning models on the dogs’ calls instead of on chemical identifications.

What happens when the samples are unfamiliar

Kevin Elliker and colleagues at Cambridge and Nottingham set out to train ten dogs on prostate cancer urine. Three reached the second stage of training, and two of those learnt to pick the cancer sample out of a four-position array often enough to look convincing, one of them at around 76 per cent against a chance rate of 25. Then came three double-blind tests using donors the dogs had never met. The nine-year-old Labrador scored 2 of 15 and 2 of 16 arrays, a sensitivity of 0.13. The Border Collie managed 4 of 16, a sensitivity of 0.25. Neither beat chance, and given samples from the same donors the two showed no agreement with each other. Their conclusion, published in BMC Urology in 2014, was that the dogs had memorised the odour of individual training donors rather than generalised to anything shared by the disease.

The Guest pilot turned up a smaller version of the same difficulty. Both dogs independently flagged the same two cancer-free controls, in the same order, despite blinding and randomised positions. Sequencing later showed both had unusual bacterial profiles. Whether the urinary microbiota was contributing to the odour is unresolved, and the authors list it for larger studies.

Why these results are not a reason to skip standard screening

Nothing in this body of work supports substituting a dog, or a dog-derived test, for colonoscopy, mammography, low-dose CT or a PSA discussion with a doctor. Most of the studies above are case-control designs on selected patients, which is a different task from screening a general population. The largest recent trial, the 2024 Scientific Reports paper on the commercial SpotitEarly breath platform, tested 1,386 participants of whom 338 had cancer. Of those, 261 had one of the four cancer types the system was trained to detect, and for that subset the paper reports 93.9 per cent sensitivity and 94.3 per cent specificity. Two of the authors are SpotitEarly employees and two more were funded by the company, as the declarations state.

Run that same performance at a population prevalence of 1 per cent and, by our arithmetic, about one in seven positive results would be a true positive. The paper does not claim otherwise. Prevalence is what separates a case-control accuracy figure from a working screening test, and it is the first thing to check when a number like that reaches a headline.

What would change the picture is a compound set that can be named, measured and reproduced across laboratories with different sample handling and different patient groups. Sonoda’s team said as much in 2011, and it remains the stated goal of most of the groups working on breath volatilomics. Until someone gets there, each new accuracy figure adds to a pattern the field can demonstrate but cannot yet explain.