How desensitising toothpastes are tested: Schiff air score, tactile probe and VAS explained.
Desensitising toothpastes are tested by provoking a sensitive tooth on purpose and writing down what happens, three ways: an examiner directs a short blast of air at the exposed neck of the tooth and grades the reaction on the Schiff scale, a calibrated probe presses on the same spot with rising force until the person flinches, and the person marks their own pain on a line, the visual analogue scale. Every number you will ever read about a sensitivity toothpaste comes out of one of those three, and a scoping review of 71 studies published in 2026 found the visual analogue and Schiff scales to be the field's two commonest measures, with the reasons given for choosing them inconsistent from paper to paper1. This page reads the rulers before it reads any measurement, including ours: the human evidence behind S3 Sensitivity Science™ is an independent third-party consumer trial of 51 adults with sensitive teeth over eight weeks (ADSL, Devon, to Good Clinical Research Practice), in which panellists were asked how their teeth felt and none of the three instruments below was used.
What was checked28 peer-reviewed studies, the CAP Code on advertising substantiation, the Oral Health Foundation, S3 consumer trial (ADSL, 2026)
- The instruments and the patients often disagree: a 2006 Cochrane review of six randomised trials found potassium nitrate toothpastes improved air-blast and tactile scores at six to eight weeks and found no significant effect on the patients' own assessment2.
- A trial without a control arm shows nothing, because the control arm moves: in a six-week double-blind randomised trial of 120 adults, a plain fluoride toothpaste improved cold-air scores as much as two desensitising toothpastes did3.
- What a fair trial looks like was agreed nearly thirty years ago: parallel groups, blinded on both sides, allocated at random, with touch and cold and evaporative stimuli, a negative control and a benchmark, eight weeks of use, follow-up to see whether the change lasts, and two independent trials before a product is approved4.
- Most of this evidence is paid for by the people selling the paste: a 2026 network meta-analysis of 93 randomised trials in 9,548 participants reported that 96% of the stannous fluoride trials, 86% of the arginine trials, 76% of the potassium trials and 33% of the nano-hydroxyapatite trials were industry funded5.
- S3's declared levels are the levels the published trials used: potassium nitrate at 5%, and nano-hydroxyapatite supplied as a 10% solution.
What does the Schiff air score measure?
The examiner's reading of how a tooth reacts to a blast of air. The stimulus is standardised as far as it can be: in one validation study the air was applied to the cervical area of the tooth for one second from a distance of one centimetre6. The air is an evaporative trigger, drying and cooling exposed dentine at the same time, which is why it stands in for the cold drink and the walk to the bus stop. The important structural fact about the scale is who holds the pen. On the Schiff scale the examiner watches the person and assigns the score; on the visual analogue, numeric, verbal and faces scales the person marks their own7.
The numbers on the scale are best understood from how trials use them. A tooth is admitted to a trial when it scores in the reactive band — two or three in a three-day double-blind randomised trial of 120 adults, above two in a UK examiner-blind randomised trial over fourteen days89 — and a tooth that has stopped reacting scores zero, which is how an industry-funded double-blind randomised trial of 120 adults funded by the maker of one of the pastes defined complete relief10.
Two things follow from that, and both matter more than the scale itself. First, a trial that enrols only the teeth scoring worst has built regression to the mean into its design; some of them would have scored better at the next visit whatever was brushed on them. Second, the scale is coarse. The entry thresholds above sit at two and three and a tooth that no longer reacts scores zero, so the whole instrument has four positions on it, on the reading of the two trials above, one of them industry-funded810. What a published improvement usually amounts to, once it is averaged across a group, is a fall of half a step or one step.
One thing this page could not do is card the scale's own definition. A PubMed search for the scale by name in September 2026 returned ten records, every one of them a study that uses the scale and none that defines it; the nearest thing to a source is a 1994 twelve-week trial of a 5% potassium nitrate toothpaste in 67 adults whose first author gives the scale its name, and that paper's abstract — the only version obtainable here — reports an air-blast measurement without printing the wording of the scale's points11. So the anchors above are given as trials apply them, not quoted from an original this page has read. The one-screen definition sits on the glossary page for the Schiff sensitivity scale.
| Instrument | What it applies to the tooth | Who produces the score | The scale | What is known about how well it measures | What it cannot capture |
|---|---|---|---|---|---|
| Schiff cold-air score | One second of air from a dental syringe, about a centimetre away | The examiner, watching the person | Four steps, from no reaction to a reaction the person wants stopped | Highest specificity of five scales tested, 91%; lowest accuracy of the five, 0.729 (cross-sectional clinical study, 72 adults, tactile and ice-stick stimuli, not air)7 | Anything between the steps; anything the person feels outside the surgery |
| Yeaple tactile probe | A calibrated point pressed on the exposed dentine, force raised step by step | The examiner reads the force; the person says when it hurts | Grams of force tolerated; trials enrol between ten and 50 grams | Not measured directly; used as the benchmark against which a newer probe was validated in a clinical study whose funding is not stated12 | Cold, air, sweetness — every trigger that is not touch |
| Visual analogue scale | Whichever stimulus the trial chose | The person, marking a line | A distance along a 100 mm line | Accuracy 0.729 to 0.750 across five scales; the visual analogue scale was one of the five judged accurate for diagnosis7 | Which stimulus produced it, unless the trial says |
| Numeric, verbal and faces scales | Whichever stimulus the trial chose | The person | A number, a word or a face | Highest sensitivity of the five, 81.9%, on the numeric scale in that same clinical study, which 47.2% of participants also preferred7 | The same as the visual analogue scale |
| DHEQ-15 questionnaire | Nothing; it asks about the last month of the person's life | The person | Fifteen items, summed | Internal consistency 0.970 and test-retest reliability 0.920 in a clinical study of 300 adults6 | Anything measured in a chair, and anything that changed today |
| Electrical stimulation | A rising electrical current | The instrument | A threshold current | Not measured; used as an endpoint alongside a verbal rating in an industry-funded randomised trial13 | Whether the current is provoking the same nerve the cold drink does |
| Cold-air quantitative sensory testing | Cooled air at a controlled temperature | The person reports; the device records the temperature | The temperature that evokes moderate to strong pain | Stable within a person over three weeks in a non-randomised clinical trial (intra-class correlation 0.83, 29 adults); varies enormously between people14 | Nothing about a group; the between-person spread is the point |
| Reproducibility of all of the above | Standardised air and cold fluid stimuli | The person, twice | Agreement between the two readings | Limited within the same subject even with the stimulus controlled, in a two-part clinical study of 63 and 42 adults15 | — |
What does the tactile probe measure?
The lightest touch a tooth will put up with. The Yeaple probe presses a fine point against the exposed dentine at a set force in grams; the force is raised in steps until the person reports discomfort, and the force at which they do is the threshold. A less sensitive tooth tolerates more grams, so this is the one instrument on which improvement means the number going up.
Trials use the probe twice, to pick their teeth and to measure them: the three-day double-blind randomised trial of 120 adults above, whose funding is not stated in the record, enrolled teeth responding between ten and 50 grams, which excludes both the tooth that flinches at almost nothing and the tooth that will take a firm press8. An industry-funded double-blind randomised trial of 120 adults set the window tighter still, at ten to 20 grams10. A non-randomised clinical study that stimulated 40 adults per group used 20 grams and below16. The probe and the air score are so well established as the pair that measures this condition that when a microprocessor-controlled tactile instrument was evaluated in two eight-week parallel-design clinical studies of 100 adults each, it was validated against them rather than the other way round; funding is not stated in that record12.
What the probe does not do is move in step with the air score. In an eight-week randomised examiner-blind trial of 133 adults, funded by the manufacturer and written by its employees, an experimental potassium chloride toothpaste beat a plain fluoride toothpaste on the Schiff score at two, four and eight weeks — and the fluoride toothpaste beat it on the tactile threshold at eight weeks17. Same teeth, same visit, opposite verdicts. A page that quotes one of those numbers and not the other has not made a mistake; it has made a choice.
What does the patient's own rating add?
The only opinion in the room that belongs to the person with the tooth. The visual analogue scale is a line on which the person marks how much the stimulus hurt; the numeric scale asks for a number, the verbal scale for a word, the faces scale for a face.
There is exactly one published study of how well any of these scales tells a sensitive tooth from a healthy one, and it is small. In a cross-sectional accuracy study of 72 adults, hypersensitive and non-sensitive teeth in the same mouths were given tactile and ice-stick stimuli and scored on all five scales; area under the receiver operating characteristic curve ran from 0.729 for the Schiff scale to 0.750 for the numeric scale, the Schiff scale had the highest specificity at 91%, the numeric scale the highest sensitivity at 81.9% — sensitivity here meaning the share of genuinely sensitive teeth a scale catches, not the tooth's own — and 47.2% of participants said they preferred the numeric scale7. The authors found all five accurate enough to diagnose the condition and named the Schiff scale as their preferred assessment scale, so this is not a demolition. It is a ceiling. An area under the curve of about 0.75 is a useful test, not a precise one, and in that clinical study it was measured with a touch and an ice stick rather than the air blast the Schiff scale is normally used with7.
Accuracy is one question; repeatability is the next one, and the answer is worse. A clinical study in two parts, of 63 and 42 adults, published in 2001 with the air and cold fluid stimuli standardised, found subject-based reproducibility limited even so, and its authors concluded that this may be part of why the efficacy of desensitising agents is so hard to establish15. No later paper in this page's searches has contradicted it.
One measure in this area has had the validation work the pain scales have not: the questionnaire. The 15-item Dentine Hypersensitivity Experience Questionnaire, which asks about the last month rather than the last second, showed internal consistency of 0.970, test-retest reliability of 0.920 and a clean separation between people with the condition and people without it, in a validation study of 300 adults6. It is also the measure that most often refuses to move when the instruments do.
Do the three measures agree?
Often not, and the pattern is consistent enough to be worth stating plainly: the instruments move first and further, the patients move later and less.
| Study | What the examiner's instrument found | What the patients said | The gap, in the paper's own numbers |
|---|---|---|---|
| Cochrane review of six randomised trials, 2006, funding none stated2 | Air-blast sensitivity significantly improved at six to eight weeks, standardised mean difference −1.25; tactile likewise | Subjective assessment: no significant effect | Significant on two instruments, absent on the patients; the reviewers concluded there was no clear evidence to support potassium toothpastes |
| Earlier Cochrane review, eight randomised trials, 2001, funding none stated18 | Air-blast standardised mean difference −1.51 in the four poolable trials | Subjective assessment at six to eight weeks not significant | The same split, five years earlier, in a different set of trials |
| 133 adults, eight weeks, examiner-blind randomised, industry-funded, authors employed by the funder17 | Schiff score favoured the test paste at two, four and eight weeks | Visual analogue scale: no difference between groups at any point | And the tactile threshold favoured the control paste at eight weeks |
| 96 adults, four weeks, double-blind randomised, industry-funded, five of eight authors employees of the funder13 | Electrical stimulation reading climbed to 111.6% over the control at four weeks | Verbal evaluation scale plateaued around 37% | The instrument kept rising for three more weeks after the people stopped noticing much change |
| 164 adults, eight weeks, six-arm randomised, no competing interests declared19 | Every active arm improved on the Schiff scale and the visual analogue scale at four and eight weeks; so did the plain fluoride control, by week eight | The quality-of-life questionnaire improved in the active arms and not in the fluoride control | The control passes two instruments out of three |
| 215 adults, eight weeks, examiner-blind randomised, funded by the maker of the test paste20 | Schiff and tactile both better than the reference paste at every point from day three | Quality-of-life questionnaire improved, but the difference between treatments was not significant | A clean instrument result and an inconclusive patient result in the same trial |
| Stimulus against stimulus rather than instrument against patient: meta-analysis of six four-week randomised trials21 | Nano-hydroxyapatite better than comparators on evaporative and tactile stimuli | — | And no difference at all on cold stimuli: the choice of trigger decided the answer |
That first row is read line by line on the page about what the Cochrane review found and what it did not. Read the last row twice. It is not that the patients were wrong and the probe was right, or the other way round; it is that "did it work?" is not a single question. It decomposes into which stimulus, which scale, against what comparator, and did the person agree with the instrument — and a claim that does not say which of those it means has not said anything. Where those four questions land for the two big families of active is the subject of the page on tubule occlusion versus nerve desensitisation.
Has anyone tried to build a better ruler?
Twice that this page could find, and neither attempt displaced what it improved on.
In 2013, a group in London took the Schiff Index — which scores one tooth surface — and built a Cumulative Hypersensitivity Index to score a whole person, validating it in a cross-sectional clinical study of 350 adults recruited from hospital and general practice in south-east England, against two other ways of summarising the same scores, with correlations of 0.982 and 0.96322. It is a sensible piece of work. Thirteen years later the scoping review of 71 studies lists the visual analogue and Schiff scales as the field's measures and does not mention it1.
The second attempt was an instrument rather than an index: a pre-calibrated, microprocessor-controlled tactile probe, tested for repeatability by two examiners in twelve adults and then run alongside the Yeaple probe, the air blast and the visual analogue scale in two eight-week clinical studies of 100 adults each, where its readings correlated significantly with all of them; funding is not stated in the record12. Correlating with the existing instruments is what a new instrument has to do to be believed, and it is also, precisely, what stops it replacing them.
There is one measurement in this area that behaves the way a physical measurement ought to. Cold air delivered at a controlled temperature gives a threshold — the temperature at which a tooth starts to hurt properly — and in 29 adults monitored over three weeks that threshold was stable within a person, with an intra-class correlation of 0.83, while varying enormously between people14. The instrument is a purpose-built research device, not the syringe on a dental unit, and no toothpaste study in this page's sources has used it.
What makes a sensitivity trial fair?
There is a published answer, it is nearly thirty years old, and it is short enough to use as a checklist. The 1997 consensus guidelines for trials in this condition recommend: a parallel-group design, with both the patient and the examiner blinded and the allocation made at random; tactile, cold and evaporative stimuli, not one of them alone; a negative control and a benchmark that is known to work; a duration of eight weeks for most trials; follow-up after treatment to see whether the change persists; results expressed as clinically significant change rather than statistical significance alone; and at least two independent trials before a product is approved4. Read a claim against those seven points and most of what is printed on a box answers two or three of them.
The field has not settled the methods since. A 2015 systematic review of 105 randomised trials of eleven agents could not run a meta-analysis at all, because the trials were too heterogeneous in design, stimulus and comparator to pool23. A 2023 systematic review and network meta-analysis of 32 randomised trials in 4,638 participants, written by manufacturer-affiliated authors who declare advisor fees, lecturer fees and research grants from dentifrice makers, was still calling for standardised methodology guidelines for these trials24. Twenty-six years after the guidelines, the people writing the reviews were still asking for the guidelines.
The most striking gap was closed only in 2024. Every trial provokes the same tooth twice in one visit, at baseline and again afterwards, and until then nobody had asked how long the tooth needs in between16. When somebody finally did, in a non-randomised clinical study of 40 adults per group, the answer was about two minutes for a cold air blast and longer for a touch stimulus, and the mean Schiff score fell rather than rose as the interval shortened; the authors recommend at least three minutes in future studies16. Everyone before them chose that interval by habit.
How much of the improvement is the toothpaste?
Less than a before-and-after number implies, and this is the part of the page that applies to everybody's numbers, ours included.
Start with the cleanest demonstration. In a split-mouth, subject-blind randomised study of 22 adults, a periodontal dressing was placed over the sensitive teeth in one quadrant and beside — not over — the sensitive tooth in the other25. The dressing contained nothing active. It was found to cut pain by 95% to a thermal stimulus and by 85% to an evaporative one, and the authors concluded that the perception of pain in this condition is altered by sensory factors, so a trial response contains a stimulus response as well as a treatment response25.
Then the placebo that lasts. In a six-month single-blind randomised crossover trial of 35 patients, water used as a placebo cut air-blast pain by 20% immediately and by 36% at six months; the active paste beat it at every recall, but the water alone accounted for more than half of what the active achieved by the end26. In the six-week double-blind randomised trial of 120 adults above, all three pastes improved and no product differed significantly from another at any time point3. And a 2019 network meta-analysis of 30 randomised trials found a significant placebo effect across the network, alongside no significant difference between fluoride and placebo, or between potassium and placebo, in that pooled evidence27.
Add the regression to the mean built into the entry criteria, and the arithmetic becomes clear: a group's improvement is the active, plus the placebo, plus the sensory effect of being treated, plus the statistical pull of having enrolled the teeth that scored highest at screening. Only the difference from a proper control arm belongs to the toothpaste. The page on why relief that builds over weeks is a feature takes the time course of that difference further.
Who paid for the trial, and does it matter?
It matters, it is knowable, and the convention this page asks you to apply is the one it applies to its own citations: name the funder in the same sentence as the finding.
The scale of it is documented. The 2026 network meta-analysis of 93 randomised trials in 9,548 participants reported that 96% of the stannous fluoride trials, 86% of the arginine trials, 76% of the potassium trials and 33% of the nano-hydroxyapatite trials were industry funded, and concluded that stannous fluoride and arginine should be first-line self-care options with the choice guided by preference, tolerability and availability rather than an expectation of superior efficacy5. Industry funding is not a reason to discard a trial. Two of the most careful write-ups on this page came from manufacturers. An eight-week examiner-blind randomised trial of 215 adults, funded by the maker of the test paste, put a clean instrument result beside a quality-of-life result that did not reach significance20. A double-blind randomised trial of 120 adults, funded by the maker of one of the pastes, recorded that its own sponsor's paste came out level with the comparator by week eight10. Both told you something that cost them something.
It runs through the guidance too. The Oral Health Foundation's page on sensitive teeth, read in September 2026, carries a line thanking a toothbrush manufacturer for the educational grant that funds it28. That disclosure is the reason the page can be used, not a reason to distrust it — which is the whole argument of this section in one example.
And on the advertising side there is a rule, though it is not the rule most people imagine. The CAP Code, read in September 2026, requires that before an advertisement is published the marketer holds documentary evidence to prove claims that consumers are likely to regard as objective, and separately forbids suggesting that a claim is universally accepted when a significant division of informed or scientific opinion exists29. Note what that does and does not say. It sets a standard for evidence held; it does not define any particular phrase, and it does not require the evidence to be published where you can read it.
How has S3 been tested, and how has it not?
Two ways, neither of them one of the instruments above, and both described here as what they are.
The human evidence is a consumer trial: an independent third-party consumer trial of 51 adults with sensitive teeth over eight weeks, run in Devon to Good Clinical Research Practice. A consumer trial asks people how they feel about a product. There was no Schiff score, no Yeaple probe, no visual analogue scale against a controlled stimulus, and no negative-control arm of the kind the 1997 consensus guideline for clinical trials in this condition requires4. It measures perception, which is a real thing worth measuring, and it is not the same thing as a randomised controlled trial. The readout has not been published, so nobody outside the company can check the method, the questions or the results. That is a limitation, it is ours, and no figure from it appears anywhere on this page.
The laboratory evidence is separate. S3 has been tested at the University of Reading, and laboratory work of that kind is an example of a measurement, not a result about a person; what a microscope can and cannot establish is the subject of the companion page on which ingredients occlude dentine tubules, and how well.
That leaves the phrase itself. In S3's own claim register, "clinically proven" is permitted in exactly one place: attached to the 5% potassium nitrate dose that placebo-controlled trials have tested. It is not permitted about the finished toothpaste, and the difference is the difference this page exists to teach. An ingredient at a tested dose is a claim you can check against the literature. A finished product is a different object, and nothing about the dose carries over to it on its own.
| Kind of evidence | What it can establish | What it cannot | Where S3's own evidence sits |
|---|---|---|---|
| Systematic review or meta-analysis | The state of the evidence across trials, with a certainty rating | Anything the underlying trials did not measure | Not applicable: no review has assessed S3 |
| Randomised trial with a negative control | That an effect exceeded what the control arm did, on the instruments used | What happens beyond the trial's length, or on instruments it did not use | None: S3 has not run one |
| Single-arm clinical study | Change from baseline | That the change was caused by the product | None |
| In vitro electron microscopy | That a surface was covered, or a tubule narrowed, on a specimen | That anyone hurt less | The University of Reading laboratory work |
| In situ study | Behaviour of a material in a real mouth, on a carried specimen | A benefit to the tooth it was cut from | None |
| Consumer perception trial | What people say about a product they used | Whether a measured outcome changed, or whether the product caused what they say | The eight-week trial of 51 adults, unpublished |
Frequently asked questions
Is S3 proven in clinical trials?
No. In S3's own claim register the phrase clinically proven is attached to one thing only, the 5% potassium nitrate dose that placebo-controlled trials have tested, and never to the finished toothpaste. S3's human evidence is a consumer trial of perception — 51 adults, eight weeks, panellists asked how their teeth felt — and its laboratory evidence is bench testing at the University of Reading. Neither is a randomised controlled trial of the product against a control, and the consumer-trial readout has not been published. What is supported is the ingredient evidence: the doses in the formula match the doses used in the trials of those ingredients, and those trials are on the other pages in this section. If that answer sounds narrower than the wording on most sensitivity boxes, that is the point of the page.
Why do trials blow air at a tooth instead of giving someone a cold drink?
Because a drink cannot be standardised and a syringe can. One second of air at a fixed distance dries and cools the exposed dentine in a way the next examiner can repeat6, whereas a drink varies in temperature, volume, contact time and which teeth it touches. The trade-off is that the reader's actual complaint — the drink, the winter air, the spoonful of ice cream — is being measured by proxy, and one instrument's answer does not transfer to another's. In a meta-analysis of six four-week randomised trials, nano-hydroxyapatite came out ahead on evaporative and tactile stimuli and showed no difference at all on cold21. Same treatment, three stimuli, two answers.
If the placebo also works, does the active matter?
Yes, but less than the raw before-and-after number suggests, and only a control arm can tell you how much less. A dressing with nothing active in it was found to cut pain by 95% to a thermal stimulus, in a split-mouth randomised study of 22 adults25. Water used as a placebo cut air-blast pain by 36% at six months in a randomised crossover trial of 35 patients, while the active paste stayed ahead of it at every recall26. Those two facts sit together comfortably: the placebo response is large and real, and the difference from it is the part that belongs to the toothpaste. It is also why a bare before-and-after figure with no control arm behind it settles nothing.
Why do some trials say a toothpaste works and others say it does not?
Four reasons, in rough order of how often they apply: the trials used different stimuli, the trials used different scales, one measured the examiner's instrument and the other asked the patient, and the comparator was different — a plain fluoride paste is a much harder benchmark than a placebo, and in a six-week double-blind randomised trial of 120 adults the plain fluoride paste matched two desensitisers3. Underneath all four sits heterogeneity so severe that a systematic review of 105 randomised trials could not pool them at all23. A 2017 UK guideline review for general dental practice reaches the practical conclusion: there does not currently appear to be one ideal desensitising agent that can be recommended30.
How long should a sensitivity trial run?
Eight weeks, by the 1997 consensus guidelines, with follow-up afterwards to see whether the change persists and at least two independent trials before a product is approved4. Shorter trials are common and are not worthless, but they answer a smaller question, and the instruments do not settle at the same speed: in one four-week industry-funded double-blind randomised trial of 96 adults, the electrical-stimulation reading was still climbing at week four while the patients' verbal ratings had levelled off weeks earlier13. The same arithmetic applies if you are judging a tube at home: give it the weeks, and remember that some of what you feel in the first fortnight is not the active. The Journal's guide to how sensitive toothpastes work, and why yours might not be working for you covers what to expect over those weeks. Its round-up of the sensitivity shelf shows how the claims are usually worded.
Where S3 sits
S3 carries its two headline actives at the levels the published studies used: potassium nitrate at 5%, nano-hydroxyapatite as a 10% solution. Both of those figures are inclusion levels of the ingredient as supplied, the active hydroxyapatite content is lower, and S3 states both. The phrase clinically proven attaches to the 5% potassium nitrate dose the placebo-controlled trials tested, never to the finished toothpaste.
S3 Sensitivity Science™ pairs potassium nitrate with two hydroxyapatites and adult-strength fluoride in one daily paste.
See the toothpasteS3's formula pairs the best-established desensitiser, potassium nitrate, with two hydroxyapatites and fluoride. Three actions, one tube: it calms the nerve, strengthens the enamel surface and protects against further wear. The formula is patent-pending S3 Repair Technology™, UK application GB2604755.5. Others are dentist recommended; S3 is dentist owned, with more than 20 UK dentists having invested their own money in it. Read more about S3.