How to read a tooth sensitivity study: sample size, blinding, stimuli and what a p-value tells you.
To read a tooth sensitivity study, put eight questions to it: who was in it, what the control group got, who knew which paste was which, which stimulus and scale were used, how many finished against how many were needed, at which visits the scores were taken, what the p value and the interval say, and who paid. Put to the studies this site reads, those questions show that "works within a week" has almost no evidence behind it for any toothpaste, because few trials measured anything before week two and the field's own 1997 guideline judges a desensitising paste over eight weeks1. S3 Sensitivity Science™'s own evidence includes an independent consumer trial of 51 adults with sensitive teeth over eight weeks, and this page puts it through the same eight questions as the studies it sits beside.
What was checked24 peer-reviewed studies, S3 consumer trial (ADSL, 2026), NHS guidance
- The 1997 consensus guideline asks for parallel-group trials kept blind to participant and examiner, with a negative and a benchmark control, three kinds of stimulus and eight weeks for most trials1. A 2026 review of 58 recent randomised trials found randomisation, blinding and effect estimates reported inconsistently2.
- Only three toothpaste trials on this page measured in the first week: two were funded by Procter & Gamble and by Haleon, each the maker of a paste it tested, and the third, with no control paste free of desensitiser, found no difference between two desensitising pastes31920. The guideline, by contrast, judges a paste over eight weeks1.
- A significant p value says a difference is unlikely to be chance; it does not say the difference is large or that most people felt it, and no agreed threshold was found for how big a change in a sensitivity score a person would notice.
- Most trials of the main actives were paid for by companies that sell them: a 2026 network meta-analysis counted industry funding behind 96% of stannous fluoride studies and 76% of potassium studies4.
- S3's own consumer trial had no head-to-head arm and reports what panellists said, which is why it is described as a consumer trial and not a clinical trial.
What does the field's own rulebook ask a sensitivity study to do?
It asks for more than most papers report. The 1997 consensus guideline by Holland and colleagues (Journal of Clinical Periodontology) recommends parallel groups allocated at random and kept blind to both participant and examiner, tactile, cold and evaporative stimuli, a negative and a benchmark control, eight weeks for most trials, follow-up to see whether changes persist, and at least two independent trials before a product is approved1. Nearly three decades later, a 2026 scoping review and Delphi consensus by Pollard and colleagues (Periodontology 2000) found that definitions of the condition still varied widely across 72 papers, and that 58 recent randomised trials reported randomisation, blinding and effect estimates inconsistently against the CONSORT reporting standard2.
Two narrower reviews point the same way. In a methodological review, Matranga and colleagues (Journal of Oral Science, 2017) read 40 randomised trials from 2009 to 2014 and judged the statistical methods inappropriate in 77.1% of the 35 that used parametric tests5. Neto and colleagues (Saudi Dental Journal, 2025), in a systematic review of 16 mouthwash trials, found the randomisation sequence and blinding fully described in 37.5% of the papers, although 14 of the 16 were judged at low risk of bias6. Neither finding says the trials are worthless56. What that 2025 review shows is that every one of its 16 trials reported the result adequately, while far fewer described how the trial was run, so the reader has to go looking for the method6.
The Bristol Dental School trials group (Newcombe, Seong and West, Journal of Dentistry, 2023) has set out what it calls an efficient design and analysis, drawn from two of its own trials, and recommends it for wider use; the authors declare no competing interests for that paper, which was not funded by any specific grant, and because its full text was not available to this page, the design itself is not described here7.
Table B sets each of the eight questions beside what the guideline asks and what the later reviews found, with the page on this site that goes deeper where one exists.
| Question | What the 1997 guideline asks | What later reviews found | Where to read more |
|---|---|---|---|
| Who was in it | A clinical diagnosis, excluding people with conflicting conditions or current treatment | Definitions vary; few studies rule out every other cause of pain (Pollard 2026) | This page, next section |
| What the comparison group got | A negative control and a benchmark control | Control arms improve markedly (West 1997; Hu 2019) | Did it work or did it just feel like it? |
| Who knew which paste was which | Double-blind | Blinding fully described in 37.5% of 16 mouthwash trials (Neto 2025); reported inconsistently in 58 trials (Pollard 2026) | This page |
| Stimulus and scale | Tactile, cold and evaporative stimuli | Scale choice inconsistently referenced across 71 studies (Gupta 2026); one person's scores repeat poorly (Ide 2001) | How desensitising toothpastes are tested |
| How many finished | At least two independent trials | Several teeth per person analysed as if independent (Matranga 2017) | This page |
| When it measured | Eight weeks for most trials; follow-up for persistence | The largest network meta-analysis pooled two-week results only (Gormley 2026) | Why relief that builds over weeks is a feature, not a flaw |
| p value and interval | Clinically significant changes, no threshold given | Statistics judged inappropriate in 77.1% of 35 trials (Matranga 2017); no agreed minimal important difference found | This page |
| Who paid | Not addressed in the abstract | Industry funding behind 96% of stannous fluoride and 33% of nano-hydroxyapatite studies (Gormley 2026) | Checking a brand's evidence |
Who was in the study, and how was their sensitivity confirmed?
A study speaks only for the people it enrolled, and a sensitivity trial is only as good as its check that the pain really was dentine sensitivity. The 1997 guideline bases selection on a clinical diagnosis and excludes people with conflicting characteristics such as active dental treatment1, yet the 2026 Pollard review found that few studies applied every element of the accepted definition, especially ruling out other causes of tooth pain2. That matters, because a cavity, a cracked tooth or a leaking filling can hurt on cold as well, and no desensitising paste is meant for any of them.
Look for the entry rule in the methods section. In the trial by Biesbrock and colleagues (Journal of Periodontology, 2025), funded by Procter & Gamble, each of 120 adults needed two teeth scoring above one on the Schiff cold-air scale and between 10 g and 20 g on a touch probe, and people who had had gum surgery, orthodontic treatment or new fillings in the previous three months were excluded3. In the eight-week trial by Amaechi and colleagues (BDJ Open, 2021), sponsored by Sangi, a Japanese maker of nano-hydroxyapatite toothpastes, a tooth qualified with a score of at least 20 mm on a 100-mm visual analogue scale after a two-second air blast, and teeth with an abscess, excess mobility or pain from gum causes were excluded8.
The question to ask is plain: did an examiner confirm the pain on a test, and were the other causes ruled out? If the paper does not say, its result may include people whose pain no toothpaste was going to touch.
What did the comparison group get?
The comparison group usually improves too, which is why a study without one tells you little about the paste911. In a six-week double-blind trial of 120 adults by West and colleagues (Journal of Clinical Periodontology, 1997), potassium nitrate, strontium acetate and plain fluoride toothpastes all reduced sensitivity with no significant difference between them, and the plain fluoride group improved significantly on cold air9. In a single-blind crossover trial of 35 patients by Mehta and colleagues (Dental Materials, 2015), water placebo alone cut air-blast pain by 20% straight after application and by 36% at six months10. A 2019 network meta-analysis of 30 randomised trials by Hu and colleagues (Journal of Dentistry) measured a significant placebo effect across the whole network11. The page on whether it worked or just felt like it works through what that means for a reader's own experience (did it work or did it just feel like it?).
Four kinds of control turn up in the trials this site reads, and each allows a different sentence:
- A benchmark fluoride paste. In the Biesbrock trial, funded by Procter & Gamble, the control was an ordinary sodium monofluorophosphate paste, and that group's adjusted mean Schiff score still moved from about 2.45 at baseline to 2.22 at week eight3.
- Another desensitiser only. The Amaechi study, funded by Sangi, set three nano-hydroxyapatite toothpastes against a calcium sodium phosphosilicate one with no placebo or fluoride arm8. In the full text of that trial funded by Sangi, the 15% paste and the 10% paste with potassium nitrate did not differ significantly from the comparator at any visit, while the plain 10% paste fell behind it on cold at weeks six and eight and on air at week eight; none of this says how any of them compare with no treatment8.
- The same routine without the test product. In the eight-week trial of 191 adults by Hall and colleagues (Journal of the American Dental Association, 2019), sponsored by GlaxoSmithKline, a 3% potassium nitrate mouthrinse used with fluoride toothpaste was set against the toothpaste alone, with no placebo rinse12.
- A placebo and a fluoride paste together. The four-week double-blind trial of 105 adults by Vano and colleagues (Quintessence International, 2014) had both, the arrangement the guideline asks for; its funding is not stated in the PubMed record13.
Ask what the control group got and how much it improved. A result against a plain paste that also improved is a modest claim once that improvement is taken away, and a result against another desensitiser says nothing about a paste against no treatment at all.
Who knew which paste was which?
Blinding stops expectation from moving the score, and when pain is measured by asking people how much it hurt, expectation is part of what gets measured. "Double-blind" means neither the participant nor the examiner knew the allocation, as in the Vano trial13 and in the Biesbrock trial funded by Procter & Gamble3. "Examiner-blinded" means only the person scoring did not know: in the Hall trial, sponsored by GlaxoSmithKline, masking was single (investigator) and participants knew whether they were rinsing, because no placebo rinse existed12. A design with no dummy product cannot hide the treatment from the person reporting their pain, however carefully the examiner is kept in the dark. The Mehta crossover trial was single-blind, and its abstract does not say whether it was the patients or the assessor who were kept blind10.
Blinding is also the part of a paper most often left vague. The Neto review found it fully described in 37.5% of 16 mouthwash trials6, and the Pollard review found it reported inconsistently across 58 recent sensitivity trials2. The words to look for are not "double-blind" in the title but a sentence in the methods saying how the products were disguised and who held the code. Even then, blinding settles who could not see the allocation; it does not settle whether the comparison was fair or the sample big enough.
Which stimulus, and which scale?
The stimulus decides what "less sensitive" means, so start by finding what was done to the tooth and how the answer was scored. Three stimuli dominate. Air is a short blast from a dental syringe, scored by the examiner on the Schiff scale from zero to three, and it is the test most papers call "cold air" or evaporative. Touch is a calibrated probe drawn across the tooth, recorded as the force in grams at which the person first feels it, usually with a Yeaple probe. Cold means ice or cold water, usually scored by the person on a visual analogue scale, a line typically ten centimetres long running from no pain to the worst pain14. The live page on how desensitising toothpastes are tested explains each instrument in full, and this page does not repeat it.
How a scale was chosen is rarely justified. A 2026 scoping review of 71 studies by Gupta and colleagues (Periodontology 2000) found the visual analogue scale and the Schiff scale the most used measures, often side by side, with inconsistent references given for choosing them14. The same person's answers move between visits: in two standardised-stimulus clinical studies of 63 and 42 adults, Ide and colleagues (Journal of Clinical Periodontology, 2001) found reproducibility within the same subject limited even when the air and cold stimuli were controlled15. A tooth's score is not a person's score either; Olley and colleagues (2013) validated a cumulative index built from Schiff scores in a cross-sectional study of 350 adults, because the Schiff scale rates one tooth surface at a time16.
The word "cold" hides two different tests. In a 2019 systematic review and meta-analysis of six four-week trials by Alencar and colleagues (Journal of Dentistry), nano-hydroxyapatite did better than comparators on evaporative air (SMD −1.09) and on touch (SMD −0.93), but showed no difference on cold stimuli (SMD −0.17, p = 0.61)17. The 2026 network meta-analysis by Gormley and colleagues reported networks only for the Schiff cold-air score and the Yeaple touch score at two weeks, and none for ice or cold water4. A claim about "cold" should say whether it means a blast of air or a cold drink.
Few trials ask whether treatment changed daily life. A 2026 systematic review by Kitsaras and colleagues (Periodontology 2000) found 13 randomised trials since 2010, with 616 participants, that measured oral health-related quality of life, and rated the certainty of that evidence very low18.
How many people finished, and how many did the authors say they needed?
Count the people analysed, not the people recruited, and then look for the sample-size calculation, because a trial that does not state one cannot tell you whether a lack of difference means no effect or too few people. In the Amaechi trial, sponsored by Sangi, 105 people were recruited, 20 dropped out during the wash-in and 85 completed, in groups of 22, 19, 24 and 208. This site's reading of that paper does not include the planned sample size, so this page does not state it. In the Biesbrock trial, funded by Procter & Gamble, 120 people were randomised, 30 to each group, and 118 completed3.
Bigger is not the same as better. A large trial with an unblinded outcome, or with several teeth in one mouth analysed as though each belonged to a separate patient, can print a very small p value for a difference that is partly an artefact. Matranga and colleagues listed that last problem, no account of several teeth measured in the same person, among four statistical faults in trials from 2009 to 2014, alongside no adjustment for multiple comparisons, no test of whether the treatment effect changed over time, and nonparametric tests on large samples without the information needed to justify them5.
The guideline's answer to size is replication: at least two independent trials before a product is approved, so one trial of any size is still "one trial found"1.
When did it measure, and does anything show a paste working within a week?
No independent toothpaste trial on this page does. Of the toothpaste trials read here, three measured in the first week: the two with a control paste containing no desensitiser were funded by Procter & Gamble and by Haleon, each of which makes a paste it tested, and the third had no such control31920. Most of the studies read here first measured at two weeks or later: the Vano trial at two and four weeks13, the Amaechi trial, sponsored by Sangi, every two weeks up to eight8, the Hall trial, sponsored by GlaxoSmithKline, at four and eight weeks12, and the Alencar review pooled only four-week results17.
Of the ten rows in Table A, only the Biesbrock trial looked that early. In the Biesbrock trial, funded by Procter & Gamble, which makes the stannous fluoride paste tested, all three test pastes were ahead of the fluoride control on the cold-air score at day three, by 18.2% for stannous fluoride, 20.3% for an experimental oxalate paste and about 13% for potassium nitrate3. On touch at day three, the potassium nitrate paste's 10.7% improvement over control in that trial, funded by Procter & Gamble, was not significant (p = 0.197), while stannous fluoride and oxalate were ahead3. At the same day-three visit, in the trial funded by Procter & Gamble, nobody in the stannous fluoride or potassium nitrate groups and 3% of the oxalate group had complete relief in a test tooth; by week eight the figures were 53%, 28% and 35%3.
Two trials outside the table also looked early. In an examiner-blind trial of 215 adults funded by Haleon, which makes the product tested, a 5% calcium sodium phosphosilicate paste reduced sensitivity more than a sodium fluoride paste from day three, and the reduction in sensitivity grew over eight weeks (Creeth and colleagues, Journal of Dentistry, 2026)19. In a double-blind trial of 60 adults, Anand and colleagues (Acta Medica, 2017) found that a nano-hydroxyapatite paste and an 8% arginine paste did not differ at five minutes after a supervised application, at one week or at four weeks; there was no non-desensitising control, the outcome was an electrical pulp-tester threshold rather than an air or touch score, and the funding is not stated in the record20.
The careful reading, then: a few trials measured in the first week, two of the three read here were paid for by a company that makes a product under test, the third had no control without a desensitiser, and the field's guideline judges a paste over eight weeks1. A pack or web page that says a sensitivity paste "works within a week" is leaning on one or two early visits, usually in a manufacturer's own trial. Why relief tends to build over weeks is covered on the page on why relief that builds over weeks is a feature, not a flaw.
What do a p value and a confidence interval tell you, and what do they not?
A p value tells you how surprising a difference would be if the products really did the same thing, and a confidence interval tells you the range of effect sizes the data are compatible with; neither tells you whether you would notice the difference. The Alencar review's cold result, p = 0.61, is not evidence that nano-hydroxyapatite does nothing on cold: it says six four-week trials could not separate it from comparators on that stimulus17.
At the other end, the Hall trial, sponsored by GlaxoSmithKline, reported P < .0001 for most comparisons between the rinse and toothpaste alone12. That says the difference is very unlikely to be chance. It says nothing about size, which is why the registry record for that trial, sponsored by GlaxoSmithKline, is more useful: the week-eight difference on the zero-to-three Schiff scale was −1.10 points (95% CI −1.28 to −0.92), with the rinse group's score falling by 1.47 and the toothpaste-alone group's by 0.3712.
An interval that crosses zero means the data fit no effect as well as some effect. In the 2015 meta-analysis of 31 randomised trials by Bae and colleagues (Journal of Clinical Periodontology), potassium toothpastes had a standardised mean difference against placebo of −1.28, with an interval of −2.05 to −0.51 that stays below zero, while strontium toothpastes, at 0.05, showed no significant effect, and heterogeneity between trials was high, with I² of 86% to 95% for the five kinds of paste that showed an effect21. A standardised mean difference divides by the spread of scores so that trials using different scales can be pooled, which is why that −1.28 has no unit and cannot be read as points on any scale21.
A relative percentage is a ratio of changes, not a share of people. The 2023 systematic review and meta-analysis of 44 clinical trials by Limeback and colleagues (Biomimetics), two of whose three authors are employees of Dr Kurt Wolff GmbH, a German maker of hydroxyapatite toothpastes, reported a 39.5% greater reduction with hydroxyapatite products than with placebo (95% CI 30.1 to 48.9), which is not 39.5% of people, and found no significant difference against other desensitising agents22.
A mean difference on a named scale keeps the unit: the 2026 Gormley network meta-analysis of 93 randomised trials gave stannous fluoride a difference of −0.85 on the Schiff score against a benchmark fluoride paste at two weeks (95% CI −1.08 to −0.62; ten studies; high confidence)4. A ranking probability is not a difference at all: in Hu's 2019 network meta-analysis, nano-hydroxyapatite had a 60% and a 67% probability of ranking first at two and four weeks, and the abstract reports no significant pairwise superiority11. Many p values from one trial also raise the chance that one is small by luck, which is why the Matranga review counted missing adjustment for multiple comparisons among its four faults5.
No agreed threshold was found for how big a drop on the Schiff scale or a visual analogue scale a person would notice. The 1997 guideline asks for "clinically significant changes" without defining one1. The 2026 Gormley systematic review set its own working rule, treating a mean difference of 0.5 on the Schiff score or five on the Yeaple score as clinically meaningful, a judgement made by the reviewers rather than a threshold derived from patients4. A screen of the titles from PubMed searches for a minimal clinically important difference in dentine hypersensitivity, run for this site in September 2026, found no paper deriving one for the Schiff score or a visual analogue scale. So no study on this page, and no figure from S3's own consumer trial, can say its result is one a person would feel.
| Measure | What it means | Example from a study on this page | The misreading it invites |
|---|---|---|---|
| p value | How surprising the difference would be if the products were equal | Alencar 2019: cold stimulus p = 0.61 | Reading a large p as "no effect", or a small one as "a big effect" |
| Confidence interval | The range of effects the data fit | Bae 2015: potassium −2.05 to −0.51 | Ranking products whose intervals overlap |
| Relative percentage | Change in one group relative to change in another | Limeback 2023: 39.5% versus placebo | Reading it as the share of people helped |
| Standardised mean difference | Difference divided by the spread of scores, no unit | Bae 2015: potassium −1.28 | Turning it back into points on a scale |
| Mean difference on a named scale | Difference in the scale's own units | Gormley 2026: stannous fluoride −0.85 Schiff points at two weeks | Assuming a person notices a change of that size |
| Ranking probability | Chance a product ranks first in a network | Hu 2019: nano-hydroxyapatite 67% at four weeks | Treating a ranking as a significant difference |
Who paid, and does it matter?
Most trials of the main actives were paid for by the companies that sell them, and a reader should know the funder of every study they lean on, but funding alone does not show that a result is wrong4. The 2026 Gormley review counted industry funding behind 96% of the stannous fluoride studies it included, 86% of arginine, 76% of potassium and 33% of nano-hydroxyapatite studies4.
Name the funder and what it sells: Procter & Gamble funded the Biesbrock trial and makes the stannous fluoride paste it tested3; Sangi, a maker of nano-hydroxyapatite toothpastes, funded the Amaechi study and made its test toothpastes8; GlaxoSmithKline was the sponsor of the Hall rinse trial12; Haleon funded the Creeth trial of its own calcium sodium phosphosilicate paste19; and two authors of the Limeback review are employed by Dr Kurt Wolff GmbH, a maker of hydroxyapatite toothpastes22. The trial that looks most like this company's own pairing of ingredients, nano-hydroxyapatite with potassium nitrate, was sponsored by the maker of the pastes it tested8.
Laboratory work can have a supplier behind it too. In the 2017 in vitro study of 62 dentine blocks by Jena and colleagues (Journal of Conservative Dentistry), the authors declared no funding or conflicts, but the nanoXIM nano-hydroxyapatite test paste was provided by its supplier, Fluidinova23. nanoXIM CarePaste is also the nano-hydroxyapatite used in S3, which is why this page says so wherever it cites that study.
Some methods papers declare no competing interests, as the Newcombe short communication does7, and the Pollard group declares none on its 2026 scoping review and Delphi paper, although the same Bristol group declares dentifrice-manufacturer fees and grants on its efficacy papers2. What funding changes is how much independent replication a result needs before you rely on it, not whether to read it. The live guide on checking a brand's evidence runs that kind of check on UK brands; this page checks studies, not brands.
How do ten real studies answer the eight questions?
Table A runs the eight questions on nine published studies and on S3's own consumer trial of 51 adults over eight weeks. Every cell comes from this site's study card for that paper or, for the last row, from S3's claims register; where neither records an answer, the cell says so rather than guessing.
| Study | Who, and how confirmed | Comparison group | Blinding | Stimulus and scale | Size: planned and analysed | First and last time point | Effect measure, and what the p or interval says | Funder |
|---|---|---|---|---|---|---|---|---|
| Limeback 2023, systematic review and meta-analysis | Pooled trials of hydroxyapatite products in people; entry rules vary by trial | Placebo, fluoride or other desensitisers, pooled | Varies by trial; more than half of trials rated high quality (GRADE) | Not in the card | 44 clinical trials; participant total not given | Not in the card; no time to effect reported | Relative difference 39.5% versus placebo (95% CI 30.1 to 48.9); versus other desensitisers not significant | No external funding; two of three authors employed by Dr Kurt Wolff GmbH, a hydroxyapatite toothpaste maker |
| Amaechi 2021, double-blind RCT | Adults with a tooth scoring at least 20 mm on a 100-mm VAS to a two-second air blast; gum and other causes excluded | Calcium sodium phosphosilicate paste; no placebo or fluoride arm | Double-blind | Ice-cold and air; VAS | Planned: not in the card; 105 recruited, 85 completed | Week 2 to week 8, every two weeks | Change from baseline P < 0.001 in all arms; 15% and 10% + potassium nitrate pastes not significantly different from the comparator; plain 10% paste behind it on cold at weeks 6 and 8 and on air at week 8 (full text) | Sponsored by Sangi, maker of the nano-hydroxyapatite test pastes |
| Kulal 2016, in vitro (SEM) | No people: 40 dentine discs from extracted teeth | Untreated discs in saline | Not in the card | No stimulus; tubules occluded under a microscope | 10 discs per group | Seven days of application | Percentages; one p value in the abstract differs from the full-text table | None declared |
| Vano 2014, double-blind RCT | Adults; how sensitivity was confirmed is not in the card | Fluoride paste and placebo | Double-blind | Cold air and touch, plus VAS | Planned: not in the card; 105 (35 per group) | Weeks 2 and 4 | P < .001 against both controls at both visits | Funding not stated in the PubMed record |
| Biesbrock 2025, double-blind RCT | Two teeth with Schiff above 1 and probe 10–20 g; recent dental treatment excluded | Sodium monofluorophosphate fluoride paste | Double-blind | Cold air (Schiff) and touch (Yeaple) | Planned: not in the card; 120 randomised, 118 completed | Day 3 to week 8, then week 11 after three weeks on the control paste | Percentage improvement over control; week-8 differences between the actives not significant | Funded by Procter & Gamble, maker of the stannous fluoride paste; seven of eight authors employees |
| Bae 2015, systematic review and meta-analysis | Pooled randomised trials; entry rules vary | Placebo | Not in the card | Mixed scales combined as SMD | 31 randomised trials | Not in the card | SMD, e.g. potassium −1.28 (95% CI −2.05 to −0.51); strontium not significant; I² 86–95% across the five paste types with an effect | Funding not stated in the PubMed record; no manufacturer among the affiliations |
| Alencar 2019, systematic review and meta-analysis | Pooled four-week trials | Placebo, negative control or other treatments, pooled | Not in the card | Evaporative, tactile and cold, analysed separately; SMD | Six trials; participant total not stated | Four weeks only | SMD by stimulus; cold −0.17, p = 0.61, not significant | Funding not stated in the PubMed record |
| Hall 2019, examiner-blinded RCT (mouthrinse) | Adults; entry score not in the card | Fluoride toothpaste alone, no placebo rinse | Examiner only; participants knew whether they rinsed | Air (Schiff), touch (Yeaple), visual rating scale | Planned: not in the card; 191 (95 and 96 per arm) | Weeks 4 and 8 | Week-8 Schiff difference −1.10 (95% CI −1.28 to −0.92); P < .0001 | Sponsored by GlaxoSmithKline |
| Jena 2017, in vitro (SEM) | No people: 62 dentine blocks from extracted molars | Untreated blocks (2) and three other desensitising pastes | Not in the card | No stimulus; tubules occluded under a microscope | 15 blocks per test group, 2 controls | 14 days of brushing | Percentages; compare arms only inside this paper | None declared; nanoXIM test paste provided by its supplier, Fluidinova |
| S3 consumer trial (ADSL) | 51 adults with sensitive teeth; recruitment criterion still to be confirmed in writing by ADSL | Not stated in the register; no head-to-head arm | Not stated in the register | Panellists' own reports; no examiner stimulus stated in the register | Planned: not stated in the register; 51 in the final report | Over eight weeks | No figure quoted on this page; results are reported in perception wording only | Not stated in the register; run by ADSL, an independent third party, in Devon to Good Clinical Research Practice |
Read across rather than down, the table shows that only one of the ten measured before week two, and that one was funded by the maker of one of its test pastes3. Two rows are laboratory work on extracted teeth: Kulal and colleagues' 2016 study of 40 dentine discs, whose abstract repeats one p value (0.235) for a comparison that the full-text table prints as 0.154, is a reminder to check the table24. Neither laboratory row describes a benefit to a person, and an occlusion percentage from one laboratory protocol cannot be set beside one from another.
How does S3's own trial answer the eight questions?
S3's own trial answers some of the eight questions and leaves others blank. Its design, as S3's claims register records it, is an independent third-party consumer trial of 51 adults with sensitive teeth over eight weeks, run by ADSL in Devon to Good Clinical Research Practice. Put through the same questions as Table A:
- Who was in it: 51 adults with sensitive teeth; the register notes that the recruitment criterion is still to be confirmed in writing by ADSL, so this page does not describe the panel further.
- What the comparison group got: the register records no head-to-head arm, and whether any control group existed is not stated.
- Who knew which paste was which: not stated in the register.
- Stimulus and scale: results are worded as what panellists said or reported, which is self-report, not an examiner's air or touch score.
- How many finished, against how many were needed: 51 in the final report; a planned sample size is not stated.
- When it measured: over eight weeks.
- What the numbers say: this page quotes no figure from it, because the question here is how the study was built, not its outcome.
- Who paid: not stated in the register; the trial was run by an independent third party.
S3's own trial is a consumer trial of 51 adults with no head-to-head arm, and the register does not state whether it was blinded or had a control group; on those two questions it answers less well than several of the studies in Table A. None of the published trials or laboratory studies on this page tested this company's toothpaste, and none tested potassium nitrate, hydroxyapatite and fluoride together in one paste. The gaps are requests rather than verdicts: blinding, a control arm, a planned size and the funder may sit in the full ADSL report without yet being in the register. How a consumer trial differs from a clinical trial is the subject of a separate page in this site's science section.
When is a study not the thing to rely on at all?
This page is written by the company that makes the toothpaste in the last row of Table A, so weigh what follows knowing that. A study is the wrong guide when the pain is not dentine sensitivity, when the evidence points most firmly to an active our paste does not contain, or when all you have is a laboratory picture.
When the pain may be something else. Pain that lingers after the cold has gone, wakes you at night or sits in one tooth is a question for a dentist, not for any trial. The NHS says to see a dentist about toothache that has gone on for more than two days, that painkillers do not settle, or that comes with a high temperature, pain on biting, red gums or a bad taste, and to go to A&E if swelling reaches the eye or neck or makes it hard to breathe, swallow or speak25.
When you want the highest-confidence two-week evidence. Look at stannous fluoride: in the 2026 Gormley network meta-analysis its cold-air result carried a high confidence rating, and the authors name stannous fluoride and arginine pastes as the first-line self-care choices4. Neither stannous fluoride nor arginine is among S3's actives, which are potassium nitrate and two forms of hydroxyapatite, with fluoride kept in. We have an obvious reason to prefer our own formula, and that review could not rank it, because no trial in it tested this combination.
When a dentist can treat the tooth directly. A treatment applied in the chair can be the better route: in one single-blind crossover trial of 35 patients, a professionally applied calcium phosphate paste relieved air-blast pain more than water placebo at every recall up to six months, although the placebo alone cut that pain by 36%10.
When all you have is a laboratory picture. A microscope image of plugged tubules on an extracted tooth is never a reason to buy a paste, whoever's paste it is, and the two laboratory rows in Table A measured discs and blocks, not people. If a paste has not helped by eight weeks, the span the guideline sets for most trials, the page on how long to try a sensitivity toothpaste before seeing a dentist sets out what to do next, and the Sensitivity Journal explains how sensitive toothpastes work, and why yours might not be working for you.
Frequently asked questions
Is a bigger sensitivity study always a better one?
No. Size helps only if the people analysed were the right people, the comparison was fair and the analysis treated several teeth in one mouth as one person's data, and the Matranga review judged the statistical methods inappropriate in 77.1% of 35 parametric trials5. The field's guideline asks for replication, at least two independent trials, rather than one big one1.
What does p < 0.05 mean in a toothpaste trial?
It means a difference at least that large would turn up less than one time in twenty if the pastes were really equal. It does not mean the paste works for most people, and it does not mean the difference is big enough to feel, because no agreed threshold for a noticeable change on the Schiff scale or a visual analogue scale was found.
Does any toothpaste work within a week?
The evidence cannot say so for any toothpaste. The one trial in this page's table that measured at day three, funded by Procter & Gamble, found three actives ahead of a fluoride control on cold air, but almost nobody had complete relief in a test tooth that early, and the field's guideline judges a paste over eight weeks13.
What kind of study is S3's own trial?
S3's claims register describes a consumer trial in which ADSL, an independent third party working in Devon to Good Clinical Research Practice, followed 51 adults who have sensitive teeth for eight weeks. With no head-to-head arm and results worded as what panellists reported, it is called a consumer trial and not a clinical trial, and the register is silent on blinding, a control group, a planned sample size and the funder.
Why do so many sensitivity trials have a manufacturer behind them?
Trials cost money, and the companies that sell the pastes have the most direct reason to pay for them: the 2026 Gormley review counted industry funding behind 96% of stannous fluoride and 86% of arginine studies4. That is a reason to look for independent replication and to read the methods closely, not a reason to dismiss a result on its funder alone.
Where S3 sits
S3 has been tested at the University of Reading and in an independent consumer trial of 51 adults with sensitive teeth over eight weeks. Neither is a randomised controlled trial of S3, which is why this site reads its ingredients' trials the way this page describes.
S3 Sensitivity Science™ pairs potassium nitrate with two hydroxyapatites and adult-strength fluoride in one daily paste.
See the toothpasteOne tube, three actives: S3 Sensitivity Science™ puts potassium nitrate to work on the nerve, nano-hydroxyapatite inside the tubule and biomimetic hydroxyapatite on the surface, with 1450 ppm fluoride kept in. Calm, strengthen, protect: the three actions sensitive teeth need, in one daily toothpaste. The formula is patent-pending S3 Repair Technology™, UK application GB2604755.5. More than 20 practising UK dentists own a stake in S3, and nine founding dentists advise on the formulation. Read more about S3.