How to Read a Non-Inferiority Trial in Botulinum Toxin Research

A practical guide to understanding non-inferiority trials, margins, confidence intervals, ITT and per-protocol analyses, with examples from randomized Neuronox studies.

9/23/202610 min read

worm's-eye view photography of concrete building
worm's-eye view photography of concrete building

What Is a Non-Inferiority Trial?

A non-inferiority trial asks whether a new or test treatment is not unacceptably worse than an active comparator for a prespecified primary outcome.

The phrase “not unacceptably worse” is important.

A non-inferiority trial does not begin by assuming that two treatments must produce exactly the same result.

Instead, the study defines in advance how much worse the test treatment could be before that difference would be considered clinically unacceptable.

That boundary is the non-inferiority margin.

A valid interpretation therefore requires more than reading the response percentages.

The reader needs to know:

  • the primary endpoint;

  • the direction of the treatment difference;

  • the prespecified non-inferiority margin;

  • the confidence interval;

  • the analysis population; and

  • whether the trial's statistical assumptions were appropriate.

Why Use a Non-Inferiority Design?

When an effective active treatment already exists, a placebo-controlled superiority study may be inappropriate, unethical in some settings, or less clinically informative.

A non-inferiority design allows investigators to ask whether a test treatment retains an acceptable level of efficacy relative to an established active comparator.

However, the design is statistically demanding.

A poorly chosen margin, inappropriate comparator, excessive protocol deviations, missing data or methodological weaknesses can make a non-inferiority conclusion difficult to interpret.

Non-Inferiority Is Not the Same as “No Significant Difference”

This is one of the most important concepts.

Suppose two treatments are compared and the difference between them has a p-value greater than 0.05.

That does not prove non-inferiority.

A conventional non-significant superiority test means the study did not establish a statistically significant difference.

It does not establish that the treatments are sufficiently similar.

To demonstrate non-inferiority, the study must be designed as a non-inferiority trial and evaluate the treatment difference against a prespecified non-inferiority margin.

Therefore:

“No statistically significant difference” ≠ “non-inferior.”

What Is the Non-Inferiority Margin?

The non-inferiority margin, often represented by Δ, defines the maximum clinically acceptable loss of efficacy relative to the active comparator.

The margin must be specified before the results are known.

It should also have a clinical and statistical justification.

A very wide margin makes non-inferiority easier to demonstrate but may allow a clinically important loss of effect.

A very narrow margin demands greater precision and often requires a larger sample size.

For this reason, the margin is not simply a technical number hidden in the statistical section.

It is central to the clinical meaning of the study.

Why Can Different Trials Use Different Margins?

Because the margin depends on factors such as:

  • the endpoint;

  • measurement scale;

  • expected comparator effect;

  • historical evidence;

  • clinical context; and

  • what degree of efficacy preservation is considered acceptable.

A margin of −20 percentage points for a binary response endpoint and a margin of 0.45 points for a continuous clinical scale are fundamentally different statistical quantities.

They should not be compared as though one trial used a “better” or “stricter” number simply because it is numerically smaller.

Why Does the Confidence Interval Matter?

Non-inferiority is generally evaluated using the confidence interval around the treatment difference.

The relevant boundary depends on how the difference is defined.

For example, if the treatment difference is calculated as:

Neuronox − comparator

and larger response rates are better, the lower confidence bound may be compared with a negative non-inferiority margin.

If the outcome is defined so that a larger numerical value represents worse performance, the relevant boundary may instead be the upper confidence bound.

Therefore, readers should not memorize a rule such as “always look at the lower CI.”

The first question should be:

How was the treatment difference defined, and which direction represents worse efficacy?

A Simple Example

Imagine a response-rate study where:

Test treatment: 90%

Comparator: 92%

The observed difference is:

90% − 92% = −2 percentage points

Suppose the prespecified non-inferiority margin is −10 percentage points.

If the lower confidence bound around the treatment difference is −7%, the entire relevant confidence interval remains above the −10% boundary.

The trial may therefore meet the non-inferiority criterion.

But if the lower bound is −13%, the confidence interval includes differences worse than the allowed −10% margin.

Non-inferiority has not been established, even if the observed response rates look numerically similar.

Why Is the Point Estimate Alone Not Enough?

Because a point estimate does not show the uncertainty around the estimate.

Two treatments might differ by only 1% in the observed sample, but if the confidence interval is very wide, clinically important inferiority may still be compatible with the data.

This is why confidence intervals are central to non-inferiority interpretation.

Why Are ITT and Per-Protocol Analyses Important?

In superiority trials, intention-to-treat analysis is often considered conservative because non-adherence and crossover may reduce apparent differences between groups.

In non-inferiority trials, that same tendency toward similarity can potentially make non-inferiority easier to demonstrate.

Per-protocol analysis evaluates participants who sufficiently followed the trial protocol, but it can also introduce selection bias.

For this reason, regulatory and statistical guidance emphasizes examining both analysis approaches in non-inferiority research when appropriate.

When ITT/full-analysis-set and per-protocol results lead to consistent conclusions, confidence in the robustness of the finding is increased.

What Is the ITT Population?

An intention-to-treat approach generally aims to analyze participants according to the treatment group to which they were randomized.

The exact implementation varies among trials.

Some publications use terms such as:

ITT

modified ITT

or

full analysis set.

These should not automatically be treated as identical.

The study's methods should be checked to understand exactly who was included.

What Is the Per-Protocol Population?

The per-protocol population generally includes participants who met predefined criteria for sufficient adherence to the study protocol.

Exclusions may involve major protocol violations, insufficient treatment exposure, missing primary-outcome data or other prespecified reasons.

Per-protocol analysis can provide a useful complementary view in a non-inferiority study.

However, because participants are excluded after randomization, the analysis can also be affected by selection bias.

It should therefore not be interpreted in isolation.

How Did the Neuronox Glabellar Trial Use Non-Inferiority?

The randomized, double-blind, active-controlled Phase III study enrolled 314 participants with moderate-to-severe glabellar lines. [1]

Participants received:

NBoNT/Neuronox 20 U

or

onabotulinumtoxinA 20 U

under the study protocol.

The primary endpoint was investigator-assessed response at maximum frown at week 4.

The observed responder rates were:

Neuronox/NBoNT: 93.7% (133/142)

onabotulinumtoxinA: 94.5% (138/146). [1]

The study's primary analysis met its prespecified non-inferiority criterion.

The correct interpretation is:

Neuronox demonstrated non-inferiority to the active comparator for the study-defined primary endpoint under the studied protocol.

It does not mean the two products were proven identical in every clinical property.

What Did the Blepharospasm Trial Show?

The randomized comparative trial evaluated Meditoxin, identified in the publication as another name for Neuronox, against BOTOX in essential blepharospasm. [2]

The ITT population included 60 participants:

Meditoxin/Neuronox: 31

BOTOX: 29

Improvement in severity of spasm at week 4 was:

Neuronox/Meditoxin: 90.3%

BOTOX: 86.2%.

The prespecified non-inferiority margin was −15 percentage points.

The lower confidence limit was:

−1.76% in the ITT analysis

and

−1.64% in the per-protocol analysis. [2]

Both remained above the −15% non-inferiority boundary.

Therefore, the study met its non-inferiority criterion in both analysis populations.

Neuronox Had a Higher Percentage in That Trial. Does That Mean Superiority?

No.

The observed response percentage was numerically higher for Neuronox/Meditoxin.

But the study's statistical objective was non-inferiority.

A numerically larger percentage alone does not establish superiority.

A superiority conclusion requires appropriate hypothesis testing and interpretation according to the study's statistical plan.

This distinction is essential:

numerically higher ≠ statistically proven superior.

What Happened in the Cerebral-Palsy Trial?

The multicentre randomized double-blind trial included 119 participants in the ITT population:

Neuronox: 60

BOTOX: 59. [3]

The primary endpoint was Physicians' Rating Scale response at week 12.

Response rates were:

Neuronox: 48.3%

BOTOX: 49.2%.

The prespecified non-inferiority margin was −20 percentage points.

The lower confidence bound in the ITT analysis was −12.57%, which remained above the −20% boundary. [3]

The per-protocol analysis also met the criterion, with a lower confidence bound of −11.58%.

Thus, the primary analysis supported non-inferiority under the study protocol.

Why Is the Cerebral-Palsy Example Useful?

Because the raw response rates alone do not tell the whole story.

48.3% and 49.2% look close.

But “they look close” is not the statistical reason non-inferiority was demonstrated.

The key evidence is that the confidence interval remained within the prespecified non-inferiority boundary.

That is the difference between visual similarity and a formal non-inferiority conclusion.

What Did the Post-Stroke Trial Do Differently?

The post-stroke upper-limb spasticity trial used a continuous endpoint rather than a response percentage. [4]

The primary endpoint was change from baseline in wrist-flexor Modified Ashworth Scale at week 4.

Results were:

Neuronox: −1.39 ± 0.79

BOTOX: −1.56 ± 0.81.

The prespecified non-inferiority margin was 0.45 MAS points.

The relevant upper confidence limit for the treatment difference was 0.40. [4]

Because 0.40 remained below the 0.45 non-inferiority boundary, the primary analysis met the criterion.

Why Did This Trial Use an Upper Confidence Limit Instead?

Because the direction and definition of the treatment difference were different.

For this endpoint, the concern was whether the Neuronox result could be worse than the comparator by more than the positive margin of 0.45 points.

Therefore, the relevant boundary was the upper confidence limit.

This is a useful reminder that non-inferiority cannot be interpreted by mechanically looking at one side of every confidence interval.

The direction of the endpoint must be understood first.

What Do the Four Neuronox Trials Show Together?

The four trials used different:

  • populations;

  • indications;

  • primary endpoints;

  • outcome scales;

  • non-inferiority margins; and

  • follow-up schedules. [1–4]

Yet their primary analyses shared one important feature:

each met the prespecified non-inferiority criterion defined for that trial.

This supports the evidence-based statement:

The primary analyses of the four reviewed randomized comparative Neuronox studies met their respective prespecified non-inferiority criteria.

This is stronger and more accurate than simply saying:

“Neuronox is the same as BOTOX.”

Does Non-Inferiority Mean Equivalent?

No.

Non-inferiority and equivalence are different statistical designs.

A non-inferiority trial primarily aims to rule out an unacceptable degree of worse efficacy in one direction.

An equivalence trial generally aims to show that the treatment difference lies within prespecified boundaries in both directions.

Therefore:

non-inferiority ≠ equivalence.

Does Non-Inferiority Mean Interchangeable?

No.

A non-inferiority result for a defined endpoint does not establish universal interchangeability of:

  • Units;

  • doses;

  • indications;

  • injection protocols;

  • formulations;

  • safety profiles; or

  • clinical performance in other populations.

The conclusion must remain linked to the study that produced it.

Does Non-Inferiority Mean Same Safety?

No.

Efficacy non-inferiority does not automatically demonstrate safety equivalence.

Safety analyses may have different statistical power, and uncommon adverse events may not be detectable in relatively small clinical trials.

Safety data should therefore be evaluated separately.

Does Non-Inferiority Mean Same Duration?

No.

Unless duration was the prespecified endpoint being tested with an appropriate non-inferiority design, efficacy non-inferiority at one time point does not prove identical duration.

For example, a week-4 non-inferiority result does not automatically prove equal persistence at week 16.

Does Non-Inferiority Mean Same Diffusion?

No.

The four core Neuronox studies did not use their primary non-inferiority analyses to establish a unique diffusion profile.

Therefore, their non-inferiority results should not be converted into claims of identical or superior spread.

Can a Non-Inferiority Trial Also Demonstrate Superiority?

Potentially, but this requires appropriate statistical planning and analysis.

One should not infer superiority simply because the test treatment's observed result is numerically better.

The publication and statistical plan must support the superiority conclusion.

For the core Neuronox evidence discussed here, the key supported conclusion is non-inferiority for the prespecified primary endpoints.

Why Does the Choice of Comparator Matter?

A non-inferiority trial depends on the active comparator having an established and predictable effect under the study conditions.

If the comparator performs unexpectedly poorly, two ineffective treatments could appear similar.

This is one reason the validity of a non-inferiority trial depends not only on the final confidence interval but also on study design, conduct and historical evidence supporting the comparator effect.

What Is Assay Sensitivity?

Assay sensitivity refers to the trial's ability to distinguish an effective treatment from a less effective or ineffective treatment.

In a superiority trial, a clear difference can demonstrate that the study was capable of detecting a treatment effect.

In a two-arm non-inferiority trial without placebo, assay sensitivity cannot simply be observed from a test-versus-placebo difference.

It must be supported by the historical evidence for the active comparator and by the similarity of the current study to conditions under which the comparator's efficacy was previously established.

What Is the Constancy Assumption?

The constancy assumption is the assumption that the active comparator would retain an effect in the current trial similar enough to the effect demonstrated in historical evidence used to justify the non-inferiority margin.

If patient populations, endpoints, concomitant treatments or trial conduct change substantially, this assumption may become less secure.

This is another reason margin selection should not be viewed as an arbitrary number.

What Should Readers Check in Every Non-Inferiority Paper?

A practical checklist is:

1. What is the primary endpoint?

2. Is higher or lower better?

3. How is the treatment difference calculated?

4. What is the prespecified non-inferiority margin?

5. Why was that margin chosen?

6. Which confidence interval is used?

7. Which boundary must remain within the margin?

8. What is the point estimate?

9. What are the ITT/full-analysis-set results?

10. What are the per-protocol results?

11. Are the analyses consistent?

12. How much missing data or protocol deviation occurred?

13. Was the active comparator appropriate?

14. Does the conclusion remain limited to the studied endpoint, population and protocol?

What Are Common Misinterpretations?

“There was no significant difference, so they are non-inferior.”

Incorrect.

“The percentages are almost the same, so they are equivalent.”

Incorrect.

“Neuronox had the higher percentage, so it was superior.”

Not established from the non-inferiority result alone.

“The trial passed non-inferiority, so the products are interchangeable.”

Incorrect.

“All non-inferiority margins can be compared numerically.”

Incorrect. Margins depend on endpoint scale and study design.

“If the primary endpoint is non-inferior, all secondary outcomes must also be non-inferior.”

Incorrect.

Each outcome requires its own appropriate analysis.

Why Is This Especially Important When Reading Neuronox Research?

Because Neuronox has several direct randomized comparative trials.

That is a meaningful strength of the evidence base.

But using those studies correctly requires preserving what they actually demonstrated.

The strongest evidence-based statement is not:

“Neuronox and BOTOX are identical.”

It is:

“Across four published randomized comparative studies reviewed here, the primary analyses for Neuronox met the respective prespecified non-inferiority criteria in glabellar lines, essential blepharospasm, cerebral-palsy-associated spastic equinus and post-stroke upper-limb spasticity.” [1–4]

This statement tells the reader exactly why the evidence is important without turning non-inferiority into an unsupported superiority or interchangeability claim.

Bottom Line

A non-inferiority trial is not a study showing that “nothing is different.”

It is a trial designed to determine whether a test treatment can rule out an unacceptable loss of efficacy, according to a margin specified before the results are known.

To interpret it correctly, read:

the endpoint

the margin

the treatment-difference direction

the confidence interval

and

the analysis population.

For Neuronox, four published randomized comparative studies in the current evidence set used non-inferiority frameworks with different endpoints and margins.

Their primary analyses met their respective prespecified criteria. [1–4]

That provides meaningful product-specific comparative evidence.

The scientifically accurate conclusion remains:

non-inferiority under the studied protocol—not universal equality, superiority or interchangeability.

References

  1. Won CH, Lee HM, Lee WS, Kang H, Kim BJ, Kim WS, Lee JH, Lee DH, Huh CH. Efficacy and safety of a novel botulinum toxin type A product for the treatment of moderate to severe glabellar lines: a randomized, double-blind, active-controlled multicenter study. Dermatol Surg. 2013;39(1 Pt 2):171–178. doi:10.1111/dsu.12072. PMID:23301821.

  2. Yoon JS, Kim JC, Lee SY. Double-blind, randomized, comparative study of Meditoxin versus Botox in the treatment of essential blepharospasm. Korean J Ophthalmol. 2009;23(3):137–141. doi:10.3341/kjo.2009.23.3.137. PMID:19794937.

  3. Kim K, Shin HI, Kwon BS, Kim SJ, Jung IY, Bang MS. Neuronox versus BOTOX for spastic equinus gait in children with cerebral palsy: a randomized, double-blinded, controlled multicentre clinical trial. Dev Med Child Neurol. 2011;53(3):239–244. doi:10.1111/j.1469-8749.2010.03830.x. PMID:21087238.

  4. Seo HG, Paik NJ, Lee SU, Oh BM, Chun MH, Kwon BS, Bang MS. Neuronox versus BOTOX in the treatment of post-stroke upper limb spasticity: a multicenter randomized controlled trial. PLoS One. 2015;10(6):e0128633. doi:10.1371/journal.pone.0128633. PMID:26030192.

  5. U.S. Food and Drug Administration. Non-Inferiority Clinical Trials to Establish Effectiveness: Guidance for Industry. November 2016.

  6. International Council for Harmonisation. ICH E9: Statistical Principles for Clinical Trials. 1998.

  7. Piaggio G, Elbourne DR, Pocock SJ, Evans SJW, Altman DG; CONSORT Group. Reporting of noninferiority and equivalence randomized trials: extension of the CONSORT 2010 statement. JAMA. 2012;308(24):2594–2604. doi:10.1001/jama.2012.87802.

• COMPLETE SITE INDEX

BotoxWiki Reference Sitemap

BotoxWiki provides evidence-focused educational content with transparent sourcing and editorial standards.

Clinical & Anatomy

Evidence & Review

Editorial & Trust

Reference Standard

BotoxWiki provides evidence-based content, medical review, and a neutral editorial approach to botulinum toxin formulations and anatomical guidance.

© BotoxWiki Clinical Reference Directory.