1. Home
  2. Drugs
  3. Science and Research | Drugs
  4. Regulatory Science in Action
  5. A Novel Stratum-Specific Trial Design Improves Statistical Power for Rare Diseases with Heterogenous Clinical Symptoms
  1. Regulatory Science in Action

A Novel Stratum-Specific Trial Design Improves Statistical Power for Rare Diseases with Heterogenous Clinical Symptoms

CDER researchers developed a new stratum-specific statistical design shown to offer substantial increases in statistical power for clinical trials with small subpopulations that may be beneficial for use in drug clinical trials for many rare diseases. When used with specific data analysis methods, the trial design allows sponsors to enroll the same number of patients as they would for a traditionally designed trial—and offers a higher chance for gathering the evidence needed to support regulatory approval of a new drug. This article reflects the views of the researchers and should not be construed to represent FDA’s policies.

Background and regulatory science challenge

In drug development for a given rare disease, there can be challenges with statistical power when using traditional clinical trial designs. This is partly due to the small number of patients willing and able to participate in clinical trials due to the fewer number of people living with rare diseases compared with common diseases. Additionally, people living with a rare disease often do not uniformly experience the same symptoms, which are evaluated as endpoints in clinical trials, and this heterogenous clinical presentation can limit clinical trial enrollment in an already small patient population further diminishing the statistical power of the trial. New research from the Center for Drug Evaluation and Research (CDER) addresses these challenges through the development of a stratum-specific trial design.

A novel trial design applicable to many rare diseases

CDER statisticians in the Office of Biostatistics have recently identified a common challenge in rare disease drug development: not being able to confidently prespecify a single primary efficacy endpoint applicable to enough trial participants to sufficiently power a trial. Due to differing clinical manifestations within a rare disease population, many patients often do not have one or more of the symptoms under evaluation in a clinical trial, which diminishes the power of a traditional trial design. Reducing enrollment to patients who experience all the symptoms under investigation, or a particular subset of them, could limit the trial’s power.

How can this research advance drug development for rare diseases?

Clinical trials for rare diseases often lack the number of participants needed to create a study with sufficient power that can then determine the efficacy of a new drug. Also, trial designs typically evaluate a specific symptom or set of symptoms that not all trial patients may be experiencing. To alleviate these problems for rare disease clinical trials, CDER statisticians developed the stratum-specific trial design. The design allows sponsors to enroll a broad population of patients and increases the power of a clinical trial in which the set of symptoms experienced by patients would be expected to differ across patients at the time they start the trial.

 

To address this common challenge in rare disease drug development, CDER statistician-researchers created and evaluated a stratum-specific design, where patients within each subpopulation can be randomly assigned to the control or treatment arm. The evaluation of treatment efficacy relies only on endpoints that measure the symptoms that patients in each subpopulation are experiencing at baseline. Statistical testing uses all information relevant to each patient—and only that information.

Using computer simulations, the researchers examined trial designs using global tests. Included in the evaluation was the stratum specific design and two traditional trial designs in which all patients are considered a single population. Traditional trial 1 had the same endpoints regardless of the symptoms the patients experienced, and traditional trial 2 only enrolled patients experiencing the symptoms evaluated as endpoint. As shown in Figure 1, the hypotheses tested and the conclusions drawn from using a global test differed in the three trials:

  • Traditional trial 1 tested whether the treatment had an effect for any endpoint among all patients.
  • Traditional trial 2 tested whether the treatment had an effect for any endpoint in patients with both symptoms.
  • The stratum-specific trial tested whether the treatment had an effect for any of the baseline, symptomatic endpoints in one or more subgroups.

CDER researchers performed the tests shown in Figure 1 by using the permutation methods shown in Figure 2. A key limitation of global tests is that they cannot specify which group of patients experienced an effect on which endpoint. So once there was initial evidence that the drug is working for at least one endpoint for some group of patients, researchers performed additional statistical tests on individual endpoints to learn which patient group experienced a drug treatment effect.

Figure 1

Figure 1. Schematic comparing the stratum-specific trial design with the two traditional designs.

Figure 1. Schematic comparing the stratum-specific trial design with the two traditional designs.


Assessing the performance of the stratum-specific trial

Researchers performed extensive computer simulations to compare the performance of the global tests in the three trials, varying factors such as the number of patients in each subgroup, the magnitude of the treatment effect, the testing method used,1 and the correlation between patient responses (i.e., the extent to which a patient’s response on one endpoint predicted their response on a second endpoint). The simulations made it possible to assess the power of the stratum-specific approach compared to the traditional trial designs. The assessment observed power as a function of population size while controlling for the chance of incorrectly concluding that a treatment works when it actually has no effect.

Figure 2

Figure 2. A permutation informs  the likelihood of obtaining an result observed in a trial if there were only random noise and no real treatment difference.

Figure 2. A permutation informs  the likelihood of obtaining an result observed in a trial if there were only random noise and no real treatment difference. To do this, researchers could list every possible way to assign the patients into the treatment groups. For example, if there were a total of six patients in a trial, there would be 20 unique ways to create control and treatment groups of three patients each. Based on the observed trial data, calculating the differences in the test statistic across the two groups for each of these possible assignments would inform on the likelihood of a given result, or one more extreme, if there were no real treatment effect, only random differences. In large trials, where exact calculation for every possible assignment is prohibitive, researchers can select a large, random sample to approximate the distribution of possible outcomes and, thus, the likelihood for obtaining a given value, or one more extreme, through computer simulation as shown.


Researchers observed substantial increases in power for a given subpopulation size using the novel design relative to both traditional trials (Figure 2)—and thus, fewer total patients would be needed to achieve the same power for the stratum-specific trial compared with traditional trial 1. Although the total sample size for a given level of power is lower in traditional trial 2 (because there is only one subpopulation), recruiting such a homogeneous population when the total number of patients is limited may be difficult. Gains in power from using the stratum-specific design depend on the treatment effect size for each endpoint, and the power diminished when the correlation between two endpoints in a patient increased (i.e., when higher treatment responses for one endpoint were associated with higher responses for the other endpoint). In the simulations, when statisticians used the Hochberg method2 for controlling for multiple testing after a global test, they observed the maintenance of reasonable statistical power to draw further meaningful conclusions about a treatment’s effect on specific symptoms, although with lower statistical certainty.

CDER statisticians suggest that their proposed stratum-specific trial design may be valuable to rare disease researchers as they plan clinical studies and that sponsors can construct many other trial designs based on the same paradigm by varying the number of subpopulations and endpoints. Global tests can be applied after adjustment for patient characteristics and are not limited to a particular type of endpoints. The increased statistical power for the proposed global test means a higher chance of the trial providing evidence to support a regulatory approval decision for truly effective products.

The new stratum-specific design has important limitations. The interpretation of a significant global test result is weaker than in traditional trial 1, and the interpretation of a global test is not the same as in traditional statistical tests. Further, in their research, CDER statisticians assumed well-defined subpopulations and accurate categorization of patients into subpopulations based on the trial’s baseline timepoint (Figure 3). The design does not address misclassification and post-randomization changes, such as a patient developing additional symptoms during the trial, though CDER statisticians continue to explore these topics.

CDER researchers encourage sponsors to learn more about the stratum-specific design. For more information about drug development for rare diseases, drug developers should refer to the Learning and Education to ADvance and Empower Rare Disease Drug Developers (LEADER 3D) initiative. Additionally, sponsors can refer to the Complex Innovative Trial Design Meeting Program for an additional avenue to discuss clinical trial designs.

Figure 3

Figure 3. Relationship of power to subpopulation size for the stratum-specific design trial compared with two traditional trials.

Figure 3. Relationship of power to subpopulation size for the stratum-specific design trial compared with two traditional trials.


Footnotes

1. CDER researchers examined the percentile mean method in which the relative magnitude of treatment response is represented using a percentile (or for patients in the subgroup with two endpoints, the mean percentile of combined responses to both endpoints); the standardized score method, which reflects how far an individual’s response (or mean response across both endpoints) is from the corresponding subgroup mean: and methods that select for each patient the maximum percentile or standardized score when there are two or more endpoints. In an additional method for assessing treatment response, t-tests comparing treatment groups were conducted for each subpopulation and corresponding endpoint(s) and the t-test statistics summed before testing. In each case, the statistical significance of differences in treatment groups was tested using a permutation test. See Shives, Emily, et al. "Novel Clinical Trial Design with Stratum‐Specific Endpoints and Global Test Methods for Rare Diseases with Heterogeneous Clinical Manifestations." Statistics in Medicine 44.18-19 (2025). 

2. See page 19 in the FDA guidance for industry Multiple Endpoints in Clinical Trials (October 2022).


References

  1. Ristl, S. Urach, G. Rosenkranz, and M. Posch, “Methods for the Analysis of Multiple Endpoints in Small Populations: A Review,” Journal of Biopharmaceutical Statistics 29, no. 1 (2019): 1–29.
  2. FDA Center for Drug Evaluation and Research, “Application Number 761278Orig1s000 Integrated Review,” (2023) 
  3. Tandon and E. Kakkis, “The Multi‐Domain Responder Index: A Novel Analysis Tool to Capture a Broader Assessment of Clinical Benefit in Heterogeneous Complex Rare Diseases,” Orphanet Journal of Rare Diseases 16, no. 1 (2021): 183.
  4. C. O'Brien, “Procedures for Comparing Samples With Multiple Endpoints,” Biometrics 40, no. 4 (1984): 1079–1087.
  5. FDA guidance for industry, “Multiple Endpoints in Clinical Trials,” (October 2022).
  6. ClinicalTrials.gov, “Phase 3 Study to Evaluate Intravenous Trappsol(R) Cyclo(TM) in Pediatric and Adult Patients With Niemann‐Pick Disease Type C1 (TransportNPC),” (2023).
  7. ClinicalTrials.gov, “Arimoclomol Prospective Study in Participants Diagnosed With Niemann‐Pick Disease Type C,” (2023).
  8. Wang, “Global Tests for Multiple Endpoints in Rare Disease Clinical Trials,” (2021), New England Rare Disease Statistics Workshop presentation slides.
  9. Logan and A. Tamhane, On O'Brien's OLS and GLS Tests for Multiple Endpoints. Lecture Notes‐Monograph Series, vol. 47, Institute of Mathematical Statistics, (2004).
  10. Li, C. M. McDonald, G. L. Elfring, et al., “Assessment of Treatment Effect With Multiple Outcomes in 2 Clinical Trials of Patients With Duchenne Muscular Dystrophy,” JAMA Network Open 3, no. 2 (2020): e1921306.
  11. Zhao, Q. Yu, and S. L. Lake, “A Flexible Multi‐Domain Test With Adaptive Weights and Its Application to Clinical Trials,” Pharmaceutical Statistics 19, no. 3 (2020): 315–325.
  12. Hochberg, “A Sharper Bonferroni Procedure for Multiple Tests of Significance,” Biometrika 75, no. 4 (1988): 800–802.
  13. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. (Routledge, 1988).
  14. Wang, “Statistical Considerations in Rare Disease Clinical Trials,” (2022), CDER-NCATS Workshop Slides, Slides 184- 209.
  15. P. Morris, I. R. White, and M. J. Crowther, “Using Simulation Studies to Evaluate Statistical Methods,” Statistics in Medicine 38, no. 11 (2019): 2074–2102.
  16. FDA guidance for industry, “E9(R1) Statistical Principles for Clinical Trials: Addendum: Estimands and Sensitivity Analysis in Clinical Trials,” (2021).
  17. Shives, Y. Gurmu, W. Lee, W. Morris, and Y. Wang, “Novel Clinical Trial Design With Stratum-Specific Endpoints and Global Test Methods for Rare Diseases With Heterogeneous Clinical Manifestations,” Statistics in Medicine 44, no. 18-19 (2025): e70206.
Back to Top