Abstract
Background: Alzheimer's disease clinical trials have had an excessively high failure rate over the past 15 years, with 99.6% of drugs failing in 2002–2012 (Cummings 2014). Although many of these drugs were likely inactive, a 2-sided 0.05 alpha should result in a 2.5% success rate on the primary endpoint by chance alone. Other addressable factors such as patient heterogeneity, variable outcomes and variable measurement processes contribute to this particularly high failure rate. Methods: Public results from several failed clinical trials were evaluated to assess potential contributors to failure. The loss of power associated with each reason for failure and the potential impact of a solution for recovering that power were estimated. The risk and benefit of these solutions were also evaluated to determine whether they can be practically implemented. Results: Many failed studies showed no evidence of a treatment benefit; however, several studies illustrated known problems in AD research that could be addressed. Some failed studies blame population heterogeneity for failure, and target a more homogeneous population in future studies (e.g.: tramiprosate). Phase 2 studies often use pivotal-study standards for success, including co-primary endpoints, endpoints for mild/moderate disease in mild-only populations, models that don't account for patient heterogeneity and fitting separate visit means rather than a linear time model. Many studies are underpowered, resulting in equivocal results when effects are smaller than expected. Elementary issues, such as equivalence of forms for psychometric tests are often assumed, and not checked. For example, a glance at the Expedition 1, 2, and 3 ADAS-cog figures reveals an abnormally high score at Month 9 for all 3 studies, likely due to an easier wordlist at that visit. Conclusions: Successful Alzheimer's clinical trials require an active compound and successful study design demonstrated by narrow confidence intervals. Phase 2 studies should use different standards for success than phase 3 studies, and statistical and psychometric issues should be fully considered. Accurately identifying compounds with small effect sizes is the first step toward developing better treatments with larger effect sizes. Narrow confidence intervals indicate more precision in treatment effect size estimates leading to failing ineffective treatments and success for effective treatments.