Search NASA⌕ Search

SEARCH · Search NASA

Results for “analysis and statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Searching for Neutrino Tridents in the NOvA Near Detector

This dissertation presents a search for neutrino trident production in the NOvA near detector through the coherent ``dimuon" channel: $\nu_\mu +\hspace{1pt}\text{X} \rightarrow \nu_\mu + \mu^- + \mu^+ +\hspace{1pt}\text{X}$. Trident production is a rare, purely electroweak process with sensitivity to physics beyond the Standard Model. The theoretical background, motivation for studying the process, and previous experimental measurements are reviewed. The analysis uses data collected by the NOvA near detector (ND) from Fermilab's Neutrinos at the Main Injector (NuMI) beam between November 2014 and February 2024, corresponding to an exposure of $25.5\times 10^{20}$ protons on target. The ND is a segmented tracking calorimeter located 800~m from the beam target, receiving neutrinos with a mean energy of 2~GeV. A multi-pass background reduction strategy is implemented, including the development of a novel dimuon-specific tracking technique. Trident candidates are identified using a boost ed decision tree classifier trained on simulated signal and background events. Limited background Monte Carlo statistics necessitate the use of functional fits to sideband data, which are extrapolated to estimate backgrounds in the signal region. The unblinded data contain 9 trident-like events, with an estimated background of 5.66 $\pm$ 5.15 events. This yields a best fit estimate of 3.34 tridents compared to the Standard Model prediction of 4.66. A profiled Feldman-Cousins method is used to determine a 90\% confidence interval of [0,9.1] on the number of signal events, corresponding to an upper limit of 1.95$\times$ the Standard Model prediction. This result represents the lowest energy search for trident events to date, and the first experimental contribution to the process in 27 years.

Bowles, Reed Scott [Indiana U.]↗

Going Off Grid: A Comparative Study of the Lagrangian and Eulerian Perspectives of New Particle Formation Events

New particle formation and growth (NPF&G) is the process by which ultrafine particles are formed from gas-phase precursors. NPF&G is the dominant source of global aerosol number with important influences on climate. Most observations of NPF&G events are conducted at stationary sites; however, NPF&G observed from stationary sites is influenced by gradual or rapid changes in the air masses passing over the site, complicating NPF&G analysis. In this work, we use observations and a 3D aerosol model to compare aerosol size distributions at a stationary site (Southern Great Plains [SGP] observatory, Oklahoma, USA) and along Lagrangian trajectories crossing the site. The model simulates the NPF&G events reasonably well at SGP. Using the model to compare the Lagrangian and stationary perspectives, we can explain previously unanalyzable days with some evidence of NPF&G as either non-event or analyzable NPF&G days. We find most of the unanalyzable NPF&G days are due to isolated and inhomogeneous NPF&G occurring upwind of the stationary site, often in the outflow of urban regions. Finally, we compare formation rates of 3 nm particles, growth rates, and the survival probability of 3 nm particles growing to 25 nm between the stationary and Lagrangian perspectives. Because of the much larger number of analyzable days along the Lagrangian trajectories, this perspective potentially provides more robust statistics and better characterization of NPF&G event extremes. Our method for extracting chemical/physical properties along Lagrangian trajectories from 3D models can be applied to a wide range of science questions.

O’Donnell, Samuel E. [Colorado State Univ., Fort C↗

Random insights into the complexity of two-dimensional tensor network calculations

Projected entangled pair states (PEPS) offer memory-efficient representations of some quantum many-body states that obey an entanglement area law and are the basis for classical simulations of ground states in two-dimensional (2d) condensed matter systems. However, rigorous results show that exactly computing observables from a 2d PEPS state is generically a computationally hard problem. Yet approximation schemes for computing properties of 2d PEPS are regularly used, and empirically seen to succeed, for a large subclass of (“not too entangled”) condensed matter ground states. Adopting the philosophy of random matrix theory, in this work, we analyze the complexity of approximately contracting a 2d random PEPS by exploiting an analytic mapping to an effective replicated statistical mechanics model that permits a controlled analysis at a large bond dimension. Through this statistical-mechanics lens, we argue that (i) although approximately sampling wave-function amplitudes of random PEPS faces a computational-complexity phase transition above a critical bond dimension, and (ii) one can generically efficiently estimate the norm and correlation functions for any finite bond dimension. Furthermore, these results are supported numerically for various bond-dimension regimes. It is an important open question whether the above results for random PEPS apply more generally also to PEPS representing physically relevant ground states.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

First observation of antiproton annihilation at rest on argon in the LArIAT experiment

We report the first observation and measurement of antiproton annihilation at rest on argon track and shower multiplicities and particle identification conducted with the LArIAT experiment. Stopping antiprotons from the Fermilab Test Beam Facility’s charged particle test beam are identified using beamline instrumentation and LArIAT’s liquid argon time projection chamber (LArTPC). The charged particle multiplicity from the annihilation vertex is manually evaluated via hand scanning, yielding a mean of 3.2 ± 0.4 tracks and a standard deviation of 1.3 tracks, consistent with a semiautomated reconstruction resulting in 2.8 ± 0.4 tracks and a standard deviation of 1.2 tracks. Both methods are consistent with Monte Carlo simulations within statistical uncertainty. The shower multiplicities and particle identification for outgoing tracks are also consistent with eant4 model predictions. These results, obtained from a low-statistics sample, provide a foundation for higher-statistics studies in larger LArTPCs, which could refine modeling of intranuclear annihilation on argon and inform scenarios such as neutron-antineutron oscillations. Published by the American Physical Society 2025

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Cross sections for the formation of Rb84m,g, Rb83, and Rb82m in Sr86(d,x) reactions up to deuteron energies of 49 MeV: Competition between α-particle and multinucleon emission processes

Cross sections of Sr86(d,x) reactions leading to the products Rb84m,g, Rb83, and Rb82m were measured by the stacked-sample activation technique up to deuteron energies of 49 MeV. Nuclear model calculations were performed using the codes talys and empire, which combine the statistical, precompound, and direct interaction components. In all cases, the empire results were much higher than the talys calculation. Fairly good agreement was obtained between measured data and the talys calculation after some optimization of the input model parameters. Insight into competition between α-particle and multinucleon emission in the Y88 compound-nucleus system was also gained.

59 ≤ A ≤ 89↗

Enhanced accuracy through ensembling of randomly initialized auto-regressive models for dynamical systems

Computational mechanics simulations using traditional finite element methods (FEM) require prohibitively expensive computational resources for real-time engineering applications, design optimization, and digital twin implementations. While machine learning (ML) surrogate models offer significant computational speedups, autoregressive ML models for time-dependent mechanical systems suffer from error accumulation that compromises long-term prediction reliability - a critical concern for engineering applications where accuracy over extended time horizons is essential for safety and performance assessments. Here, we propose a deep ensemble framework specifically designed to address this challenge in computational mechanics applications, where multiple ML surrogate models with random weight initializations are trained in parallel and their predictions aggregated during inference. This approach leverages statistical diversity to maximize information gain from a fixed set of training data and to mitigate error propagation, while maintaining the computational efficiency that makes ML surrogates attractive for engineering practice. We validate the framework on three representative problems spanning critical areas of computational mechanics: stress field evolution in heterogeneous microstructures under complex loading (relevant to advanced materials design and composite analysis), planetary-scale shallow water dynamics (applicable to environmental and geotechnical engineering), and Gray-Scott reaction-diffusion systems (relevant to mass transport and chemical process engineering). Across all test cases, the ensemble approach demonstrates consistent error reduction of 15-33% compared to individual models. The codes for this work are available on GitHub (https://github.com/Graham-Brady-Research-Group/AutoregressiveEnsemble_SpatioTemporal_Evolution).

autoregressive prediction↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Creating high-precision reference gas standards of 85Kr for groundwater age-dating

Absolute gas counting (AGC) was applied to two gas blends of 85Kr in argon-methane (P10) counting gas to establish a high-precision specific activity (Bq/cm3) reference value for characterizing 85Kr detection efficiency for groundwater age dating measurements. The AGC or length-compensated technique has been utilized by the metrology community for decades and is an accepted method for developing radioactive gas standards. The AGC capability at Pacific Northwest National Laboratory (PNNL) uses a set of nine unequal-length proportional counters with precisely-measured internal volumes, and a gas loading system with high-precision pressure and temperature sensors. A series of AGC measurements were collected at multiple pressures to determine the inverse pressure relationship (1/P) for 85Kr and define a wall-effect correction that accounts for events decaying into the detector wall and not depositing sufficient energy in the gas to be detected. In addition to the wall-effect, two additional corrections were evaluated and are discussed in detail. Specifically, the threshold effect which accounts for events deposited below the analysis threshold and a detection efficiency as a function of detector volume effect that was observed during analysis. A robust uncertainty model was developed using the Guide to the expression of Uncertainty in Measurements (GUM) approach. The combination of carefully scrutinized correction factors, precise measurements of pressure, temperature and detector volume, and robust counting statistics resulted in the determination of high-precision specific activity values with 0.50% or less total combined uncertainty for two Kr-in-P10 reference gas standards (KP10) that will enable new groundwater age-dating measurements at PNNL.

Kr-85↗

Real-Time event reconstruction for Nuclear Physics Experiments using Artificial Intelligence

Charged track reconstruction is a critical task in nuclear physics experiments, enabling the identification and analysis of particles produced in high-energy collisions. Machine learning (ML) has emerged as a powerful tool for this purpose, addressing the challenges posed by complex detector geometries, high event multiplicities, and noisy data. Traditional methods rely on pattern recognition algorithms like the Kalman filter, but ML techniques, such as neural networks, graph neural networks (GNNs), and recurrent neural networks (RNNs), offer improved accuracy and scalability. By learning from simulated and real detector data, ML models can identify and classify tracks, predict trajectories, and handle ambiguities caused by overlapping or missing hits. Moreover, ML-based approaches can process data in near-real-time, enhancing the efficiency of experiments at large-scale facilities like the Large Hadron Collider (LHC) and Jefferson Lab (JLAB). As detector technologies and computational resources evolve, ML-driven charged track reconstruction continues to push the boundaries of precision and discovery in nuclear physics. In these proceedings, we highlight advancements in charged track identification leveraging Artificial Intelligence within the CLAS12 detector, achieving a notable enhancement in experimental statistics compared to traditional methods. Additionally, we showcase real-time event reconstruction capabilities, including the inference of charged particle properties, such as momentum, direction, and species identification, at speeds matching data acquisition rates. These innovations enable the extraction of physics observables directly from the experiment in real-time.

Gavalian, Gagik (ORCID:0000000267385457)↗

Genesis Mission-Enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE)

Argonne National Laboratory is supporting the U.S. Department of Transportation’s (USDOT’s) Bureau of Transportation Statistics (BTS) with collaborative research on development and application of privacy preserving AI frameworks that leverage unmatched AI expertise and secure computing resources made available through the U.S. Genesis Mission1 . This research advances U.S. energy security goals by supporting a safe offshore energy industry with secure, domain-specific AI tools to analyze confidential industry datasets collected by BTS to rapidly improve identification of hazards, precursors, and systemic safety risks in high-risk operational environments. The staged, security-first approach begins with development and testing of Argonne’s Genesis Mission-enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE) framework within Argonne’s accredited secure computing enclave (ABLE) leveraging Argonne’s AI scientific assistant substrate (AISAC). Methods to build synthetic datasets were developed together with BTS for use in preparing synthetic datasets that can be used to validate data containment, governance, and security controls in the ABLE environment. Future research directions would focus on applying the Genesis-SAFE framework to CIPSEA-protected datasets entirely within ABLE to support confidentiality-preserving analysis of safety risks, trends, and contributing factors.

Kim, Hyekyung [Argonne National Laboratory (ANL), ↗

Full Modeling and parameter compression methods in configuration space for DESI 2024 and beyond

In the contemporary era of high-precision spectroscopic surveys, led by projects like DESI, there is an increasing demand for optimizing the extraction of cosmological information from clustering data. This work conducts a thorough comparison of various methodologies for modeling the full shape of the two-point statistics in configuration space. We investigate the performance of both direct fits (Full Modeling) and the parameter compression approaches (ShapeFit and Standard). We utilize the ABACUS-SUMMIT simulations, tailored to exceed DESI's precision requirements. Particularly, we fit the two-point statistics of three distinct tracers (LRG, ELG, and QSO), by employing a Gaussian Streaming Model in tandem with Convolution Lagrangian Perturbation Theory and Effective Field Theory. We explore methodological setup variations, including the range of scales, the set of galaxy bias parameters, the inclusion of the hexadecapole, as well as model extensions encompassing varying ns and allowing for w 0 w a CDM dark energy model. Throughout these varied explorations, while precision levels fluctuate and certain configurations exhibit tighter parameter constraints, our pipeline consistently recovers the parameter values of the mocks within 1σ in all cases for a 1-year DESI volume. Additionally, we compare the performance of configuration space analysis with its Fourier space counterpart using three models: PyBird, FOLPS and velocileptors, presented in companion papers. We find good agreement with the results from all these models.

79 ASTRONOMY AND ASTROPHYSICS↗

Part distortion monitoring in additive manufacturing using machining

In additive manufacturing, accumulation of residual stresses can result in severe part distortion from the desired preform shape. Current methods for in-situ part distortion monitoring in additive manufacturing typically require expensive sensors, or capital equipment, and require time-consuming post-processing to understand the shape deviation. This paper presents an in-situ method, in the context of hybrid manufacturing, for part distortion detection using machining of additively manufactured parts. As a surrogate, three test artifacts were used to represent different distorted geometries. The tool axis positions from the machine tool controller and the cutting power were monitored during a facing operation. Cutting power data was used to detect the tool entry and exit in the workpiece using a novel approach with power standard deviation metric. The workpiece geometry and distorted configuration was subsequently predicted for positional and rotational deviations to within 2 mm accuracy using synchronized tool position data with cutting power. The proposed method can be used in a hybrid (additive and subtractive) machine tool to periodically check part distortion in the additive build. The method is applicable for any additive process and is low-cost and computationally inexpensive.

36 MATERIALS SCIENCE↗

Statistical inference of collision frequencies from x-ray Thomson scattering spectra

Thomson scattering spectra measure the response of plasma particles to incident radiation. In warm dense matter, which is opaque to visible light, x-ray Thomson scattering (XRTS) enables a detailed probe of the electron distribution and has been used as a diagnostic for electron temperature, density, and plasma ionization. In this work, we examine the sensitivities of inelastic XRTS signatures to modeling details, including the dynamic collision frequency and the electronic density of states. Applying verified Monte Carlo inversion methods to dynamic structure factors obtained from time-dependent density functional theory, we assess the utility of XRTS signals as a way to inform the dynamic collision frequency, especially its direct-current limit, which is directly related to the electrical conductivity.

Collision frequency↗

Evaluating User Errors and Temporal Trends in Marine Fish Communities Using 360-Degree Underwater Photography

The use of environmental DNA (eDNA) sampling has been proposed as a complementary method to monitor fish species in marine environments, offering a non-invasive and potentially more efficient approach to marine species observations. eDNA monitoring could be especially useful in and around sites targeted for marine energy generation as these regions need regular monitoring that would be impractical with traditional techniques. Before we can fully rely upon eDNA, we must first verify its accuracy against other proven methods, such as the use of underwater photography. In this study, I deployed a 360-degree camera in the tidal channel of Sequim Bay once a month during several hours overlapping slack tide. I investigated how having multiple people identify and count fish on underwater images could affect the overall results. Using chi square tests in R, I compared my fish identifications and counts to those made by another intern on the same images recorded in August. I found significant differences in the number of species identified and the total individual counts between the two different datasets. I also tested the statistical differences in both Shannon diversity and Pielou evenness indices between the August, September, and November camera deployments using a Hutcheson t-test. Only one significant difference was found in the Shannon index comparisons, and none were found between the Pielou evenness comparisons. These findings show that if multiple identifiers are used to process underwater images, quality control checks must be made to reduce the potential for error. This also points toward the possibility to leverage more advanced image analysis processes, such as automated image analysis software. The findings from this study also show that the dynamics of marine fish communities can vary over a few months; however, further analysis is needed to determine the extent of the seasonal changes in Sequim Bay.

59 BASIC BIOLOGICAL SCIENCES↗

Exploring HOD-dependent systematics for the DESI 2024 Full-Shape galaxy clustering analysis

We analyze the robustness of the DESI 2024 cosmological inference from the full shape of the galaxy power spectrum to uncertainties in the Halo Occupation Distribution (HOD) model of the galaxy-halo connection and the choice of priors on nuisance parameters. We assess variations in the recovered cosmological parameters across a range of mocks populated with different HOD models and find that shifts are often greater than 20% of the expected statistical uncertainties from the DESI data. We encapsulate the effect of such shifts in terms of a systematic covariance term, C HOD , and an additional diagonal contribution quantifying the impact of our choice of nuisance parameter priors on the ability of the effective field theory (EFT) model to correctly recover the cosmological parameters of the simulations. These two covariance contributions are designed to be added to the usual covariance term, C stat , describing the statistical uncertainty in the power spectrum measurement, in order to fairly represent these sources of systematic uncertainty. This novel approach should be more general and robust to the choice of model or additional external datasets used in cosmological fits than the alternative approach of adding systematic uncertainties to the recovered marginalised parameter posteriors. We compare the approaches within the context of a fixed ΛCDM model and demonstrate that our method gives conservative estimates of the systematic uncertainty that nevertheless have little impact on the final posteriors obtained from DESI data.

79 ASTRONOMY AND ASTROPHYSICS↗

Impact of survey spatial variability on galaxy redshift distributions and the cosmological 3 × 2-point statistics for the Rubin Legacy Survey of Space and Time (LSST)

We investigate the impact of spatial survey non-uniformity on the galaxy redshift distributions for forthcoming data releases of the Rubin Observatory Legacy Survey of Space and Time (LSST). Specifically, we construct a mock photometry data set degraded by the Rubin OpSim observing conditions, and estimate photometric redshifts of the sample using a template-fitting photo-z estimator, BPZ, and a machine learning method, FlexZBoost. We select the Gold sample, defined as $i\lt 25.3$ for 10 yr LSST data, with an adjusted magnitude cut for each year and divide it into five tomographic redshift bins for the weak lensing lens and source samples. We quantify the change in the number of objects, mean redshift, and width of each tomographic bin as a function of the coadd i-band depth for 1-yr (Y1), 3-yr (Y3), and 5-yr (Y5) data. In particular, Y3 and Y5 have large non-uniformity due to the rolling cadence of LSST, hence provide a worst-case scenario of the impact from non-uniformity. We find that these quantities typically increase with depth, and the variation can be $10\!-\!40~{{\rm per\ cent}}$ at extreme depth values. Using Y3 as an example, we propagate the variable depth effect to the weak lensing $3\times 2$ pt analysis, and assess the impact on cosmological parameters via a Fisher forecast. We find that galaxy clustering is most susceptible to variable depth, and non-uniformity needs to be mitigated below 3 per cent to recover unbiased cosmological constraints. There is little impact on galaxy–shear and shear–shear power spectra, given the expected LSST Y3 noise.

cosmology↗

Site-specific Design Case Study for Wet Waste Hydrothermal Liquefaction and Biocrude Upgrading to Hydrocarbon Fuels

Hydrothermal liquefaction (HTL) is a thermal process that converts wet biomass to renewable hydrocarbon fuel blendstocks (i.e., renewable naphtha, renewable diesel, and sustainable aviation fuel (SAF)). It can utilize a wide range of pure and blended wet feedstocks, including sewage sludge from water resource recovery facilities (WRRF), food and agriculture wastes, algae, fats, oils and greases (FOG) and blends of dry and wet wastes/feedstocks. Historically, techno-economic analysis (TEA) and annual state of technology (SOT) assessments with standard economic assumptions used by the Bioenergy Technologies Office (BETO) were conducted for the wet waste HTL pathway leveraging experimental data collected from Pacific Northwest National Laboratory’s (PNNL) continuous flow reactor systems. The objective of the SOT assessment has been to guide and track progress of BETO’s HTL research and development (R&D) toward reduced cost and greenhouse gas (GHG) emissions for the pathway. However, gaps exist between BETO’s traditional SOT updates and the needs of key external stakeholders that – if addressed – will accelerate technology adoption. This Business Case Study aims to bridge this gap by providing an updated design, TEA, and LCA based on PNNL’s FY23 R&D with added analyses and information that provide enhanced relevance for stakeholders of the HTL technology. This includes specific siting, regional wet waste resource inventory and transportation cost analyses, fuel market information, sustainable fuel policy impacts, economic metrics of net present value (NPV) and internal rate of return (IRR), greenhouse gas (GHG) emissions analysis, and statistical analysis of cost and technical uncertainties of the HTL plant design. The study focuses on the “Detroit combined statistical area (CSA)” region for siting of a wet waste HTL plant adjacent to the Great Lakes Water Authority (GLWA) facility with guidance from industry participants. Regional resource and siting analyses were conducted to identify feedstock availability, scale, and cost, as well as a beneficial site location. TEA with detailed rigorous capital cost estimation for the specific site application was conducted to evaluate the key economic metrics of most value to industrial partners. These include total capital investment, operating costs, minimum fuel selling price (MFSP) of the biocrude and fuel blendstock, and NPV and internal rate of return IRR with sustainable fuel credits. Life cycle analysis was conducted to evaluate the supply chain greenhouse gas (GHG) emissions for the wet waste HTL process as compared with petroleum derived diesel. This study is also informed by years of R&D and process de-risking learnings and was conducted with a basic engineering HTL plant design and costing that akin to a “first-of-a-kind” plant economics. This differs from our conventional “nth plant ” SOT assessments. Specifically, the HTL process model has been updated with more operationally reliable methods for feed heating and phase separations. Further, we have implemented additional spare equipment for redundancy, a more rigorous installed equipment cost estimation approach, and additional costs associated with feed formatting and delivery, building, piping and site development. An Excel-based cost sheet based on the basic engineering design is also released alongside the report that allows users to conduct customized TEA with their own feed composition and financial assumptions.

09 BIOMASS FUELS↗

Pivotal trial characteristics and types of endpoints used to support Food and Drug Administration rare disease drug approvals between 2013 and 2022

Background/aims Rare disease drug development faces unique challenges, such as genotypic and phenotypic heterogeneity within small patient populations and a lack of established outcome measures for conditions without previously successful drug development programs. These challenges complicate the process of selecting the appropriate trial endpoints and conducting clinical trials in rare diseases. In this descriptive study, we examined novel drug approvals for non-oncologic rare diseases by the U.S. Food and Drug Administration’s Center for Drug Evaluation and Research over the past decade and characterized key regulatory and trial design elements with a focus on the primary efficacy endpoint utilized as the basis of approval. Methods Using the Food and Drug Administration’s Data Analysis Search Host database, we identified novel new drug applications and biologics license applications with orphan drug designation that were approved between 2013 and 2022 for non-oncologic indications. From Food and Drug Administration review documents and other external databases, we examined characteristics of pivotal trials for the included drugs, such as therapeutic area, trial design, and type of primary efficacy endpoints. Differences in trial design elements associated with primary efficacy endpoint type were assessed such as randomization and blinding. Then, we summarized the primary efficacy endpoint types utilized in pivotal trials by therapeutic area, approval pathway, and whether the disease etiology is well defined. Results One hundred and seven drugs that met our inclusion criteria were approved between 2013 and 2022. Assessment of the 107 drug development programs identified 150 pivotal trials that were subsequently analyzed. The pivotal trials were mostly randomized (80%) and blinded (69.3%). Biomarkers (41.1%) and clinical outcomes (42.1%) were commonly utilized as primary efficacy endpoints. Analysis of the use of clinical trial design elements across trials that utilized biomarkers, clinical outcomes, or composite endpoints did not reveal statistically significant differences. The choice of primary efficacy endpoint varied by the drug’s therapeutic area, approval pathway, and whether the indicated disease etiology was well defined. For example, biomarkers were commonly selected as primary efficacy endpoints in hematology drug approvals (70.6%), whereas clinical outcomes were commonly selected in neurology drug approvals (69.6%). Further, if the disease etiology was well defined, biomarkers were more commonly used as primary efficacy endpoints in pivotal trials (44.7%) than if the disease etiology was not well defined (27.3%). Discussion In the past 10 years, numerous novel drugs have been approved to treat non-oncologic rare diseases in various therapeutic areas. To demonstrate their efficacy for regulatory approval, biomarkers and clinical outcomes were commonly utilized as primary efficacy endpoints. Biomarkers were not only frequently used as surrogate efficacy endpoints in accelerated approvals, but also in traditionally approved rare disease drugs. The choice of primary efficacy endpoints varied by therapeutic area, approval pathway, and understanding of disease etiology.

Hong, Kyungwan [Rare Diseases Team, Office of New ↗