Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

Constraining Present‐Day Anthropogenic Total Iron Emissions Using Model and Observations

Abstract Iron emissions from human activities, such as oil combustion and smelting, affect the Earth's climate and marine ecosystems. These emissions are difficult to quantify accurately due to a lack of observations, particularly in remote ocean regions. In this study, we used long‐term, near‐source observations in areas with a dominance of anthropogenic iron emissions in various parts of the world to better estimate the total amount of anthropogenic iron emissions. We also used a statistical source apportionment method to identify the anthropogenic components and their sub‐sources from bulk aerosol observations in the United States. We find that the estimates of anthropogenic iron emissions are within a factor of 3 in most regions compared to previous inventory estimates. Under‐ or overestimation varied by region and depended on the number of sites, interannual variability, and the statistical filter choice. Smelting‐related iron emissions are overestimated by a factor of 1.5 in East Asia compared to previous estimates. More long‐term iron observations and the consideration of the influence of dust and wildfires could help reduce the uncertainty in anthropogenic iron emissions estimates.

Meteorology & Atmospheric Sciences↗

Entropy-based feature selection for capturing impacts in Earth system models with abrupt forcing

This paper presents the development of a new entropy-based feature selection method for identifying and quantifying impacts. Here, impacts are defined as statistically significant differences in spatio-temporal fields when comparing datasets with and without an external forcing in an Earth system model. Temporal feature selection is performed by first computing the cross-fuzzy entropy to quantify similarity of patterns between two datasets and then applying changepoint detection to identify regions of statistically constant entropy. The method is used to capture temperate north surface cooling from a 9-member simulation ensemble of the Mt. Pinatubo volcanic eruption, which injected 10 Tg of SO 2 into the stratosphere. The results estimate a mean difference decrease in near surface air temperature of -0.560 K with a 99% confidence interval between -0.864 K and -0.257 K between April and November of 1992, one year following the eruption. A sensitivity analysis with decreasing SO 2 injection revealed that the impact is statistically significant at 5 Tg but not at 3 Tg. Using identified features, a dependency graph model based on a 9-day lag had significantly fewer nodes than a graph based on monthly means. Furthermore, this demonstrates our method’s ability to perform dimension reduction while still uncovering source-to-impact pathways.

Changepoint detection↗

Expanded method for the determination of burnup in nuclear fuels using multiple neodymium isotopes

Established methods for the determination of burnup in nuclear fuels commonly rely on the measurement of 148Nd in the spent fuel in concert with known fission product yields to determine the atom percent of fissions in the fuel. This isotope of neodymium is used for various reasons, including chemical and radioactive stability, ease of measurement, and low rates of formation and destruction due to neutron flux apart from fission. However, careful calculation of effective cumulative fission yields and correction factors for (n,γ) capture reactions allows for additional stable and long-lived isotopes of neodymium to be used to provide additional independent measurements of burnup, reducing statistical uncertainty. This method was developed and successfully applied to measure the burnup of compacts from the Advanced Gas Reactor (AGR) Fuel Development and Qualification Program. The mean burnup measured using 143Nd, 145Nd, 146Nd, 148Nd, and 150Nd was statistically observed to be the same as that measured using 148Nd alone, but the statistical uncertainty in the measurement was reduced by a factor of 2, providing a tighter confidence interval in the final results.

Helmreich, Grant [ORNL] (ORCID:0000000330464394)↗

End-Use Savings Shapes Measure Documentation: Dispatch Schedule Generation for Demand Flexibility Measures

This supplemental document describes the methodology used for determining the dispatch timing of various EUSS demand flexibility measures. Demand flexibility measures are designed to reduce/dispatch electricity demand in buildings during especially beneficial/critical times. The method used in this work utilizes predictions of building loads to generate a schedule that reflects the periods when the building's daily peak load occurs to support decision making in demand flexibility measures. The dispatch schedule generation method described in this document creates an hourly schedule that includes a load dispatch (peak) window for each day for a whole year based on load prediction, with options using different prediction methods: perfect prediction, bin-sampling method, fixed schedule, and outdoor air temperature (OAT)-based prediction method. The perfect prediction method performs a simulation to obtain the annual load profile as predicted load, representing the scenario of perfect load prediction. The bin-sampling method (1) categorizes days into representative bins by temperature characteristics, (2) performs simulations on sample days from each of those bins to create representative (or predicted) load, and (3) assigns representative loads for all days in a year based on the bin categorization. The fixed schedule method defines uniform start and end time of peak window with assumed fixed daily peak time, for all days in a season or a year. The OAT-based prediction method uses the statistics of OAT (minimum and maximum) as the indicators of peak load, with specified delay response time from building loads to temperature. Given the load prediction, daily peak periods are determined as a time window with specified length in each day that include the predicted daily peak load and with a secondary rule such as maximizing energy saving potential. The dispatch schedule generation method is not a standalone measure and is intended to be combined with other demand flexibility measures that could leverage the peak schedule and apply demand controls on specific systems or devices for demand response, such as measures described in "Measure Documentation - Thermostat Control for Load Shedding" and "Measure Documentation - Thermostat Control for Load Shifting".

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Development of Solar Flare and Energetic Particle Prediction Portal (SEP 3 )

Solar activity is a primary factor determining the state of the Earth’s space environment, geomagnetic and ionospheric disturbances, and radiation hazards. In the current state of knowledge, machine learning (ML) methods provide essential tools for processing data, investigating relationships among various physical properties and characteristics, uncovering hidden connections, and predicting hazardous solar events. The primary difficulty in developing and applying modern machine-learning tools in heliophysics is that the essential data are scattered among over a hundred data repositories developed by instrument teams of space missions and ground-based observatories. In addition, statistical and ML methods require long time series of homogeneous measurements. To facilitate ML-ready data preparation and access, we have developed an interactive database of solar flares integrating the most essential datasets (https://solarflare.njit.edu/). The database performs an initial data processing and is automatically updated. In addition, we are developing the Solar Energetic Particle Prediction Portal (SEP3, https://sun.njit.edu/SEP3), which hosts web applications that allow users to retrieve the database records. The Portal has a search page for browsing the events from the most widely used catalogs and a dedicated space to share the most recent achievements of the team. The interactive widget can display soft X-ray and proton flux time series from GOES satellites and the flare records. The data portal has been used to evaluate the forecasts of solar proton events and investigate machine-learning approaches to SEP prediction.

SMD↗

Dynamic response and input identification of MDOF structures subjected to coupled random vector inputs

A method of random dynamic analysis is presented which is based on the classical approach applied to a discretized structure. The method assumes that the system identification is available in the form of natural modes and frequencies. These modes and frequencies can be found from either available solutions or approximately from finite element programs. A computer program has been developed to perform the computations required in the analysis. All computations are performed with transformed modal variables, which results in significant economy since the number of modal degrees of freedom is almost always less than the number of physical degrees of freedom. The program computes the response to force and base inputs which are statistically coupled. A method is also presented for predicting one of the inputs if the second input and the acceleration response at a point on the structure are known. Finally, results are presented for the random response of a rectangular plate subjected to a random pressure and a random base input. The inputs are considered individually and with various degrees of statistical coupling.

Ocallahan, J. C.↗

A statistical model of the human core-temperature circadian rhythm

We formulate a statistical model of the human core-temperature circadian rhythm in which the circadian signal is modeled as a van der Pol oscillator, the thermoregulatory response is represented as a first-order autoregressive process, and the evoked effect of activity is modeled with a function specific for each circadian protocol. The new model directly links differential equation-based simulation models and harmonic regression analysis methods and permits statistical analysis of both static and dynamical properties of the circadian pacemaker from experimental data. We estimate the model parameters by using numerically efficient maximum likelihood algorithms and analyze human core-temperature data from forced desynchrony, free-run, and constant-routine protocols. By representing explicitly the dynamical effects of ambient light input to the human circadian pacemaker, the new model can estimate with high precision the correct intrinsic period of this oscillator ( approximately 24 h) from both free-run and forced desynchrony studies. Although the van der Pol model approximates well the dynamical features of the circadian pacemaker, the optimal dynamical model of the human biological clock may have a harmonic structure different from that of the van der Pol oscillator.

NASA Discipline Regulatory Physiology↗

Statistical fluctuations in Monte Carlo calculations

The time counter and modified Nanbu simulation techniques are analyzed, with emphasis placed on the convergence of the calculations to a steady macroscopic state. Such variables as translational and rotational temperature, and flow velocity, sampled at several points in the flowfield, are considered. Both macroscopic averages and molecular distribution functions are analyzed. The calculation of inelastic collisions, in which transfer of energy between translational and internal energy modes is performed, is achieved through the use of the Larsen-Borgnakke phenomenological model. It is noted that, with reference to translational temperature, the time counter method shows less statistical scatter than that found with the modified Nanbu simulation technique.

Boyd, I. D.↗

Truncated ARQ Statistical Link Analysis for Dynamic Links

The future deep space links are migrating towards higher frequency bands such as Ka band and optical. These links are susceptible to non Gaussian and non linear effects such as atmospheric turbulence, scintillation, antenna mis-pointing, jitter, etc. These dynamic links thus will experience various degrees of fading loss, and some of these link disruptions cannot be effectively mitigated by forward error correction coding and/or interleaving. One effective way to ensure reliable communication is by using Automatic Repeat Request (ARQ) protocol, where the receiver acknowledges to the transmitter whether or not a data unit is successfully received. If a data unit is not successfully received (such as after a pre-set time-out), the transmitter would then re-transmit the lost data unit to the receiver. In a previous paper, we derived a statistical link analysis method of finding the optimal operating Signal-to-Noise Ratio (SNR) and estimating the latency of an ARQ scheme. In a more recent paper, we demonstrated the above method using the SNR distribution constructed from the Ka-band (32 GHz) flight data. To simplify the discussion, we considered the academic approach that the ARQ scheme allows for an infinite number of retransmissions. In this paper, we consider the more practical case of a truncated ARQ scheme, where there is a limit on the number of retransmissions. We derive the error probability, the optimal SNR setting, and the latency statistics of the correctly received frames of the truncated ARQ schemes. We first discuss the truncated ARQ link analysis principles using the Gaussian assumption for SNR distribution with a large variance. Next, we demonstrate the statistical truncated ARQ link analysis using the SNR distribution constructed from the Ka-band flight data. The results in this paper can be applied in the design of reliable communication systems such as the Consultative Committee for Space Data System (CCSDS) File Transfer Protocol (CFTP) and the Delay Tolerant Network (DTN).

Morabito, David↗

Methods for Probabilistic Uncertainty Analysis and Bayesian Analysis with Examples of Statistically Analyzing Data to Revise MMOD Risk Estimates and Compare Models

Probabilistic methods are presented for characterizing and quantifying uncertainties in models and model predictions. Techniques are given for constructing specific uncertainty distributions based on available information. Alternative techniques are given for propagating uncertainties in model inputs to obtain the uncertainty in the model result or prediction. Bayesian techniques are also described for utilizing data and information to update and revise model results and predictions. The focus is on applications with numerous specific examples given. The use of data to revise Micrometeoroid and Orbital Debris (MMOD) risk prediction models are among the examples given.

Risk↗

Forecast of an exceptionally large even-numbered solar cycle

Using the 'dynamo theory' method to predict solar activity, an accurate prediction was made for solar cycle 21 by Schatten et al. (1978). Using the same dynamo technique for solar cycle 22, a value for the smoothed sunspot number of 170 + or - 25 is obtained. This large sunspot number is expected to peak in 1990 + or - 1 year. The F(10.7) radio flux is expected to reach a smoothed value of 210 + or - 25 flux units. Since this value is larger than values obtained with prediction schemes based upon 'statistical' and 'periodicity' methods, it provides a useful test for the current methodology, based upon the strength of the sun's polar field near solar minimum. The predicted degree of solar activity is expected to enhance the density and temperature of the earth's thermosphere to values somewhat larger than those found in solar cycle 21. This will have an impact on the orbital lifetime of low altitude satellites.

Schatten, Kenneth H.↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Uncertainty analysis of Raman spectra for measuring ortho-parahydrogen compositions

Determining ortho-parahydrogen compositions at cryogenic temperatures is an important quality control for liquid hydrogen custody exchange. High orthohydrogen compositions lead to an exothermic reaction resulting in increased boil off and increased venting losses of liquid hydrogen product by either the supplier or consumer. Traditional methods for measuring ortho-parahydrogen compositions such as hot-wire anemometry, nuclear magnetic resonance, and infrared spectroscopy are typically inconvenient at best. As the number of cryogenic hydrogen systems continue to increase, there is an increased need for flexible, standardized techniques for post processing composition measurements. A Raman spectrometer implemented in the Cryo-Catalysis Hydrogen Experimental Facility (CHEF) has demonstrated incredible flexibility for composition measurements of cryogenic hydrogen flows through a wide range of temperatures and pressures with the Raman probe in-situ at cryogenic temperatures. This article analyzes the uncertainty of Raman spectroscopy data of cryogenic hydrogen equilibrated at temperatures in IONEX catalyst. Spectra from hydrogen catalyzed to the equilibrium ortho-parahydrogen composition are processed and compared to expected statistical distributions for method verification.

08 HYDROGEN↗

Statistical Calibration and Validation of a Homogeneous Ventilated Wall-Interference Correction Method for the National Transonic Facility

Wind tunnel experiments will continue to be a primary source of validation data for many types of mathematical and computational models in the aerospace industry. The increased emphasis on accuracy of data acquired from these facilities requires understanding of the uncertainty of not only the measurement data but also any correction applied to the data. One of the largest and most critical corrections made to these data is due to wall interference. In an effort to understand the accuracy and suitability of these corrections, a statistical validation process for wall interference correction methods has been developed. This process is based on the use of independent cases which, after correction, are expected to produce the same result. Comparison of these independent cases with respect to the uncertainty in the correction process establishes a domain of applicability based on the capability of the method to provide reasonable corrections with respect to customer accuracy requirements. The statistical validation method was applied to the version of the Transonic Wall Interference Correction System (TWICS) recently implemented in the National Transonic Facility at NASA Langley Research Center. The TWICS code generates corrections for solid and slotted wall interference in the model pitch plane based on boundary pressure measurements. Before validation could be performed on this method, it was necessary to calibrate the ventilated wall boundary condition parameters. Discrimination comparisons are used to determine the most representative of three linear boundary condition models which have historically been used to represent longitudinally slotted test section walls. Of the three linear boundary condition models implemented for ventilated walls, the general slotted wall model was the most representative of the data. The TWICS code using the calibrated general slotted wall model was found to be valid to within the process uncertainty for test section Mach numbers less than or equal to 0.60. The scatter among the mean corrected results of the bodies of revolution validation cases was within one count of drag on a typical transport aircraft configuration for Mach numbers at or below 0.80 and two counts of drag for Mach numbers at or below 0.90.

Walker, Eric Lee↗

Statistical Calibration and Validation of a Homogeneous Ventilated Wall-Interference Correction Method for the National Transonic Facility

Wind tunnel experiments will continue to be a primary source of validation data for many types of mathematical and computational models in the aerospace industry. The increased emphasis on accuracy of data acquired from these facilities requires understanding of the uncertainty of not only the measurement data but also any correction applied to the data. One of the largest and most critical corrections made to these data is due to wall interference. In an effort to understand the accuracy and suitability of these corrections, a statistical validation process for wall interference correction methods has been developed. This process is based on the use of independent cases which, after correction, are expected to produce the same result. Comparison of these independent cases with respect to the uncertainty in the correction process establishes a domain of applicability based on the capability of the method to provide reasonable corrections with respect to customer accuracy requirements. The statistical validation method was applied to the version of the Transonic Wall Interference Correction System (TWICS) recently implemented in the National Transonic Facility at NASA Langley Research Center. The TWICS code generates corrections for solid and slotted wall interference in the model pitch plane based on boundary pressure measurements. Before validation could be performed on this method, it was necessary to calibrate the ventilated wall boundary condition parameters. Discrimination comparisons are used to determine the most representative of three linear boundary condition models which have historically been used to represent longitudinally slotted test section walls. Of the three linear boundary condition models implemented for ventilated walls, the general slotted wall model was the most representative of the data. The TWICS code using the calibrated general slotted wall model was found to be valid to within the process uncertainty for test section Mach numbers less than or equal to 0.60. The scatter among the mean corrected results of the bodies of revolution validation cases was within one count of drag on a typical transport aircraft configuration for Mach numbers at or below 0.80 and two counts of drag for Mach numbers at or below 0.90.

Walker, Eric L.↗

Data for Transposon Signatures of Allopolyploid Genome Evolution

Hybridization brings together chromosome sets from two or more distinct progenitor species. Genome duplication associated with hybridization, or allopolyploidy, allows these chromosome sets to persist as distinct subgenomes during subsequent meioses. Here, we present a general method for identifying the subgenomes of a polyploid based on shared ancestry as revealed by the genomic distribution of repetitive elements that were active in the progenitors. This subgenome-enriched transposable element signal is intrinsic to the polyploid, allowing broader applicability than other approaches that depend on the availability of sequenced diploid relatives. We develop the statistical basis of the method, demonstrate its applicability in the well-studied cases of tobacco, cotton, and Brassica napus, and apply it to several cases: allotetraploid cyprinids, allohexaploid false flax, and allooctoploid strawberry. These analyses provide insight into the origins of these polyploids, revise the subgenome identities of strawberry, and provide perspective on subgenome dominance in higher polyploids.

Genomics↗

A Bayesian Framework for Spectral Reprojection

Abstract Fourier partial sum approximations yield exponential accuracy for smooth and periodic functions, but produce the infamous Gibbs phenomenon for non-periodic ones. Spectral reprojection resolves the Gibbs phenomenon by projecting the Fourier partial sum onto a Gibbs complementary basis, often prescribed as the Gegenbauer polynomials. Noise in the Fourier data and the Runge phenomenon both degrade the quality of the Gegenbauer reconstruction solution, however. Motivated by its theoretical convergence properties, this paper proposes a new Bayesian framework for spectral reprojection, which allows a greater understanding of the impact of noise on the reprojection method from a statistical point of view. We are also able to improve the robustness with respect to the Gegenbauer polynomials parameters. Finally, the framework provides a mechanism to quantify the uncertainty of the solution estimate.

Li, Tongtong (ORCID:0000000276644764)↗