Search NASA⌕ Search

SEARCH · Search NASA

Results for “Factor Analysis, Statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Technical Evaluation of the NASA Model for Cancer Risk to Astronauts Due to Space Radiation

At the request of NASA, the National Research Council's (NRC's) Committee for Evaluation of Space Radiation Cancer Risk Model reviewed a number of changes that NASA proposes to make to its model for estimating the risk of radiation-induced cancer in astronauts. The NASA model in current use was last updated in 2005, and the proposed model would incorporate recent research directed at improving the quantification and understanding of the health risks posed by the space radiation environment. NASA's proposed model is defined by the 2011 NASA report Space Radiation Cancer Risk Projections and Uncertainties 2010 (Cucinotta et al., 2011). The committee's evaluation is based primarily on this source, which is referred to hereafter as the 2011 NASA report, with mention of specific sections or tables cited more formally as Cucinotta et al. (2011). The overall process for estimating cancer risks due to low linear energy transfer (LET) radiation exposure has been fully described in reports by a number of organizations. They include, more recently: (1) The "BEIR VII Phase 2" report from the NRC's Committee on Biological Effects of Ionizing Radiation (BEIR) (NRC, 2006); (2) Studies of Radiation and Cancer from the United Nations Scientific Committee on the Effects of Atomic Radiation (UNSCEAR, 2006), (3) The 2007 Recommendations of the International Commission on Radiological Protection (ICRP), ICRP Publication 103 (ICRP, 2007); and (4) The Environmental Protection Agency s (EPA s) report EPA Radiogenic Cancer Risk Models and Projections for the U.S. Population (EPA, 2011). The approaches described in the reports from all of these expert groups are quite similar. NASA's proposed space radiation cancer risk assessment model calculates, as its main output, age- and gender-specific risk of exposure-induced death (REID) for use in the estimation of mission and astronaut-specific cancer risk. The model also calculates the associated uncertainties in REID. The general approach for estimating risk and uncertainty in the proposed model is broadly similar to that used for the current (2005) NASA model and is based on recommendations by the National Council on Radiation Protection and Measurements (NCRP, 2000, 2006). However, NASA's proposed model has significant changes with respect to the following: the integration of new findings and methods into its components by taking into account newer epidemiological data and analyses, new radiobiological data indicating that quality factors differ for leukemia and solid cancers, an improved method for specifying quality factors in terms of radiation track structure concepts as opposed to the previous approach based on linear energy transfer, the development of a new solar particle event (SPE) model, and the updates to galactic cosmic ray (GCR) and shielding transport models. The newer epidemiological information includes updates to the cancer incidence rates from the life span study (LSS) of the Japanese atomic bomb survivors (Preston et al., 2007), transferred to the U.S. population and converted to cancer mortality rates from U.S. population statistics. In addition, the proposed model provides an alternative analysis applicable to lifetime never-smokers (NSs). Details of the uncertainty analysis in the model have also been updated and revised. NASA's proposed model and associated uncertainties are complex in their formulation and as such require a very clear and precise set of descriptions. The committee found the 2011 NASA report challenging to review largely because of the lack of clarity in the model descriptions and derivation of the various parameters used. The committee requested some clarifications from NASA throughout its review and was able to resolve many, but not all, of the ambiguities in the written description.

Source record↗

Novel insights enabled by combining mouse muscle datasets from the Rodent Research-1 mission

Biological space experiments are often expensive and difficult to conduct. As such, it is critical to maximize the value of the data that is collected during these experiments. One way to do this is to combine multiple–previously separate–datasets. This can increase the number of replicates for the conditions of interest (and hence statistical power), allow new multi-factor questions to be asked, and potentially highlight new patterns that otherwise would not have been identified from single-dataset studies. However, the process of combining datasets introduces noise due to inherent technical variations between experiments. To better understand the insights that can be gained from multi-dataset analyses and the problems that may arise from joining multiple datasets, several mouse muscle RNA-Seq datasets from the Rodent Research-1 mission were first selected. Then, using the R package DESeq2, principal component analysis (PCA) plots and differentially expressed gene (DEG) lists between ground and flight muscle samples were generated for individual datasets and for different pairwise combinations of datasets. Several new DEGs were identified in the combined datasets, and patterns in the PCA plots were affected depending on which datasets were joined. Understanding the results of this work will be critical for future studies that seek to perform multi-dataset analyses.

spaceflight↗

Recognition and characterization of hierarchical interstellar structure. II - Structure tree statistics

A new method of image analysis is described, in which images partitioned into 'clouds' are represented by simplified skeleton images, called structure trees, that preserve the spatial relations of the component clouds while disregarding information concerning their sizes and shapes. The method can be used to discriminate between images of projected hierarchical (multiply nested) and random three-dimensional simulated collections of clouds constructed on the basis of observed interstellar properties, and even intermediate systems formed by combining random and hierarchical simulations. For a given structure type, the method can distinguish between different subclasses of models with different parameters and reliably estimate their hierarchical parameters: average number of children per parent, scale reduction factor per level of hierarchy, density contrast, and number of resolved levels. An application to a column density image of the Taurus complex constructed from IRAS data is given. Moderately strong evidence for a hierarchical structural component is found, and parameters of the hierarchy, as well as the average volume filling factor and mass efficiency of fragmentation per level of hierarchy, are estimated. The existence of nested structure contradicts models in which large molecular clouds are supposed to fragment, in a single stage, into roughly stellar-mass cores.

Houlahan, Padraig↗

Constraining neutrino-nucleon form factors with charged-current scattering at the Electron-Ion Collider

Next-generation neutrino oscillation experiments such as the Deep Underground Neutrino Experiment require percent-level knowledge of neutrino-nucleon interaction cross sections. The nucleon axial form factor 𝐹 𝐴 ⁡(𝑄 2 ), parametrized by the axial mass 𝑀 𝐴 , is the dominant source of uncertainty in the quasielastic channel, and the parity-violating structure function 𝑥⁢𝐹 3 is poorly constrained on free nucleons. We propose using charged-current (CC) electron-proton scattering at the Electron-Ion Collider (EIC) to address both problems simultaneously. The measurement exploits three key features of the EIC: (1) helicity-selective electron bunches provide in situ electromagnetic background rejection; (2) a longitudinally polarized proton target enables extraction of 𝐹 𝐴 ⁡(𝑄 2 ) through the target-spin asymmetry 𝐴 𝑈⁢𝐿 ; and (3) the 𝑦-distribution leverage in CC deep inelastic scattering (DIS) separates 𝐹 2 and 𝑥⁢𝐹 3 on a free proton, without nuclear corrections. Using a Fisher information analysis at $\sqrt{𝑠}$ =141 GeV with 500 fb −1 of integrated luminosity, we project the Cramér-Rao statistical floor of 𝛿⁢𝑀 𝐴 ≈0.03 GeV (3%). Incorporating first-order realistic detector effects, such as zero-degree calorimeter acceptance, 𝑄 2 smearing (5%), and background noise from helicity subtraction, the projected sensitivity is severely background-limited due to the small signal-to-background ratio (𝑆/𝐵 ≈ 3 ×10 −4 ) in the elastic channel. Achieving competitive sensitivity (𝛿⁢𝑀 𝐴 ≈ 0.14 GeV) would require ∼10 −7 background suppression, 3 orders of magnitude beyond current projections. The CC DIS 𝑦 distribution provides subpercent extraction of 𝑥⁢𝐹$^{𝑊^{−}}_{3}$ over 0.05 < 𝑥 < 0.5, representing the most robust electroweak measurement in the near term.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Study of Uncertainties of Predicting Space Shuttle Thermal Environment

Quantitative estimates of the uncertainty in predicting aerodynamic heating rates for a fully reusable space shuttle system are developed and the impact of these uncertainties on Thermal Protection System (TPS) weight are discussed. The study approach consisted of statistical evaluations of the scatter of heating data on shuttle configurations about state-of-the-art heating prediction methods to define the uncertainty in these heating predictions. The uncertainties were then applied as heating rate increments to the nominal predicted heating rate to define the uncertainty in TPS weight. Separate evaluations were made for the booster and orbiter, for trajectories which included boost through reentry and touchdown. For purposes of analysis, the vehicle configuration is divided into areas in which a given prediction method is expected to apply, and separate uncertainty factors and corresponding uncertainty in TPS weight derived for each area.

Fehrman, A. L.↗

Assessment of Mars Atmospheric Temperature Retrievals from the Thermal Emission Spectrometer Radiances

Motivated by the needs of Mars data assimilation. particularly quantification of measurement errors and generation of averaging kernels. we have evaluated atmospheric temperature retrievals from Mars Global Surveyor (MGS) Thermal Emission Spectrometer (TES) radiances. Multiple sets of retrievals have been considered in this study; (1) retrievals available from the Planetary Data System (PDS), (2) retrievals based on variants of the retrieval algorithm used to generate the PDS retrievals, and (3) retrievals produced using the Mars 1-Dimensional Retrieval (M1R) algorithm based on the Optimal Spectral Sampling (OSS ) forward model. The retrieved temperature profiles are compared to the MGS Radio Science (RS) temperature profiles. For the samples tested, the M1R temperature profiles can be made to agree within 2 K with the RS temperature profiles, but only after tuning the prior and error statistics. Use of a global prior that does not take into account the seasonal dependence leads errors of up 6 K. In polar samples. errors relative to the RS temperature profiles are even larger. In these samples, the PDS temperature profiles also exhibit a poor fit with RS temperatures. This fit is worse than reported in previous studies, indicating that the lack of fit is due to a bias correction to TES radiances implemented after 2004. To explain the differences between the PDS and Ml R temperatures, the algorithms are compared directly, with the OSS forward model inserted into the PDS algorithm. Factors such as the filtering parameter, the use of linear versus nonlinear constrained inversion, and the choice of the forward model, are found to contribute heavily to the differences in the temperature profiles retrieved in the polar regions, resulting in uncertainties of up to 6 K. Even outside the poles, changes in the a priori statistics result in different profile shapes which all fit the radiances within the specified error. The importance of the a priori statistics prevents reliable global retrievals based a single a priori and strongly implies that a robust science analysis must instead rely on retrievals employing localized a priori information, for example from an ensemble based data assimilation system such as the Local Ensemble Transform Kalman Filter (LETKF).

Hoffman, Matthew J.↗

Error Estimation of An Ensemble Statistical Seasonal Precipitation Prediction Model

This NASA Technical Memorandum describes an optimal ensemble canonical correlation forecasting model for seasonal precipitation. Each individual forecast is based on the canonical correlation analysis (CCA) in the spectral spaces whose bases are empirical orthogonal functions (EOF). The optimal weights in the ensemble forecasting crucially depend on the mean square error of each individual forecast. An estimate of the mean square error of a CCA prediction is made also using the spectral method. The error is decomposed onto EOFs of the predictand and decreases linearly according to the correlation between the predictor and predictand. Since new CCA scheme is derived for continuous fields of predictor and predictand, an area-factor is automatically included. Thus our model is an improvement of the spectral CCA scheme of Barnett and Preisendorfer. The improvements include (1) the use of area-factor, (2) the estimation of prediction error, and (3) the optimal ensemble of multiple forecasts. The new CCA model is applied to the seasonal forecasting of the United States (US) precipitation field. The predictor is the sea surface temperature (SST). The US Climate Prediction Center's reconstructed SST is used as the predictor's historical data. The US National Center for Environmental Prediction's optimally interpolated precipitation (1951-2000) is used as the predictand's historical data. Our forecast experiments show that the new ensemble canonical correlation scheme renders a reasonable forecasting skill. For example, when using September-October-November SST to predict the next season December-January-February precipitation, the spatial pattern correlation between the observed and predicted are positive in 46 years among the 50 years of experiments. The positive correlations are close to or greater than 0.4 in 29 years, which indicates excellent performance of the forecasting model. The forecasting skill can be further enhanced when several predictors are used.

Shen, Samuel S. P.↗

Fundamental Mistuning Model for Probabilistic Analysis Studied Experimentally

The Fundamental Mistuning Model (FMM) is a reduced-order model for efficiently calculating the forced response of a mistuned bladed disk. FMM ID is a companion program that determines the mistuning in a particular rotor. Together, these methods provide a way to acquire mistuning data in a population of bladed disks and then simulate the forced response of the fleet. This process was tested experimentally at the NASA Glenn Research Center, and the simulated results were compared with laboratory measurements of a "fleet" of test rotors. The method was shown to work quite well. It was found that the accuracy of the results depends on two factors: (1) the quality of the statistical model used to characterize mistuning and (2) how sensitive the system is to errors in the statistical modeling.

Griffin, Jerry H.↗

Comparison of Comet Enflow and VA One Acoustic-to-Structure Power Flow Predictions

Comet Enflow is a commercially available, high frequency vibroacoustic analysis software based on the Energy Finite Element Analysis (EFEA). In this method the same finite element mesh used for structural and acoustic analysis can be employed for the high frequency solutions. Comet Enflow is being validated for a floor-equipped composite cylinder by comparing the EFEA vibroacoustic response predictions with Statistical Energy Analysis (SEA) results from the commercial software program VA One from ESI Group. Early in this program a number of discrepancies became apparent in the Enflow predicted response for the power flow from an acoustic space to a structural subsystem. The power flow anomalies were studied for a simple cubic, a rectangular and a cylindrical structural model connected to an acoustic cavity. The current investigation focuses on three specific discrepancies between the Comet Enflow and the VA One predictions: the Enflow power transmission coefficient relative to the VA One coupling loss factor; the importance of the accuracy of the acoustic modal density formulation used within Enflow; and the recommended use of fast solvers in Comet Enflow. The frequency region of interest for this study covers the one-third octave bands with center frequencies from 16 Hz to 4000 Hz.

Grosveld, Ferdinand W.↗

Landslides in West Coast Metropolitan Areas: The Role of Extreme Weather Events

Rainfall-induced landslides represent a pervasive issue in areas where extreme rainfall intersects complex terrain. A farsighted management of landslide risk requires assessing how landslide hazard will change in coming decades and thus requires, inter alia, that we understand what rainfall events are most likely to trigger landslides and how global warming will affect the frequency of such weather events. We take advantage of 9 years of landslide occurrence data compiled by collating Google news reports and of a high-resolution satellite-based daily rainfall data to investigate what weather triggers landslide along the West Coast US. We show that, while this landslide compilation cannot provide consistent and widespread monitoring everywhere, it captures enough of the events in the major urban areas that it can be used to identify the relevant relationships between landslides and rainfall events in Puget Sound, the Bay Area, and greater Los Angeles. In all these regions, days that recorded landslides have rainfall distributions that are skewed away from dry and low-rainfall accumulations and towards heavy intensities. However, large daily accumulation is the main driver of enhanced hazard of landslides only in Puget Sound. There, landslide are often clustered in space and time and major events are primarily driven by synoptic scale variability, namely "atmospheric rivers" of high humidity air hitting anywhere along the West Coast, and the interaction of frontal system with the coastal orography. The relationship between landslide occurrences and daily rainfall is less robust in California, where antecedent precipitation (in the case of the Bay area) and the peak intensity of localized downpours at sub-daily time scales (in the case of Los Angeles) are key factors not captured by the same-day accumulations. Accordingly, we suggest that the assessment of future changes in landslide hazard for the entire the West Coast requires consideration of future changes in the occurrence and intensity of atmospheric rivers, in their duration and clustering, and in the occurrence of short-duration (sub-daily) extreme rainfall as well. Major regional landslide events, in which multiple occurrences are recorded in the catalog for the same day, are too rare to allow a statistical characterization of their triggering events, but a case study analysis indicates that a variety of synoptic-scale events can be involved, including not only atmospheric rivers but also broader cold- and warm-front precipitation. That a news-based catalog of landslides is accurate enough to allow the identification of different landslide/ rainfall relationships in the major urban areas along the US West Coast suggests that this technology can potentially be used for other English-language cities and could become an even more powerful tool if expanded to other languages and non-traditional news sources, such as social media.

landslides↗

Power flow analysis of two coupled plates with arbitrary characteristics

The limitation of keeping two plates identical is removed and the vibrational power input and output are evaluated for different area ratios, plate thickness ratios, and for different values of the structural damping loss factor for the source plate (plate with excitation) and the receiver plate. In performing this parametric analysis, the source plate characteristics are kept constant. The purpose of this parametric analysis is to be able to determine the most critical parameters that influence the flow of vibrational power from the source plate to the receiver plate. In the case of the structural damping parametric analysis, the influence of changes in the source plate damping is also investigated. As was done previously, results obtained from the mobility power flow approach will be compared to results obtained using a statistical energy analysis (SEA) approach. The significance of the power flow results are discussed together with a discussion and a comparison between SEA results and the mobility power flow results. Furthermore, the benefits that can be derived from using the mobility power flow approach, are also examined.

Cuschieri, J. M.↗

Ultra heavy cosmic ray experiment (A0178)

The Ultra Heavy Cosmic Ray Experiment (UHCRE) is based on a modular array of 192 side viewing solid state nuclear track detector stacks. These stacks were mounted in sets of four in 48 pressure vessels using 16 peripheral LDEF trays. The geometry factor for high energy cosmic ray nuclei, allowing for Earth shadowing, was 30 sq m sr, giving a total exposure factor of 170 sq m sr y at an orbital inclination of 28.4 degs. Scanning results indicate that about 3000 cosmic ray nuclei in the charge region with Z greater than 65 were collected. This sample is more than ten times the current world data in the field (taken to be the data set from the HEAO-3 mission plus that from the Ariel-6 mission) and is sufficient to provide the world's first statistically significant sample of actinide cosmic rays. Results are presented including a sample of ultra heavy cosmic ray nuclei, analysis of pre-flight and post-flight calibration events and details of track response in the context of detector temperature history. The integrated effect of all temperature and age related latent track variations cause a maximum charge shift of + or - 0.8e for uranium and + or - 0.6e for the platinum-lead group. Astrophysical implications of the UHCRE charge spectrum are discussed.

Thompson, A.↗

The LDEF ultra heavy cosmic ray experiment

The Long Duration Exposure Facility (LDEF) Ultra Heavy Cosmic Ray Experiment (UHCRE) used 16 side viewing LDEF trays giving a total geometry factor for high energy cosmic rays of 30 sq m sr. The total exposure factor was 170 sq m sr y. The experiment is based on a modular array of 192 solid state nuclear track detector stacks, mounted in sets of 4 pressure vessels (3 experiment tray). The extended duration of the LDEF mission has resulted in a greatly enhanced potential scientific yield from the UHCRE. Initial scanning results indicate that at least 2000 cosmic ray nuclei with Z greater than 65 were collected, including the world's first statistically significant sample of actinides. Postflight work to date and the current status of the experiment are reviewed. Provisional results from analysis of preflight and postflight calibrations are presented.

Osullivan, D.↗

Understanding software faults and their role in software reliability modeling

This study is a direct result of an on-going project to model the reliability of a large real-time control avionics system. In previous modeling efforts with this system, hardware reliability models were applied in modeling the reliability behavior of this system. In an attempt to enhance the performance of the adapted reliability models, certain software attributes were introduced in these models to control for differences between programs and also sequential executions of the same program. As the basic nature of the software attributes that affect software reliability become better understood in the modeling process, this information begins to have important implications on the software development process. A significant problem arises when raw attribute measures are to be used in statistical models as predictors, for example, of measures of software quality. This is because many of the metrics are highly correlated. Consider the two attributes: lines of code, LOC, and number of program statements, Stmts. In this case, it is quite obvious that a program with a high value of LOC probably will also have a relatively high value of Stmts. In the case of low level languages, such as assembly language programs, there might be a one-to-one relationship between the statement count and the lines of code. When there is a complete absence of linear relationship among the metrics, they are said to be orthogonal or uncorrelated. Usually the lack of orthogonality is not serious enough to affect a statistical analysis. However, for the purposes of some statistical analysis such as multiple regression, the software metrics are so strongly interrelated that the regression results may be ambiguous and possibly even misleading. Typically, it is difficult to estimate the unique effects of individual software metrics in the regression equation. The estimated values of the coefficients are very sensitive to slight changes in the data and to the addition or deletion of variables in the regression equation. Since most of the existing metrics have common elements and are linear combinations of these common elements, it seems reasonable to investigate the structure of the underlying common factors or components that make up the raw metrics. The technique we have chosen to use to explore this structure is a procedure called principal components analysis. Principal components analysis is a decomposition technique that may be used to detect and analyze collinearity in software metrics. When confronted with a large number of metrics measuring a single construct, it may be desirable to represent the set by some smaller number of variables that convey all, or most, of the information in the original set. Principal components are linear transformations of a set of random variables that summarize the information contained in the variables. The transformations are chosen so that the first component accounts for the maximal amount of variation of the measures of any possible linear transform; the second component accounts for the maximal amount of residual variation; and so on. The principal components are constructed so that they represent transformed scores on dimensions that are orthogonal. Through the use of principal components analysis, it is possible to have a set of highly related software attributes mapped into a small number of uncorrelated attribute domains. This definitively solves the problem of multi-collinearity in subsequent regression analysis. There are many software metrics in the literature, but principal component analysis reveals that there are few distinct sources of variation, i.e. dimensions, in this set of metrics. It would appear perfectly reasonable to characterize the measurable attributes of a program with a simple function of a small number of orthogonal metrics each of which represents a distinct software attribute domain.

Munson, John C.↗

Effects of Natural Variability on the Use of Standard Deviation to Represent Measurement Uncertainties in Atmospheric Composition Studies

Measurement uncertainty is defined as a “non-negative parameter characterizing the dispersion of the quantity values being attributed to a measurand”. It is most common that the uncertainties of GAW hourly measurements, such as greenhouse gas measurements, are reported as standard deviations derived from individual sampling at a higher time resolution (e.g., 1 min). In contrast, the uncertainties of GAW measurements of reactive gases and aerosol properties are reported in percentiles covering the same probability. A quick look at hourly CO 2 data from Cape Grim, Australia yielded some interesting findings: the hourly standard deviation is, on average, more than a factor of 10 higher for the measurements under non-background conditions (over 50% observations), while the difference in average CO 2 amount fraction was less than 2 ppm. The dramatic contrast cannot be explained by the difference in measurement uncertainties, but can largely be attributed the natural variability, or episodic ambient CO 2 variation reflecting changes in meteorological conditions or emissions. These initial findings motivated a more in-depth analysis of the ground-based measurements of trace gases and aerosol properties. This analysis will be using continuous 1 min ground site observations to construct time averaged statistical indicators to evaluate whether the standard deviation is adequate to represent the dispersion, especially under marked influence by natural variability. The suitability of this representation can be determined by examining the difference between the standard deviation and percentiles encompassing the same probability. We will examine time intervals of 1 hour, 3 hours, and 24 hours, with the latter time intervals chosen to match those commonly used in model assessments. We will also investigate how natural variability can alter the probability distribution of the measurands and how adequate the quadrature propagation of uncertainties is under these conditions. The results will include several trace gases (e.g., CO 2 , CO, O 3 , and NO 2 ) with a range of measurement techniques (e.g., PANDORA, in situ), atmospheric lifetimes, and aerosol properties (e.g. scattering coefficient). The findings from this analysis should provide some useful feedback on the best practices for uncertainty reporting.

Measurement Uncertainty↗

Parallelization of the Physical-Space Statistical Analysis System (PSAS)

Atmospheric data assimilation is a method of combining observations with model forecasts to produce a more accurate description of the atmosphere than the observations or forecast alone can provide. Data assimilation plays an increasingly important role in the study of climate and atmospheric chemistry. The NASA Data Assimilation Office (DAO) has developed the Goddard Earth Observing System Data Assimilation System (GEOS DAS) to create assimilated datasets. The core computational components of the GEOS DAS include the GEOS General Circulation Model (GCM) and the Physical-space Statistical Analysis System (PSAS). The need for timely validation of scientific enhancements to the data assimilation system poses computational demands that are best met by distributed parallel software. PSAS is implemented in Fortran 90 using object-based design principles. The analysis portions of the code solve two equations. The first of these is the "innovation" equation, which is solved on the unstructured observation grid using a preconditioned conjugate gradient (CG) method. The "analysis" equation is a transformation from the observation grid back to a structured grid, and is solved by a direct matrix-vector multiplication. Use of a factored-operator formulation reduces the computational complexity of both the CG solver and the matrix-vector multiplication, rendering the matrix-vector multiplications as a successive product of operators on a vector. Sparsity is introduced to these operators by partitioning the observations using an icosahedral decomposition scheme. PSAS builds a large (approx. 128MB) run-time database of parameters used in the calculation of these operators. Implementing a message passing parallel computing paradigm into an existing yet developing computational system as complex as PSAS is nontrivial. One of the technical challenges is balancing the requirements for computational reproducibility with the need for high performance. The problem of computational reproducibility is well known in the parallel computing community. It is a requirement that the parallel code perform calculations in a fashion that will yield identical results on different configurations of processing elements on the same platform. In some cases this problem can be solved by sacrificing performance. Meeting this requirement and still achieving high performance is very difficult. Topics to be discussed include: current PSAS design and parallelization strategy; reproducibility issues; load balance vs. database memory demands, possible solutions to these problems.

Larson, J. W.↗

Continuing Long-term Global SO 2 Data Record with JPSS OMPS Instruments

NASA’s long-term Earth Observing System (EOS) SO 2 climate data record (CDR) started with Aura/Ozone Monitoring Instrument (OMI, launched in 2004) and is now being continued with the SNPP/Ozone Mapping and Profiler Suite (OMPS, launched in 2011). Both OMI and SNPP/OMPS SO 2 CDRs are produced with the Goddard principal component analysis (PCA) spectral fitting algorithm. By inherently accounting for various instrumental factors, the PCA technique enables highly consistent retrievals between different instruments. In this presentation, we will provide an overview on our effort to further extend the EOS SO 2 CDR, by implementing the PCA SO 2 algorithm with multiple OMPS instruments flying on the Joint Polar Satellite System (JPSS) constellation, including NOAA-20 (launched in 2017) and NOAA-21 (launched in 2022). We will present results analyzing our new NOAA-20/OMPS PCA SO 2 EOS continuity product, to be publicly released in fall of 2023. We will show statistical analyses on the quality of NOAA-20 PCA SO 2 product, such as retrieval noise, biases over background areas, and long-term stability. We will employ a previously established top-down method to estimate SO2 emissions from selected large point sources, using NOAA-20 SO 2 retrievals and assimilated wind fields as input. The SO 2 emission estimates derived from NOAA-20 retrievals will be compared with those from OMI, SNPP/OMPS, and S5P/TROPOMI (TROPOspheric Monitoring Instrument). We will also demonstrate the application of a new machine learning technique that further reduces the noise of NOAA-20 SO 2 retrievals. Finally, we will present preliminary PCA SO 2 retrievals from recently launched satellite sensors, including NOAA-21/OMPS and NASA’s geostationary TEMPO (Tropospheric Emissions: Monitoring of Pollution) instrument.

SO2↗

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator↗