Search NASA⌕ Search

SEARCH · Search NASA

Results for “pca”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Predicting the Seawater Chemistry of an Ocean World Using Machine Learning on Isotopic Measurements of Volatile CO2

Introduction: Given the long time intervals required for data transmission to and from ocean worlds targets, low bandwidth for data transmission, time required for data processing and analysis, and potentially extreme radiation environments (e.g., Europa), it is clear that ocean worlds missions will need more autonomous flight instruments and software in order to achieve established science goals. Protracted time intervals for data analysis (e.g., Europa Lander) strongly motivates the development of rapid, consistent and streamlined methods for interpreting data from flight mass spectrometers to e.g., determine how mass spectra from a plume or surface liquid/ice relates to the surface/subsurface. Since mass spectrometry also has the potential to correctly identify biosignatures[1], it is imperative that such methods for interpreting data are consistent and accurate. We used 848 isotope ratio mass spectra from laboratory analyses of CO2 that interacted with ocean worlds-relevant seawaters as a ‘training’ dataset for ‘unsupervised’ machine learning. In unsupervised learning, characteristics of the data are not labeled or linked, and any similarities found only result from the neural network. CO2 isotopologues analyzed for this dataset mimic the remote measurements of CO2 by a flight mass spectrometer, and are detailed in Theiling [2]. From this dataset, we used measured features of the spectra, such as retention time, intensity, and (isotopologue) mass ratios as inputs for our autoencoder neural network. Our neural network was trained to find similarities in these and other spectral features for seawaters of a particular composition and amount of initial CO2. Successful training then created an output of these similarities for various seawaters, which included MgSO4, Na2SO4, NaCl, MgCl2, KCl, and NaHCO3, and combinations of these salts. We then applied dimensionality reduction techniques such as Principal Component Analysis (PCA), T-Distributed Stochastic Neighbor Embedding (TSNE), and Uniform Manifold Approximation and Projection (UMAP) to demonstrate latent data features as a two-dimensional projection in a unitless, high-dimensional space. In this projection, a data point represents the combined effect of spectral features such as intensity, retention time, and isotope ratio. Our initial UMAP demonstrates data clustering (organization of the data by the neural network) based on the amount of CO2 that had initially interacted with each seawater. Further training using more ‘supervised’ learning techniques demonstrate strong clustering of preliminary data based on initial CO2 concentration, seawater chemical composition, and ionic strength (salinity). Our preliminary work therefore suggests that machine learning has the potential to identify compositional variants of an ocean world seawater based on mass spectra from volatile CO2 measurements. Acknowledgments: This work was funded through a Strategic Task Group at NASA Goddard Space Flight Center. The training dataset was collected through funding from the Oklahoma Space Grant Consortium. References: [1] Pappalardo, R. et al. (2013) Astrobiology, 13, 740–773. [2] Theiling (2020) Icarus, 114216.

Europa↗

The Abundances of F, Cl, and H2O in 4Vesta from Eucrites

The abundance and distribution of magmatic volatiles (i.e., H, C, N, F, S, and Cl) within the silicate portion of a differentiated planetary body has important consequences on its thermochemical evolution. However, the abundances of magmatic volatiles within differentiated bodies are difficult to quantify, and they are often depleted by varying degrees relative to CI chondrites. The mechanisms of depletion are not well constrained and could relate to intrinsic volatile depletion of the building blocks that formed the bodies, high temperature processes that result from accretion, post-accretion loss through parent body geological processes and large-scale impacts, and/or redistribution within a parent body through processes like core formation [1–4]. In the present study, we aim to constrain the abundances of F, Cl, and H2O in eucrites to better understand the magnitude of volatile depletion on 4Vesta. To accomplish this objective, we report electron microprobe analyses of apatite from seven unbrecciated, non-cumulate eucrites (i.e., CMS 04049,GRA 98098, LEW 88010, MAC 02522, MAC 041169,QUE 94484, and QUE 97053) and two monomict, non-cumulate eucrites (i.e., Berthoud and Stannern). In combination with previously published data on eucrite-hosted apatite, we determine Cl/F and H2O/F ratios in bulk rock eucrites through the application of apatite-based melt hygrometry and chlorometry [e.g., 5–7].Additionally, we estimate the bulk rock abundances of F in six non-cumulate eucrites (i.e., GRA 98098, MAC041169, PCA 91078, QUE 97053, Stannern, and Berthoud), which we combine with previously published bulk rock F data on non-cumulate eucrites[8] to constrain the abundances of F, Cl, and H2O in 4Vesta using appropriately paired volatile/refractory element ratios for F, followed by Cl/F and H2O/F ratios for Cl and H2O, respectively.

F M McCubbin↗

Reducing OCO-2 regional biases through novel 3D cloud, albedo, and meteorology estimation

This ROSES project tests the addition of novel 3D cloud, albedo fine structure, temperature and water vapor profile to the OCO-2 retrieved system. All new parameters have been added to the ReFRACtor / TROPESS system for testing. Currently, we are comparing 3d-cloud retrievals from OCO-2 radiances to the MODIS-derived EAR3T cloud properties. The water and temperature profile retrievals have been implemented and result in improved XCO2 comparisons to TCCON, and now are implementing principal-component analysis (PCA) to reduce the number of parameters added. We are also studying ECOSTRESS spectral library, AVIRIS-ng, OCO-2 spectral residuals, and albedo retrievals with additional parameters to determine the appropriate characterization of albedo. We find that the current albedo parametrization (2nd order polynomial) is adequate for ocean and vegetation, but not adequate for soil and rock observations.

Reducing OCO-2↗

SWIPE: Spectral Water Inversion Processor and Emulator

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will be discussing the progress made developing SWIPE: Spectral Water Inversion Processor and Emulator. SWIPE is a platform for advanced modeling of coastal and inland aquatic habitats. The goal is create a comprehensive and cohesive system to leverage recent advancements in computation and machine learning to develop a synthetic training ground for sensitivity studies and algorithm development. The four principal facets of SWIPE include: 1. Advanced two-layer coated sphere bio-optical modeling and GPU radiative transfer modeling, 2. Big Data involving massive synthetic spectral libraries of optical properties of various global aquatic particles, surface reflectance, and top-of-atmosphere reflectance, all at hyperspectral resolution leveraging high-end computing systems at NASA Ames Research Center, 3. Deep Learning for algorithm development for water quality inversion of concentrations of common biogeophysical variables as well as optics, full uncertainty characterization by water type, and forward emulation, and lastly, 4. Image Processing for application of developed retrieval algorithms for both hyperspectral and multispectral sensors with experimental corrections for global adjacency, noise, sunglint, and benthic reflectance. This presentation will demonstrate the Equivalent Algal Populations (EAP) two-layer coated sphere scattering model which has been used develop spectral libraries of hyperspectral inherent optical properties of roughly 80 species of phytoplankton, covering 15 different classes and nine taxonomic functional types. The EAP model was also used to derive spectral properties of 10 different non-algal particle functional types. Examples of how the SMART-G (Speed-up Monte-carlo Advanced Radiative Transfer using GPU) radiative transfer code is used to model optically complex aquatic signals will be presented and discussed in the context of creating a massive synthetic database which can leverage the full power of next generation machine learning techniques and high end computing for water quality inversion. We will discuss our active investigation in things like appropriate model architectures, dimensionality reduction techniques such as PCA and autoencoders, uncertainty quantification and abstaining, and which variables actually benefit most from hyperspectral information versus multispectral resolution. We are also curious about questions relating to cost/benefit analysis in terms of computation resources, neural network complexity, and data volumes. Answers to these questions will hopefully elaborate on cost efficiency for potential future sensor design considerations.

SWIPE↗

In Situ Measurements of Surface Texture with Virtual Environments Support Science-Driven Human Surface Operations on the Moon and Beyond

Visualization tools enabling real-time scientific analysis are important for supporting future astronaut operations on the lunar surface. Such tools can be built into virtual environments to support scientific investigations, as well as situational awareness, real-time decision making, and efficient communication between astronauts and ground and support systems. Understanding how these tools can be optimized for science is essential for upcoming Artemis missions. In this contribution, we discuss how measurements of surface texture at multiple length scales can greatly enhance in situ science on/of the Moon, and eventually Mars, asteroids, and beyond. Roughness measurements at various wavelengths directly support objectives defined in the Artemis Science Plan, including (O1) “understanding planetary processes,” (O2) “understanding volatile cycles,” and (O3) “interpreting the impact history of the Earth-Moon system” . Key scientific analyses enabled by texture measurements at different length scales include: ● Sub-centimeter scales: Texture measurements can help constrain lava flow crystallinity, lava rheology, emplacement flow dynamics, and cooling histories (O1). Measurements of lacunarity (voids in fractal fill space) can shed light on eruptive volatile content, residence time of migrating volatiles, and near-surface volume available for micro-cold trapping of volatiles (O1, O2). ● Centimeter–meter scales: Texture measurements can be used for the differentiation of individual lava flows, the reconstruction of local stratigraphies and emplacement sequences, characterization of post-emplacement surface modification processes (O1, O3). Derived roughness (polarization) metrics can be used in the detection of water ice and characterization of ice properties (e.g., purity, grade, depth, abundance). ● Hectometer–Kilometer scales: Texture measurements can be used to differentiate major geologic surface units and surface structures (O1), constrain the presence of abundant ground ices (O2), and analyze surface modification and estimate surface age (O3). Real-time measurements of surface texture across these multiple length scales will enable efficient sample identification and scientific investigations by future astronauts. To support these investigations and the objective classification of surface texture, virtual environments employed by astronauts should be able to instantaneously convert raw data into processed data (e.g., digital terrain and elevation models) and derived metrics (e.g., RMS, std, Hurst, CPR) and perform statistical analyses (e.g., PCA, outliers, correlation matrices). Such tools are being developed and tested by the Resource Exploration and Science of our Cosmic Environment (RESOURCE) team, a node of NASA’s Solar System Exploration Research Virtual Institute (SSERVI), and are an excellent example of the powerful synergies of human and robotic ground assets critical in the return of humans to the Moon.

Ariel N. Deutsch↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

A Sulfur Dioxide Covariance-Based Retrieval Algorithm (COBRA): Application to TROPOMI Reveals New Emission Sources

Sensitive and accurate detection of sulfur dioxide (SO 2 ) from space is important for monitoring and estimating global sulfur emissions. Inspired by detection methods applied in the thermal infrared, we present here a new scheme to retrieve SO 2 columns from satellite observations of ultraviolet back-scattered radiances. The retrieval is based on a measurement error covariance matrix to fully represent the SO 2 -free radiance variability, so that the SO 2 slant column density is the only retrieved parameter of the algorithm. We demonstrate this approach, named COBRA, on measurements from the TROPOspheric Monitoring Instrument (TROPOMI) aboard the Sentinel-5 Precursor (S-5P) satellite. We show that the method reduces significantly both the noise and biases present in the current TROPOMI operational DOAS SO 2 retrievals. The performance of this technique is also benchmarked against that of the Principal Component Algorithm (PCA) approach. We find that the quality of the data is similar and even slightly better with the proposed COBRA approach. The ability of the algorithm to retrieve SO 2 accurately is also further supported by comparison with ground-based observations. We illustrate the great sensitivity of the method with a high-resolution global SO 2 map, considering two and a half years of TROPOMI data. In addition to the known sources, we detect many new SO 2 emission hotspots worldwide. For the largest sources, we use the COBRA data to estimate SO 2 emission rates. Results are comparable to other recently published TROPOMI-based SO 2 emissions estimates, but the associated uncertainties are significantly lower than with the operational data. Next, for a limited number of weak sources, we demonstrate the potential of our data for quantifying SO 2 emissions with a detection limit of about 8 kt yr -1 , a factor of 4 better than the emissions derived from the Ozone Monitoring Instrument (OMI). We anticipate that the systematic use of our TROPOMI COBRA SO 2 column data set at a global scale will allow identifying and quantifying missing sources, and help improving SO 2 emission inventories.

SO2↗

Variability in Mt. Sharp Group Bedrock as Seen By ChemCam Passive and Active Spectra

The Curiosity rover landed in Gale crater in August 2012 and has since been travelling up the central sedimentary mound known as Mt. Sharp. The ChemCam instrument on Curiosity was designed primarily for the use of Laser Induced Breakdown Spectroscopy (LIBS), where a laser ablates a small amount of material from the target and the spectrum of the resulting plasma yields elemental abundance data. ChemCam’s three spectrometers range from 240-905 nm and can also take passive spectra (without the use of the laser). The spectral range ChemCam passive spectra observe is sensitive to charge-transfer and crystal field absorptions related to iron-bearing minerals. In the first 2934 sols of Curiosity’s mission, 9,400 passive spectra were taken of bedrock targets in Mt. Sharp’s Murray and Carolyn Shoemaker formations. We examine these spectra using spectral slope/ratio and band depth calculations as well as Principal Component Analysis (PCA). For the first time, paired passive spectra and LIBS elemental abundances are compared on a large scale. Finally, CheMin data are compared to ChemCam passive observations to understand sources of spectral variability.

H T Manelski↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

COMPACT KNN V2: Analogy-Based Cost Estimation Model for CubeSats

The CubeSat Or Microsat Probabilistic and AnalogiesCost Tool, or COMPACT, is a NASA Headquarters fundedeffort to fill the gap in cost estimating capabilities for CubeSats,as well as other microsat spacecraft. The COMPACT team hasfocused mainly on CubeSats to date, and has collected technical,programmatic and cost data on dozens of flown CubeSatsmissions led by NASA, research labs, and universities. In late2019, the team released the first tool prototype which uses a nonparametricregression technique, k-Nearest Neighbors (KNN),on actual data from historical CubeSat missions to produceearly ballpark analogy-based cost estimates for new CubeSatconcepts. Since the KNN prototype was first released, theCOMPACT team has normalized 17 new missions to be addedto the model in COMPACT V2. COMPACT V2 also featureschanges to the KNN tool algorithm including the introduction ofPrincipal Component Analysis (PCA) to the model developmentprocess and changes to the input parameters which have madethe analogy results more intuitive and have improved modelperformance. This paper describes the current COMPACTKNN dataset, improvements made to the model in COMPACTV2, an assessment of current model performance, and a forwardlook at COMPACT’s planned future enhancements.

Hooke, Melissa↗

Enabling Intelligent Data Downlink Prioritization of In-Situ Observations through Generalizable and Computationally Inexpensive Anomaly Detection

High-fidelity measurements of magnetic fields and other observed properties, such as energetic particle fluxes, are a necessary component to our understanding of the highly dynamic near-Earth space environment. As our desire to study smaller-scale phenomena such as shocks and dipolorizations has increased, we have been driven to take and telemeter measurements at higher cadences. Unfortunately, many missions are unable to downlink all their captured data due to the well-known data transmission bottleneck at the DSN. These missions must then prioritize their high-cadence data such that the most scientifically useful intervals are transmitted. One simple prioritization technique uses the spacecraft position to telemeter data from only the region of interest. Although easy to implement, this method does not leverage the available scientific data and can omit intervals of useful scientific data when they lie outside the region of interest. The Magnetospheric Multiscale Mission (MMS) uses mission-specific parameterization of several data products to automatically prioritize scientifically useful intervals. Then, MMS verifies the automatically selected intervals by having a domain expert manually select intervals for downlink. The overall complexity required by this technique make it prohibitive for deployment on low-cost platforms (i.e., CubeSats) or on future missions featuring large constellations of satellites such as the Geospace Dynamics Constellation (GDC). We present preliminary results for a simple, generic, and data-driven method of downlink prioritization for magnetic field (and other) measurements. Specifically, Principal Components Analysis (PCA) and One-Class Support Vector Machines (OC-SVMs) are used to detect intervals containing anomalous activity, which can then be prioritized for subsequent downlink. The computational simplicity of this algorithm makes it an excellent candidate for implementation on spaceflight hardware, as well as provide generalizability to a broad range of missions and data products. Initial analysis of this technique has been performed using magnetic field measurements from the Magnetospheric Multiscale Mission and CASSIOP, where it automatically identified scientifically interesting intervals containing Alfvén waves and EMIC activity.

Matthew G. Finley↗

Training and Validation of Spectral Gap Filling Algorithm for Cpf-Ceres Intercalibration

The Climate Absolute Radiance and Refractivity Observatory (CLARREO) Pathfinder (CPF) mission is set to launch an SI-traceable reflective solar (RS) spectrometer aboard the International Space Station to measure Earth-reflected solar radiation with a radiometric uncertainty of 0.3% (k=1). The CPF intercalibration team has devised a cutting-edge methodology to accurately transfer the benchmark CPF calibration reference to the shortwave (SW) channel (200-5000 nm) of the Clouds and the Earth’s Radiant Energy System (CERES) instrument. The spectral range of CPF measurements spans from 350-2300 nm, while the CERES SW channel measures the Earth-reflected broadband solar radiances between 200 nm to 5 μm. To conduct precise CPF-CERES intercalibration analysis, the CPF-like spectral radiances outside the CPF spectral range need to be estimated to match the CERES SW spectral range. In response, the team has developed a fast algorithm that leverages spectrally redundant information within the CPF-measured portion through principal component analysis (PCA) and utilizes pre-established spectral correlation relationships among wavelengths to extend the CPF spectrum below 350 nm and above 2300 nm. Our results show that the algorithm achieves excellent accuracy in generating the missing energy in the UV and IR portions. The RMS error in the UV region is less than 4.5x10-3 W/m2/sr/nm, while in the IR region, it is smaller than 8x10-5 W/m2/sr/nm. Our methodology was validated using measured EMIT radiance data, which covers the spectral range from 0.381 μm to 2.493 μm. We employed EMIT radiances within the wavelength range of 0.43 – 2.25 μm to generate radiances for both the shorter wavelength range (0.381 – 0.43 μm) and longer wavelength range (2.25 – 2.493 μm). The generated radiances agree very well with the measured EMIT radiances. The standard deviation in the integrated broadband radiances was about 0.1%, and the bias is less than 0.004% for over 1.5 million EMIT measured samples. These statistics show that the spectral gap filling algorithm is robust and effective in substantially reducing the spectral difference-induced uncertainty in the CPF-CERES intercalibration samples.

Qiguang Yang↗

Development of a Data Fusion Methodology for Lineload Aerodynamic Databases for a Launch Vehicle during Liftoff and Transition

The need for databases for the distributed loading on launch vehicles during the early portion of flight necessitates the use of expensive computational flows in regimes where wake effects dominate. While also being expensive, this is a regime that computational tools tend to historically have problems simulating accurately. To help tackle this problem, a method of data fusion to combine computational results to wind tunnel derived force and moment data is developed. Using this method, significant reduction in computational costs and increases in confidence of the final product is possible and has been used to generate several databases for the Space Launch System (SLS) at NASA. While the full details of database generation are not part of this work, the crucial method at its core is developed here. Two SLS geometries are used throughout the work to demonstrate the techniques. These are two of the larger geometries and represent both planned crewed missions to the Moon as well as potential cargo missions to deep space. The method uses principal component analysis (PCA) to generate a reduced ordered model (ROM) to help fill in the full parameter space. Other similar techniques are explored, but were not found to have a significant result on the predictions of the ROM. Because the full number of components are kept to generate the model, this lack of difference is expected. This method is then extended to ensure that predicted surfaces match trusted force and moment data derived from wind tunnel testing. This extension is done by setting up a constrained optimization problem in order to minimize the deviation from the surface resolved computational data while still integrating to the desired values. When generating the constrained optimization problem, a weighting factor to balance these competing needs is introduced. The work compares previously introduced weighting terms from similar work to the proposed terms and shows that the previously used terms do not have as desirable behavior in this flow regime. This method is then expanded by developing a technique to incorporate uncertainty quantification into the developed data fusion methodology. This expansion takes a two pronged approach. One examines transferring the uncertainties in the force and moment database and characterizes how those adjustments change the predicted lineloads. The second looks at model form error and looks how rebuilding the model using slightly different data changes the predictions. These two terms are then combined in order to create an uncertainty model that takes both effects into account. The limitations of the proposed methods is then discussed as well as possible techniques to address these shortcomings.

Launch Vehicles↗

Enabling Intelligent Data Downlink Prioritization of In-Situ Observations through Generalizable and Computationally Inexpensive Anomaly Detection

High-fidelity measurements of magnetic fields and other observed properties, such as energetic particle fluxes, are a necessary component to our understanding of the highly dynamic near-Earth space environment. As our desire to study smaller-scale phenomena such as shocks and dipolorizations has increased, we have been driven to take and telemeter measurements at higher cadences. Unfortunately, many missions are unable to downlink all their captured data due to the well-known data transmission bottleneck at the DSN. These missions must then prioritize their high-cadence data such that the most scientifically useful intervals are transmitted. One simple prioritization technique uses the spacecraft position to telemeter data from only the region of interest. Although easy to implement, this method does not leverage the available scientific data and can omit intervals of useful scientific data when they lie outside the region of interest. The Magnetospheric Multiscale Mission (MMS) uses mission-specific parameterization of several data products to automatically prioritize scientifically useful intervals. Then, MMS verifies the automatically selected intervals by having a domain expert manually select intervals for downlink. The overall complexity required by this technique make it prohibitive for deployment on low-cost platforms (i.e., CubeSats) or on future missions featuring large constellations of satellites such as the Geospace Dynamics Constellation (GDC). We present preliminary results for a simple, generic, and data-driven method of downlink prioritization for magnetic field (and other) measurements. Specifically, Principal Components Analysis (PCA) and One-Class Support Vector Machines (OC-SVMs) are used to detect intervals containing anomalous activity, which can then be prioritized for subsequent downlink. The computational simplicity of this algorithm makes it an excellent candidate for implementation on spaceflight hardware, as well as provide generalizability to a broad range of missions and data products. Initial analysis of this technique has been performed using magnetic field measurements from the Magnetospheric Multiscale Mission and CASSIOP, where it automatically identified scientifically interesting intervals containing Alfvén waves and EMIC activity.

Matthew G. Finley↗

Anomaly Detection for the Roman Space Telescope Wide Field Instrument’s Science Data Processing Pipeline

The Roman Space Telescope (RST) Wide Field Instrument (WFI) will be utilizing a preliminary Science Data Processing (SDP) pipeline during its Integration and Test, and to some extent during Operations, to track basic statistics and identify known features such as cosmic rays, snowballs as well as possible anomalies in raw detector data. In our detectors, these anomalies appear as jumps in the ramp of a readout and are classified as cosmic rays if they appear as a streak or snowballs if they’re more circular. The WFI employs an array of 18 H4RG-10 detectors that collect image samples. Each set of raw frames within a non-destructive exposure is packaged by the SDP pipeline into image cubes for each detector. Each cube is a time series of 4096 × 4096 accumulating pixel frames. The preliminary analysis pipeline is used to locate anomalies in these time-series accumulation frames and identify the type of anomaly, either natural phenomena or detector characteristic. To compare different methods, we’ve implemented both heuristic-based and data-driven methods to identify anomalies. For the heuristic-based approach, we identify snowballs and cosmic rays by the size and shape of outlier pixel clusters between consecutive frames. For data driven methods, we evaluated a Convolutional Neural Network (CNN) model, and more traditional methods like Principal Component Analysis (PCA). CNN is a supervised learning/classification method. Thus, we used a labeled dataset of anomalies to perform segmentation of the image and identify anomalies. We used previously identified cosmic rays and snowballs to measure the accuracy and efficiency of the mentioned approaches. In evaluating these methods, we aim to pick the best fit for the SDP pipeline’s anomaly detection in terms of both performance and runtime.

Paul Horton↗

DROP DURABILITY ASSESSMENT OF ELECTRONIC ASSEMBLIES UNDER OFF-AXIS LOADING WITH SKEWED FIXTURES

This thesis studies drop durability of electronic assemblies when the acceleration vector is oriented at 45° to the out-of-plane direction of the circuit card. The off-axis drop tests are accomplished with a skewed fixture and are conducted as a proxy for multiaxial drop testing. Advanced shock testing and vibration test methods have been developed over the last few decades to better represent real-world field environments during ground-based laboratory testing. However, many of these test methods require expensive and specialized equipment not available in most laboratories. An alternative approach for approximating simultaneous loading along multiple axes on conventional equipment utilizes skewed fixtures which have seen use in off-axis random vibration and drop impact testing. These methods generally rely on the conversion of a uniaxial input load from the test equipment (using a uniaxial drop tower or shaker) into a multiaxial load when resolved in the reference frame of the test article (mounted on a skewed fixture). Skewed fixture design is presented and recommendations for conducting skewed angle drop testing are introduced based on local measurements along the skewed face of the fixture to accurately monitor the impact event. Characterization tests were performed with a skewed fixture, at simultaneous acceleration loads from 500 to 3,000 g in two (in-plane and out-of-plane) directions, while meeting standard time domain tolerances. Upon experimental characterization, drop shock durability tests were conducted on a printed circuit assembly (PCA). Mean drops-to-failure were measured and quantified with Weibull statistics. Dominant solder joint failure modes were identified via failure analysis. Prior work on inclined angle impact testing is limited, and the majority of solder joint interconnect level fatigue studies are conducted considering perpendicular loading normal the circuit card. Low-cycle fatigue curves are generated based on plastic strain and plastic work density within the solder joint. A multiscale nonlinear finite element model is used to relate board-level flexure to solder joint interconnect level plastic strain. A high strain rate solder constitutive model allows for accurate modeling of solder plasticity resulting from high-impact drop shock. Fatigue parameters are computed from the Coffin-Manson relation and Palmgren-Miner damage accumulation. This work serves to apply established low-cycle fatigue methods for conventional drop shock loading (impact normal to circuit card) to non-perpendicular loading with a skewed fixture.

Hower, Jonathan [Kansas City National Security Cam↗

Investigation of acoustic waves under subsurface conditions to improve the predictions of rock mechanical properties and natural fracture characteristics

Mechanical properties and natural fracture characteristics are critical to investigate for subsurface engineering applications, including carbon storage, well drilling, and stimulation, as they govern rock stability, fluid flow, and mechanical behavior under stress. This dissertation integrates experimental and machine learning approaches to enhance the prediction and understanding of these properties by analyzing acoustic wave behavior under varied subsurface conditions. First, the influence of temperature, pore pressure, and supercritical CO2 (scCO2) saturation on poroelastic properties is examined using Gray Berea sandstone samples. The results show that temperature and pore pressure significantly affect the bulk modulus and Biot’s coefficient, while scCO2 saturation impacts rock compressibility, informing strategies for effective geological carbon storage. The study extends this understanding by experimentally evaluating the impact of reservoir depletion on the dynamic mechanical properties of the emerging Caney shale in South Oklahoma with the employment of unsupervised machine learning to predict static mechanical properties across the Caney shale. Integrating petrophysical data and chemostratigraphy, the workflow—featuring K-means clustering, principal component analysis (PCA), and inverse distance weighting (IDW)—improves stratigraphic characterization and the estimation of static-to-dynamic modulus ratios, which is vital for optimizing drilling and stimulation strategies. Finally, the work explores how natural fracture characteristics in shale influence acoustic waveforms and shear wave splitting (SWS) analysis. Experimental data on fractured samples under different stress and temperature conditions, combined with machine learning models such as K-nearest neighbors (KNN) and extreme gradient boosting (XGBoost), reveal key fracture properties impacting SWS and wave propagation. Together, these studies provide a comprehensive framework for linking acoustic wave behavior with rock properties, advancing the methods for monitoring and predicting geomechanical changes. The insights offered valuable implications for safer, more efficient CO2 injection, hydrocarbon extraction, and subsurface management.

Elkholy, Sherif↗