Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Multivariate Statistical Analysis Software Technologies for Astrophysical Research Involving Large Data Bases

We developed a package to process and analyze the data from the digital version of the Second Palomar Sky Survey. This system, called SKICAT, incorporates the latest in machine learning and expert systems software technology, in order to classify the detected objects objectively and uniformly, and facilitate handling of the enormous data sets from digital sky surveys and other sources. The system provides a powerful, integrated environment for the manipulation and scientific investigation of catalogs from virtually any source. It serves three principal functions: image catalog construction, catalog management, and catalog analysis. Through use of the GID3* Decision Tree artificial induction software, SKICAT automates the process of classifying objects within CCD and digitized plate images. To exploit these catalogs, the system also provides tools to merge them into a large, complex database which may be easily queried and modified when new data or better methods of calibrating or classifying become available. The most innovative feature of SKICAT is the facility it provides to experiment with and apply the latest in machine learning technology to the tasks of catalog construction and analysis. SKICAT provides a unique environment for implementing these tools for any number of future scientific purposes. Initial scientific verification and performance tests have been made using galaxy counts and measurements of galaxy clustering from small subsets of the survey data, and a search for very high redshift quasars. All of the tests were successful and produced new and interesting scientific results. Attachments to this report give detailed accounts of the technical aspects of the SKICAT system, and of some of the scientific results achieved to date. We also developed a user-friendly package for multivariate statistical analysis of small and moderate-size data sets, called STATPROG. The package was tested extensively on a number of real scientific applications and has produced real, published results.

Djorgovski, S. G.↗

Deep-learning based artificial intelligence tool for melt pools and defect segmentation

Accelerating fabrication of additively manufactured components with precise microstructures is important for quality and qualification of built parts, as well as for a fundamental understanding of process improvement. Accomplishing this requires fast and robust characterization of melt pool geometries and structural defects in images. This paper proposes a pragmatic approach based on implementation of deep learning models and self-consistent workflow that enable systematic segmentation of defects and melt pools in optical images. Deep learning is based on an image-to-image translation–conditional generative adversarial neural network architecture. An artificial intelligence (AI) tool based on this deep learning model enables fast and incrementally more accurate predictions of the prevalent geometric features, including melt pool boundaries and printing-induced structural defects. We present statistical analysis of geometric features that is enabled by the AI tool, showing strong spatial correlation of defects and the melt pool boundaries. The correlations of widths and heights of melt pools with dataset processing parameters show the highest sensitivity to thermal influences resulting from laser passes in adjacent and subsequent layer passes. The presented models and tools are demonstrated on the aluminum alloy and datasets produced with different sets of processing parameters. However, they have universal quality and could easily be adapted to different material compositions. The method can be easily generalized to microstructural characterizations other than optical microscopy.

additive manufacturing↗

Machine Learned Force Field Modeling of Metal Organic Frameworks for CO2 Direct Air Capture

Metal organic frameworks (MOFs) are a large class of porous materials and have garnered significant interest due to their large surface areas and their tunable physical and chemical properties. Numerous prior studies have been performed to screen large databases of this material class for promising DAC sorbent materials. These studies have often relied on classical model potentials. While density functional theory (DFT) calculations have been shown to be very accurate for modeling the interaction of CO2 with MOFs, such calculations are too computationally demanding for statistically significant adsorption predictions. To overcome this barrier, we developed methods for training models to achieve DFT-level accuracy for the forces and energies associated with MOF flexibility and CO2 adsorption using machine learned force fields (MLFFs). These methods were parametrized based on DFT calculations of CO2 in a flexible MOF and used to predict MOF structural properties as well as CO2 adsorption in several MOFs.

Findley, John↗

New Neural Network Cloud Mask Algorithm Based on Radiative Transfer Simulations

Cloud detection and screening constitute critically important first steps required to derive many satellite data products. Traditional threshold-based cloud mask algorithms require a complicated design process and fine tuning for each sensor, and they have difficulties over areas partially covered with snow/ice. Exploiting advances in machine learning techniques and radiative transfer modeling of coupled environmental systems, we have developed a new, threshold-free cloud mask algorithm based on a neural network classifier driven by extensive radiative transfer simulations. Statistical validation results obtained by using collocated CALIOP and MODIS data show that its performance is consistent over different ecosystems and significantly better than the MODIS Cloud Mask (MOD35 C6) during the winter seasons over snow-covered areas in the mid-latitudes. Simulations using a reduced number of satellite channels also show satisfactory results, indicating its flexibility to be configured for different sensors. Comparedto threshold-based methods and previous machine-learning approaches, this new cloud mask (i) does not rely on thresholds, (ii) needs fewer satellite channels, (iii) has superior performance during winter seasons in mid-latitude areas, and (iv) can easily be applied to different sensors.

cloud mask algorithms↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Absorption dissymmetry factor enhancement: A data-driven approach to unravel the synthesis knobs of chiral 2D perovskites

Chiral 2D metal halide perovskites (MHPs) are promising for spin-optoelectronic applications, yet their absorption dissymmetry factor (g abs ) exhibits significant variability due to complex, co-dependent structural and experimental factors. Here, we established a data-driven framework using Pearson’s correlation, ANOVA, and Gaussian process regression to identify and model key synthesis “knobs” governing these properties. The analysis revealed that solvent choice is the primary factor driving variability. For acetonitrile-based films, g abs was maximized by optimizing annealing temperature and film thickness. Conversely, films from higher boiling point solvents showed complex dependencies on annealing temperature, excitonic integral intensity, and film texture. These statistical correlations provide a roadmap for the rational design of high-performance chiral MHPs and establish a foundation for future machine learning-driven material exploration.

ANOVA↗

A Probabilistic Model for Global EMIC Wave Activity Using Van Allen Probes Observations

Electromagnetic ion cyclotron (EMIC) waves play a key role in radiation belt dynamics through resonant interactions. However, their low occurrence probability, high variability, and spatial intermittency pose challenges for accurate modeling. In this study, we present a machine learning (ML)-based global EMIC wave model built on the entire data set from the Van Allen Probes mission. To capture the distinct statistical characteristics of wave occurrence and amplitude, the model is separated into two modules: an occurrence model trained using ML techniques, and a wave amplitude model sampled from observed probability distributions. The input parameters are limited to real-time or predictable variables to ensure practical applicability. Our model shows strong performance across the entire test set and demonstrates improved predictive capability over a baseline random occurrence model, particularly during quiet geomagnetic conditions. Evaluation during both quiet and active periods confirms the model's ability to represent the clustered and intermittent nature of EMIC wave activity. Furthermore, the model provides global estimates of wave power, enabling integration with radiation belt electron data and showing signatures consistent with wave-induced scattering. We found a good correlation between the global wave activity from the model and relativistic electron observation by Van Allen Probes, regardless of the availability of in situ wave observations. The modular structure of the model also allows for straightforward expansion for additional wave properties, such as wave frequency, which can be modeled independently. This flexible, event-sensitive approach offers a promising framework for data-driven radiation belt simulations and space weather applications.

79 ASTRONOMY AND ASTROPHYSICS↗

A unified ensemble soil moisture dataset across the continental United States

Abstract A unified ensemble soil moisture (SM) package has been developed over the Continental United States (CONUS). The data package includes 19 products from land surface models, remote sensing, reanalysis, and machine learning models. All datasets are unified to a 0.25-degree and monthly spatiotemporal resolution, providing a comprehensive view of surface SM dynamics. The statistical analysis of the datasets leverages the Koppen-Geiger Climate Classification to explore surface SM’s spatiotemporal variabilities. The extracted SM characteristics highlight distinct patterns, with the western CONUS showing larger coefficient of variation values and the eastern CONUS exhibiting higher SM values. Remote sensing datasets tend to be drier, while reanalysis products present wetter conditions. In-situ SM observations serve as the basis for wavelet power spectrum analyses to explain discrepancies in temporal scales across datasets facilitating daily SM records. This study provides a comprehensive soil moisture data package and an analysis framework that can be used for Earth system model evaluations and uncertainty quantification, quantifying drought impacts and land–atmosphere interactions and making recommendations for drought response planning.

54 ENVIRONMENTAL SCIENCES↗

Wave propagation and earth satellite radio emission studies

Radio propagation studies of the ionosphere using satellite radio beacons are described. The ionosphere is known as a dispersive, inhomogeneous, irregular and sometimes even nonlinear medium. After traversing through the ionosphere the radio signal bears signatures of these characteristics. A study of these signatures will be helpful in two areas: (1) It will assist in learning the behavior of the medium, in this case the ionosphere. (2) It will provide information of the kind of signal characteristics and statistics to be expected for communication and navigational satellite systems that use the similar geometry.

Yeh, K. C.↗

Summer High School Apprenticeship Research Program (SHARP)

The summer of 1997 will not only be noted by NASA for the mission to Mars by the Pathfinder but also for the 179 brilliant apprentices that participated in the SHARP Program. Apprentice participation increased 17% over last year's total of 153 participants. As indicated by the End-of-the-Program Evaluations, 96% of the programs' participants rated the summer experience from very good to excellent. The SHARP Management Team began the year by meeting in Cocoa Beach, Florida for the annual SHARP Planning Conference. Participants strengthened their Education Division Computer Aided Tracking System (EDCATS) skills, toured the world-renowned Kennedy Space Center, and took a journey into space during the Alien Encounter Exercise. The participants returned to their Centers with the same goals and objectives in mind. The 1997 SHARP Program goals were: (1) Utilize NASA's mission, unique facilities and specialized workforce to provide exposure, education, and enrichment experiences to expand participants' career horizons and inspire excellence in formal education and lifelong learning. (2) Develop and implement innovative education reform initiatives which support NASA's Education Strategic Plan and national education goals. (3) Utilize established statistical indicators to measure the effectiveness of SHARP's program goals. (4) Explore new recruiting methods which target the student population for which SHARP was specifically designed. (5) Increase the number of participants in the program. All of the SHARP Coordinators reported that the goals and objectives for the overall program as well as their individual program goals were achieved. Some of the goals and objectives for the Centers were: (1) To increase the students' awareness of science, mathematics, engineering, and computer technology; (2) To provide students with the opportunity to broaden their career objectives; and (3) To expose students to a variety of enrichment activities. Most of the Center goals and objectives were consistent with the overall program goals. Modem Technology Systems, Inc., was able to meet the SHARP Apprentices, Coordinators and Mentors during their site visits to Stennis Space Center, Ames Research Center and Dryden Flight Research Center. All three Centers had very efficient programs and adhered to SHARP's general guidelines and procedures. MTSI was able to meet the apprentices from the other Centers via satellite in July during the SHARP Video-Teleconference(ViTS). The ViTS offered the apprentices and the NASA and SHARP Coordinators the opportunity to introduce themselves. The apprentices from each Center presented topical "Cutting Edge Projects". Some of the accomplishments for the 1997 SHARP Program year included: MTSI hiring apprentices from four of the nine NASA Centers, the full utilization of the EDCATS by apprentices and NASA/SHARP Coordinators, the distribution of the SHARP Apprentice College and Scholarship Directory, a reunion with former apprentices from Langley Research Center and the development of a SHARP Recruitment Poster. MTSI developed another exciting newsletter containing graphics and articles submitted by the apprentices and the SHARP Management Team.

Source record↗

Understanding Oceanic Heavy Precipitation Using Scatterometer, Satellite Precipitation, and Reanalysis Products

The primary aim of this study is to understand the heavy precipitation events over Oceanic regions using vector wind retrievals from space based scatterometers in combination with precipitation products from satellite and model reanalysis products. Heavy precipitation over oceans is a less understood phenomenon and this study tries to fill in the gaps which may lead us to a better understanding of heavy precipitation over oceans. Various phenomenon may lead to intense precipitation viz. MJO (Madden-Julian Oscillation), Extratropical cyclones, MCSs (Mesoscale Convective Systems), that occur inside or outside the tropics and if we can decipher the physical mechanisms behind occurrence of heavy precipitation, then it may lead us to a better understanding of such events which further may help us in building more robust weather and climate models. During a heavy precipitation event, scatterometer wind observations may lead us to understand the governing dynamics behind that event near the surface. We hypothesize that scatterometer winds can observe significant changes in the near-surface circulation and that there are global relationships among these quantities. To the degree to which this hypothesis fails, we will learn about the regional behavior of heavy precipitation-producing systems over the ocean. We use a "precipitation feature" (PF) approach to enable statistical analysis of a large database of raining features.

Winds↗

Landslide Hazard is Projected to Increase Across High Mountain Asia

High Mountain Asia has long been known as a hotspot for landslide risk, and studies have suggested that landslide hazard is likely to increase in this region over the coming decades. Extreme precipitation may become more frequent, with a nonlinear response relative to increasing global temperatures. However, these changes are geographically varied. This article maps probable changes to landslide hazard, as shown by a landslide hazard indicator (LHI) derived from downscaled precipitation and temperature. In order to capture the nonlinear response of slopes to extreme precipitation, a simple machine-learning model was trained on a database of landslides across High Mountain Asia to develop a regional LHI. This model was applied to statistically downscaled data from the 30 members of the Seamless System for Prediction and Earth System Research large ensembles to produce a range of possible outcomes under the Shared Socioeconomic Pathways 2-4.5 and 5-8.5. The LHI reveals that landslide hazard will increase in most parts of High Mountain Asia. Absolute increases will be highest in already hazardous areas such as the Central Himalaya, but relative change is greatest on the Tibetan Plateau. Even in regions where landslide hazard declines by year 2100, it will increase prior to the mid-century mark. However, the seasonal cycle of landslide occurrence will not change greatly across High Mountain Asia. Although substantial uncertainty remains in these projections, the overall direction of change seems reliable. These findings highlight the importance of continued analysis to inform disaster risk reduction strategies for stakeholders across High Mountain Asia.

Thomas A Stanley↗

Critical statistical assessment of data in metal additive manufacturing

Obtaining high quality data reflecting the relationships between the additive manufacturing (AM) process parameters, material microstructure and mechanical properties is crucial for the use of machine learning in AM. A database of over 4,000 data entries of metal AM was created thanks to a large number of literature studies on key process parameters and indicators of build quality. Meta-analysis reveals critical biases in the literature. Firstly, majority of studies report only high quality builds, these imbalances in reporting result in weak correlation between process parameters, properties and consolidation, limiting the ability of machine learning models to generalize beyond optimized conditions. Nevertheless, the trained models accurately predict yield strength ($R^2 = 0.85$), suggesting that certain process–property relationships are effectively captured within these models. Secondly, quantitative microstructural data are largely absent, limiting the learning of the microstructure-mechanical properties relationships. Finally, current process window identification is based largely on the consolidation, despite significant uncertainty in its measurement. It is important to identify the process map on the basis of not only the consolidation, but also mechanical behaviour under loading. Such a identification shows that 316 L and Inconel have much larger process map (i.e. highly printable) in comparison to the AlSi10Mg and Ti6Al4V.

Additive manufacturing↗

Evidence of short chains in liquid sulfur

High energy x-ray pair distribution function measurements show the average coordination number of the first shell in liquid sulfur is 1.86 ± 0.04 across the λ-transition, not precisely 2.0 as widely accepted. This indicates that upon melting, liquid sulfur does not comprise solely of S 8 rings but also possesses a significant number of short chains. Intensities of the pre-peak and first diffraction peak of the x-ray structure factor and third peak height of the pair distribution function all show deviations at the λ-transition temperature T λ , associated with the break-up of S 8 rings and the start of oligomer polymerization. A significant number of non-bonded or loosely bonded “interstitial atoms,” with an average coordination number of 0.20 ± 0.005, are also observed in the so-called “forbidden zone” between the first and second shells upon melting. The number of interstitial atoms is found to decrease to a minimum at the λ-transition, but the majority persist into the high temperature polymerized liquid. Furthermore, the existence of short chains and nearby interstitial atoms represent the two main factors required to initiate the S 8 -ring to chain transition, as proposed by recent molecular dynamics simulations.

Chemical bonding↗

Data-based filtered dissipation rate modelling for multi-modal turbulent combustion: evaluating a priori model generalizability

Manifold-based models offer a computationally efficient alternative to directly transporting the thermochemical state in computational simulations of turbulent reacting flows, projecting the high-dimensional thermochemical state-space onto a low-dimensional manifold. Recent efforts have yielded a manifold-based model applicable to multi-modal combustion, enabling reconstruction of the thermochemical state from solutions to two-dimensional manifold equations in mixture fraction and generalized progress variable that are parameterised by three scalar dissipation rates. In coarse-grained simulations such as Large Eddy Simulation (LES), closure of the multi-modal manifold equations and subfilter variances/covariance requires closure of three filtered scalar dissipation rates. Here, the present work adopts a data-based approach, providing closure for the three filtered scalar dissipation rates via deep neural networks (DNNs). High-fidelity datasets corresponding to an autoigniting n-dodecane jet flame and a bluff body swirl-stabilized confined lifted spray flame of two aviation fuels (Jet-A and C1) with different ignition propensities are leveraged to generate training data that spans a diverse range of thermodynamic conditions and combustion modes, including low- and high-temperature ignition regimes in addition to premixed and nonpremixed behaviour. A final DNN model is trained to enforce inherent physical constraints by learning nonlinear functional transformations of the three filtered scalar dissipation rates. The generalizability of this constrained DNN model is demonstrated a priori via conditional statistics evaluated on the lifted spray flame with C1–a configuration that had not been included in the training data. Excellent DNN agreement with conditional DNS statistics is observed, and integrated gradients are computed to identify the most sensitive input variables. The similarity of the marginal PDFs of the most informative input variables and outputs across configurations are quantified via the Wasserstein metric, demonstrating that data-based models may successfully generalize to unseen parametric conditions so long as the most informative input variables share similar distributions across training and testing datasets.

Data-based modelling↗

Battery Life Prediction Using Reduced-Order Physics Models and Machine Learning (CRADA Final Report)

Phase 1 (Original CRADA, plus no-cost extension modifications #1-3, 6/1/2017 to 3/13/2021): The Australian Department of Defence (AUDoD) is performing accelerated aging tests of Li-ion batteries to benchmark their reliability and degradation characteristics. Using its previously developed battery lifetime predictive model framework, the National Laboratory of the Rockies (NLR) will develop analytical models based the AUDoD data to predict lifetime of the multiple Li-ion battery chemistries under real-world use scenarios of interest to AUDoD. The NLR model is based on physical degradation mechanisms encountered by Li-ion batteries and has been previously validated. Phase 2 (CRADA modification #4, plus no-cost extension modification #5, 2/22/2021 to 3/30/2025): Train and support Australian Department of Defence personnel to use NLR software for model-based estimation of Li-ion battery lifetime using accelerated battery aging data collected by the Australian Department of Defence. Under separate DOE funding from 2019 to 2021, NLR enhanced its battery life-prediction software using machine learning algorithms to automate portions of the model-fitting process, requiring significantly less labor and expert judgment and also adding uncertainty quantification, increasing statistical rigor. Under Phase 2, NLR will customize NLR Software and provide it to AuDoD. NLR will enhance its NLR Model to capture aging modes of AuDoD's multi-cell modules, including cell-balancing effects. NLR will develop example single-cell and multi-cell models based on one AuDoD battery aging dataset. NLR will train AuDoD personnel on NLR Software. By the conclusion of the project, NLR will have provided AuDoD the training materials, a user manual and software needed to perform their own analysis of additional and/or future battery aging datasets.

33 ADVANCED PROPULSION SYSTEMS↗

Analyzing a Mature Software Inspection Process Using Statistical Process Control (SPC)

This paper presents a cooperative effort where the Software Engineering Institute and the Space Shuttle Onboard Software Project could experiment applying Statistical Process Control (SPC) analysis to inspection activities. The topics include: 1) SPC Collaboration Overview; 2) SPC Collaboration Approach and Results; and 3) Lessons Learned.

Barnard, Julie↗

Optimization of Airport Runway Configuration with Forecast-Augmented Offline Reinforcement Learning

Runway configuration Management (RCM) governs the optimal utilization of runways based on variables such as traffic and meteorological conditions, making it a daunting task in air traffic management due to its dependency on volatile operational and environmental factors. This paper improves upon our previous work [1] on using offline model-free reinforcement learning for creating a Runway Configuration Assistance (RCA) decision-support tool. A novel integration of forecast data from LAMP (Localized Aviation Model Output Statistics Program) and TAF (Terminal Area Forecast) is introduced, enhancing the tool’s accuracy and also its adaptability to quick wind changes. The performance is evaluated using two major US airports, Charlotte Douglas International Airport (CLT) and Denver International Airport (DEN). To counter scalability issues presented by the addition of discrete forecast variables, we transitioned to a continuous state space model, ensuring scalability and inclusion of longer forecast data. The results of our experiments reflect significant improvements in the RCA tool’s prediction accuracy.

Sumanth Nethi↗