Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data Science and Machine Learning in Education

The growing role of data science (DS) and machine learning (ML) in high-energy physics (HEP) is well established and pertinent given the complex detectors, large data, sets and sophisticated analyses at the heart of HEP research. Moreover, exploiting symmetries inherent in physics data have inspired physics-informed ML as a vibrant sub-field of computer science research. HEP researchers benefit greatly from materials widely available materials for use in education, training and workforce development. They are also contributing to these materials and providing software to DS/ML-related fields. Increasingly, physics departments are offering courses at the intersection of DS, ML and physics, often using curricula developed by HEP researchers and involving open software and data used in HEP. In this white paper, we explore synergies between HEP research and DS/ML education, discuss opportunities and challenges at this intersection, and propose community activities that will be mutually beneficial.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

Data Science and Machine Learning for Genome Security

This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.

59 BASIC BIOLOGICAL SCIENCES↗

Revolutionizing Energetic Materials Discovery and Design: The Role of Data Science and Machine Learning

Here this Special Issue of Propellants, Explosives, Pyrotechnics (PEP) is focused on energetic materials discovery and design using Data Science and Machine Learning (DS&ML). The application of DS&ML has proven to be transformative in many areas, where it has been shown to expedite analysis, enable extraction of greater quantities of information from datasets, and guide experiments. However, energetic materials and their applications present unique challenges that often hinder the use of standardized tools and practices. In spite of these challenges, important and compelling advancements are being made toward data-directed research in energetics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning and Data Science to Advance Laboratory Earthquake Prediction and Illuminate the Mechanics of Precursors to Failure

Earthquakes represent one of our greatest natural hazards and in recent years human induced seismicity is adding to the threat. Even a modest improvement in the ability to forecast devastating large earthquakes or smaller shallow events associated with fluid injection could save thousands of lives and billions of dollars. Current efforts to forecast earthquakes are limited by knowledge of earthquake physics and hampered by a lack of reliable lab or field observations. However, recent work has provided a critical opportunity for advancement. We have found: 1) clear and consistent precursors prior to earthquake-like failure in the laboratory and 2) that lab earthquakes can be predicted using machine learning (ML). These works show that stick-slip failure events –the lab equivalent of earthquakes– are preceded by a cascade of micro-failure events that radiate elastic energy in a manner that foretells catastrophic failure. Remarkably, ML predicts the fault zone stress state, the failure time and in some cases the magnitude of lab earthquakes. In addition, the observations include clear precursors to failure in the form of changes in fault zone properties prior to lab earthquakes. Precursors have been observed in previous laboratory studies but their origin is poorly understood and their possible connection to ML based earthquake prediction is unknown. The work conducted under our project has dramatically expanded these efforts. We have developed an integrated data science approach to illuminate the physics of earthquake precursors and lab earthquake prediction. Our work has accelerated the development of ML, artificial intelligence (AI), and related data science approaches by providing massive data sets that are tightly connected to critical scientific problems and by bringing together leading subject matter experts and data scientists. Earthquake physics involves phenomena that are far from equilibrium. Our work has leveraged data science methods to illuminate these phenomena and investigate how they relate to earthquake prediction. In addition to a large database with many types of labeled events that is available to everyone, our work has advanced the fundamental understanding of seismic forecasting, earthquake physics, and fault rheology

58 GEOSCIENCES↗

Data Science in Chemical Engineering: Applications to Molecular Science

Chemical engineering is being rapidly transformed by the tools of data science. On the horizon, artificial intelligence (AI) applications will impact a huge swath of our work, ranging from the discovery and design of new molecules to operations and manufacturing and many areas in between. Early adoption of data science, machine learning, and early examples of AI in chemical engineering has been rich with examples of molecular data science—the application tools for molecular discovery and property optimization at the atomic scale. Here, we summarize key advances in this nascent subfield while introducing molecular data science for a broad chemical engineering readership. We introduce the field through the concept of a molecular data science life cycle and discuss relevant aspects of five distinct phases of this process: creation of curated data sets, molecular representations, data-driven property prediction, generation of new molecules, and feasibility and synthesizability considerations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bridging the length scales in ionic separations via data-driving machine learning

We pursued a data science driven machine learning (ML) approach that blended molecular scale attributes informed from molecular dynamics (MD) simulation and materials properties to the selectivity and energy efficiency in targeted ionic separations using electric fields. The model mixtures investigated for ionic separations are pH sensitive and include organic acids, silica and boron, transition metals, such as copper and chromium. There were two major research thrusts of this project. Firstly, we investigated surrogate models and deep learning that relate material chemistries and structures to selective transport of ionic species under applied electric fields. Secondly we investigated how the bipolar junction interfacial design and water dissociation catalyst in bipolar membranes affect reverse bias polarization behavior and pH modulation in deionization platforms as a function of the platform operating parameters (e.g., cell voltage, residence time, and salt feed concentration). As a result of this work, we also were able to start a new direction, namely ML models for molecular design of surfactants.

36 MATERIALS SCIENCE↗

A baseline structure inventory with critical attribution for the US and its territories

Leveraging high performance computing, remote sensing, geographic data science, machine learning, and computer vision, Oak Ridge National Laboratory has partnered with Federal Emergency Management Agency (FEMA) to build a baseline structure inventory covering the US and its territories to support disaster preparedness, response, and recovery. The dataset contains more than 125 million structures with critical attribution, and is ready to be used by federal agencies, local government and first responders to accelerate on-the-ground response to disasters, further identify vulnerable areas, and develop strategies to enhance the resilience of critical structures and communities. Data can be freely and openly accessed through Figshare data repository, ESRI’s Living Atlas or FEMA’s Geodata platform.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

New technologies as decision aids for the advancement of ecological risk assessment

Moore's law states that the number of transistors that can be placed on an integrated circuit doubles every two years (Moore, 1975). This has led to a steady increase in the processing power of computers over time, and technology is now enhancing and advancing software and scientific applications, which has enabled computationally intensive methods such as machine learning, data science, modeling, and simulation. The advancement of computers and data-driven algorithms is profoundly impacting people's lives. It is changing the way we work, the way we learn, and the way we interact with the world around us. Here, this editorial will discuss how scientists can benefit from the latest technology advancements and related tools by incorporating them into the ecological risk assessment (ERA) to study ecosystems as a way to create refined assessments and accelerate the turnaround times.

54 ENVIRONMENTAL SCIENCES↗

An Ensemble of Neural Networks for Moist Physics Processes, Its Generalizability and Stable Integration

Abstract With the recent advances in data science, machine learning has been increasingly applied to convection and cloud parameterizations in global climate models (GCMs). This study extends the work of Han et al. (2020, https://doi.org/10.1029/2020MS002076 ) and uses an ensemble of 32‐layer deep convolutional residual neural networks, referred to as ResCu‐en, to emulate convection and cloud processes simulated by a superparameterized GCM, SPCAM. ResCu‐en predicts GCM grid‐scale temperature and moisture tendencies, and cloud liquid and ice water contents from moist physics processes. The surface rainfall is derived from the column‐integrated moisture tendency. The prediction uncertainty inherent in deep learning algorithms in emulating the moist physics is reduced by ensemble averaging. Results in 1‐year independent offline validation show that ResCu‐en has high prediction accuracy for all output variables, both in the current climate and in a warmer climate with +4K sea surface temperature. The analysis of different neural net configurations shows that the success to generalize in a warmer climate is attributed to convective memory and the 1‐dimensional convolution layers incorporated into ResCu‐en. We further implement a member of ResCu‐en into CAM5 with real world geography and run the neural‐network‐enabled CAM5 (NCAM) for 5 years without encountering any numerical integration instability. The simulation generally captures the global distribution of the mean precipitation, with a better simulation of precipitation intensity and diurnal cycle. However, there are large biases in temperature and moisture in high latitudes. These results highlight the importance of convective memory and demonstrate the potential for machine learning to enhance climate modeling.

Meteorology & Atmospheric Sciences↗

Foreword to special issue: Papers from the 63rd annual meeting of the APS Division of Plasma Physics, November 8–12, 2021

The 63rd annual meeting of the APS Division of Plasma Physics (DPP) was held on November 8–12, 2021 in Pittsburgh at the David Lawrence Convention Center with both a live (in person) component and a virtual component. Following guidance from an APS COVID task force, all in-person attendees were fully vaccinated and masked. More than 800 physicists attended, safely, in-person. With both virtual and on-site participants, discussions were lively, and the research presentations showed unmatched mastery in the modern observation, theory, simulation, and manipulation of plasma. The presentations included four invited review talks, 97 invited talks, four tutorials, and four presentations from this year's prize and award recipients. There were more than 1200 contributed poster presentations and 725 contributed oral presentations. Including both in-person and remote attendees, DPP 2021 had a record of 2232 participants. As a hybrid meeting, in-person presentations of all invited presentations were broadcast live and were accompanied by a Q&A discussion. Contributed oral and poster presentations were prerecorded along with options to schedule in-person discussions on demand. Five mini-conferences were held: “Gatekeeper Workshop: Creating a Diverse, Equitable, and Inclusive Pipeline,” “Collisionless Shocks in Laboratory and Space Plasmas,” “The High Repetition Rate Frontier in High-Energy-Density Physics,” “Measuring and Modeling Plasma Surface Interactions,” and “The Second Mini-Conference on Machine Learning, Data Science and Artificial Intelligence in Plasma Research.” Finally, on the day before the official start of the meeting, an afternoon “for students, by students” included lightning talks, plasma trivia, and an informal occasion to connect with other students, learn how to get the most from the DPP Annual Meeting, and share successful ways to connect with colleagues and advance their professional careers.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

kynema-fmb [SWR-23-07]

Kynema-FMB (FKA: Kynema) is an open-source performance portable flexible multibody (FMB) dynamics solver designed for time-domain simulations. While originally tailored for wind turbine structural dynamics, the formulation and implementation are those of a general flexible-multidbody dynamics solver that can readily be applied to a wide range of systems. Kynema was designed with a narrow focus, namely to provide a lightweight, fast, accurate FMD solver for coupling to computational-fluid-dynamics (CFD) codes, especially the CFD codes in the Kynema suite, for fluid-structure-interaction (FSI) simulations. Kynema-FMB is equipped to model systems that can be represented as a collection of beams and rigid bodies that are connected through constraints. Degrees of freedom are defined in the inertial/global frame of reference and include displacements and rotations (formally as rotation matrices, but stored as quaternions). The underlying formulation is built on a Lie-group time integrator designed for index-3 differential-algebraic equations, which is second-order accurate in time (Bruls et al., 2012). Beam models are based on geometrically exact beam theory and are discretized as high-order spectral finite elements similar to those in BeamDyn (Wang et al., 2017). The governing equations for a FMD system like a wind turbine constitute a highly nonlinear system of constrained partial-differential equations. Kynema-FMB uses analytical Jacobians in the nonlinear-system solves in each time step. Linear systems use sparse storage and several third-party sparse-linear-system solvers are enabled. Ill conditioning of linear systems is mitigated with preconditioning described in Bottasso et al, 2008. Kynema-FMB is integrated with a simple open-source controller (ROSCO). There is an application programming interface (API) for coupling to geometry-resolved CFD (like that in Sharma et al., 2023) and actuator-force CFD (like that in Kuhn et al., 2025). In the latter, for actuator-line models, Kynema-FMB includes an internal blade-element solver that depends on user-provided lookup tables for coefficients of lift and drag, i.e., aerodynamic polars. Kynema-FMB is written in C++ and leverages Kokkos and Kokkos-Kernels (KokkosEcosystem) as its performance portability layer enabling simulations on both CPU and GPU systems. The repository is equipped with extensive automated testing at the unit and regression/system levels. The following describes the high-level development objectives conceived for Kynema: *Kynema will follow modern software development best practices, including test-driven development (TDD), version control, hierarchical automated testing, and continuous integration (CI) for a robust development environment. *The core data structures are memory efficient and enable vectorization and parallelization at multiple levels. *Data structures are data-oriented to exploit methods for accelerated computing including high utilization of chip resources (e.g., single instruction multiple data (SIMD) instruction sets) and parallelization using GP-GPUs. *The computational algorithms incorporate robust open-source libraries for mathematical operations, resource allocation, and data management. *The API design considers multiple stakeholder needs and ensure integration with existing and future ecosystems for data science, machine learning, and AI. *Kynema-FMB is written in modern C++ and leverages Kokkos as its performance-portability library with inspiration from the kynema stack.

Sprague, MichaelA.↗

Opportunities for Process Intensification with Membranes to Promote Circular Economy Development for Critical Minerals

Critical minerals are essential to the future of clean energy, especially energy storage, electric vehicles, and advanced electronics. In this paper, we argue that process systems engineering (PSE) paradigms provide essential frameworks for enhancing the sustainability and efficiency of critical mineral processing pathways. As a concrete example, we review challenges and opportu-nities across material-to-infrastructure scales for process intensification (PI) with membranes. Within critical mineral processing, there is a need to reduce environmental impact, especially con-cerning chemical reagent usage. Feed concentrations and product demand variability require flex-ible, intensified processes. Further, unique feedstocks require unique processes (i.e., no one-size-fits-all recycling or refining system exists). Membrane materials span a vast design space that allows significant optimization. Therefore, there is a need to rapidly identify the best opportunities for membrane implementation, thus informing materials optimization with process and infrastructure scale performance targets. Finally, scale-up must be accelerated and de-risked across the materials-to-process levels to fully realize the opportunity presented by membranes, thereby fostering the development of a circular economy for critical minerals. Tackling these challenges requires integrating efforts across diverse disciplines. We advocate for a holistic molecular-to-systems perspective for fully realizing PI with membranes to address sustainability challenges in critical mineral processing. The opportunities for PI with membranes are excellent applications for emerging research in machine learning, data science, automation, and optimization.

Dougher, Molly↗

AMIA KDDM Working Group Collaborative Workshop: Enriching Electronic Health Records with Social Determinants of Health to Improve Outcomes and Health Equity

Prior research has demonstrated that social determinants of health (SDoH) are major drivers of health outcomes and contributors to widespread health inequities. It was estimated that, in the United States, SDoH could be responsible for up to 40% of all preventable deaths, significantly higher than the 10-15% for which better medical care is responsible. Public health interventions that target SDoH are instrumental for improving health outcomes and reducing long-standing health inequities. Currently, most mainstream EHR vendors have implemented SDoH screeners in their EHR systems. However, the utility of the screeners is low, rendering patient-level SDoH still widely unavailable in the structured fields. SDoH are sometimes mentioned in free-text clinical notes (e.g., social context section) where natural language processing (NLP) can be applied to extract relevant information. Contextual-level SDoH can be identified from multiple data sources, many of which are publicly available and spatiotemporally linked to EHR data. As such, there is an opportunity for the KDDM research community to create innovative solutions to draw meaningful insights by creating and using rich data with SDoH to improve health outcomes while reducing disparities. In this workshop organized by AMIA Knowledge Discovery and Data Mining Working Group (AMIA KDDM WG), we will invite world-leading experts from academia, national laboratories, and life science industry with varied backgrounds in biomedical informatics, epidemiology, data science, machine learning, natural language processing, and pediatric cardiology to discuss the best practice of capturing, standardizing, and using SDoH information in various applications aiming at improving outcomes and health equity.

He, Zhe↗

Final Report (October 2024): University of Tennessee, Knoxville (UTK) contribution to: FusMatML: Machine Learning Atomistic Modeling for Fusion Materials Collaborative Project led by Dr. Aidan Thompson, Sandia National Laboratory

The rapid growth of the field of Machine Learning Inter-Atomic Potentials (MLIAP) has lead to a profusion of methods, all of which have some similarity to each other, but each also restricted to particular design choices, often arrived at in a rather ad hoc fashion. Beyond anecdotal evidence, and some benchmarking studies on specific problems, little progress has been made in developing design principles for MLIAPs. The goal of this project is to use machine learning, data science, and uncertainty quantification methods to optimize the design choices for MLIAP.

Density functional theory, Helium and Hydrogen↗