Search NASA⌕ Search

SEARCH · Search NASA

Results for “Databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Quality Ranking of Unary Chloride Salt Property Data Included in MSTDB-TP

Molten salt reactor developers rely on thermal property data to design, license and operate the reactors. The Molten Salt Thermal Database-Thermophysical Properties (MSTDB-TP) was established under the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program and is managed by Oak Ridge National Laboratory to serve as a single source of thermophysical property values measured for a wide variety of molten salt systems for use by researchers, molten salt reactor developers, and regulators. These properties include density, viscosity and thermal diffusivity and conductivity. Published measurements of molten salt properties are lacking for many salts of interest and the data that are available are often inconsistent. This creates a challenge for MSR developers when determining which property values to use when designing their reactors. It is the purpose of this work to apply a consistent ranking system to all data entries that indicates the quality of property values listed in the database. These rankings will be the technical basis for down-selections by the database developers and alert users about the quality of the available property values. MSTDB-TP collects all available property data and indicates preferred data sets or correlations. However, all available data sets are included in the database. Quality assessments and rankings are being applied to data in MSTDB-TP to provide an indication of the quality of each data set independent of consistency with other data. Previous reports detailed the ranking system that was followed and assessments of unary fluoride data sets. Documentation of the quality of data in MSTDB-TP was continued by reviewing and assessing all available sources of density, viscosity and thermal diffusivity or conductivity values for unary chloride salts in MSTDB-TP V3.0 using the same criteria.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

CLEAP Project: OR-SAGE Analysis for MT, UT, and CO States

The OR-SAGE tool is designed to use industry-accepted practices in screening sites and then employ the proper array of data sources through the considerable computational capabilities of GIS technology available at ORNL. The tool was developed to screen the potential for NPP siting on a national and regional basis. However, because of the tool granularity, it is often focused specifically on the immediate area around user sites of interest. If data center siting parameters can be added to OR-SAGE, the ability to evaluate data center siting on a localized scale will be beneficial.1 More than 60 data sets have been collected and processed by ORNL to develop exclusionary, avoidance, and suitability criteria for screening sites for a variety of power generation types, including nuclear power plants. Available site evaluation parameters include population density, slope, seismic activity, proximity to cooling-water sources, proximity to hazard facilities, avoidance of protected lands and floodplains, susceptibility to landslide hazards, and many others. All siting parameters should be considered as flags to inform siting decisions and should not be used to rule in or rule out any NPP site. Once data center siting parameters are identified, appropriate data sets will be collected and processed. The OR-SAGE process is very versatile. Essentially, OR-SAGE is a visual, relational database. The database partitions the contiguous United States, a total of 720 million hectares (~1.8 billion acres), into 100-m by 100-m (1 hectare or ~2.5 acre) cells. The database is tracking just under 700 million individual land cells. Successive suitability criterion is applied to each cell in the database. User-specified thresholds can be applied to each siting parameter data layer. In this manner, a variety of scenarios can be quickly and thoroughly evaluated. Data can be added and/or revised within OR-SAGE to address user interests. Siting security assessment capability is currently being added to OR-SAGE. Security is expected to be of concern at data centers whether it is collocated with a nuclear power generating technology or not. If data center is collocated with a nuclear power generating source, the security threat attractiveness level of both will likely increase. It will be of additional benefit if a potential data center site is also assessed for security vulnerability.

97 MATHEMATICS AND COMPUTING↗

2024 Buildings Technology Baseline: Dataset Documentation

The Buildings Technology Baseline is a curated and regularly updated dataset of current and projected performance, retail, and installed price data for all major building energy technologies needed to enable cost/benefit analyses. Building technology analyses require an up-to-date understanding of installation costs and cost-effectiveness of key building energy efficiency technologies. The dataset was assembled by Guidehouse during fiscal year 2024. Data was gathered from the 2024 National Residential Efficiency Measures Database (NREMDB), the 2023 Energy Information Administration Updated Buildings Sector Appliance and Equipment Costs and Efficiencies ("EIA Building Data Report"), DOE Lighting Market Model, the 2023 RSMeans database, and the 2020 Grid-Interactive Efficient Building Technology Cost, Performance, and Lifetime Characteristics ("GEB Data Report"), Lawrence Berkeley National Laboratory, various literature, as well as new data from online retailers, stakeholder interviews, and contractor databases in 2023 and 2024. The dataset has been reviewed by subject matter experts at NREL and DOE. The 2024 dataset release is intended to be a starting point for interested users to provide feedback. This database is not intended to provide specific cost estimates for a specific project. The cost estimates do not include any rebates or tax incentives that may be available for the measures. Rather, it is meant to help determine which measures may be more cost-effective. The National Renewable Energy Laboratory (NREL) makes every effort to ensure accuracy of the data; however, NREL does not assume any legal liability or responsibility for the accuracy or completeness of the information.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

An Integrated ML/AI Framework for Digitizing, Structuring and Searching DOE U-TRU-Fuels Data with Gap Analysis of Non-DOE Records

The U.S. Department of Energy (DOE) Advanced Fuels Campaign (AFC) is advancing transmutation fuel technologies to reduce long-lived radioactive waste by converting minor actinides into shorter-lived or stable elements through irradiation in sodium-cooled fast reactors. Key experiments such as AFC-1, AFC-2, FUels for the transmutation of Trans-URanium elements In phéniX (FUTURIX)-Fortes Teneurs en Actinides (FTA), and Experimental Breeder Reactor-II (EBR-II) X501 have provided fuel fabrication, irradiation, and performance data on various transuranic-bearing fuel forms. This report documents the creation of an artificial-intelligence assisted database, which has consolidated all DOE-owned data related to Transuranic (TRU)-bearing fuel experiments and stored across it across both the Idaho National Laboratory (INL) Nuclear Data Management and Analysis System and the INL high performance computing (HPC) infrastructure. A dedicated webpage, hosted on the INL HPC system, has been developed to support role-based access and data interaction. The database architecture allows researchers to navigate large, heterogeneous archives with far greater speed and accuracy than manual search and lays the foundation for future expansion into multimodal nuclear materials analysis environments. The database represents a major step towards a nationally integrated fuels database utilizing artificial intelligence tools.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Report priority gaps in high temperature thermodynamic data (Interim Progress Report)

This interim progress report (Level 4 Milestone Number M4SF-26LL010203023) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite Host Rock Properties & Processes SF-26LL01020302. Our focus is to assess gaps in data availability and understanding for radionuclide thermodynamics within the context of a “hot repository” concept and expand SUPCRT-NE database development to address higher temperatures needed for a DPC DGR disposal concept. The database is intended to inform the argillite GDSA baseline model. The leading European thermochemical database (Thermochimie) is only applicable to temperatures below 80°C. Thus, a US effort to integrate and expand upon other international thermodynamics database efforts is needed, particular if a “hot repository” concept moves forward.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Seeing values for LSST strategy simulations

The opsim4 operations simulation program for the LSST astronomical survey uses a database of seeing values covering the range of times to besimulated. Idescribethe creation of such a database using Dual Image Motion Monitor(DIMM)datacollected at Cerro Pachon from 2004-03-17 to 2019-10-07. In times during which the data overlap, I compare the distribution of DIMM seeing values to the seeing measured in DECamimages,takenatasite 10kmaway. Becauseinstrumentalproblemsinthe DIMMmay indicate unreliablemeasurements,cutsonimagequality(asindicatedby the measured Strehlratio)wereexplored. TheDIMMhassignificantgaps,soImodel thedata(withandwithoutcutsonStrehlratio)andgenerateartificialdatainthegaps according to the model. The model consists of a sinusoidal variation with a period of one year, an autoregressive (AR1) model for variations in mean seeing from one night to the next, and another AR1 model for variations on a 5 minute timescale. I create four databases according to thisprocedure, twobasedonDIMMdatastarting 2006-01-01 (with and without a Strehl ratio cut), and two starting 2009-01-01. I then run opsim simulations using each, and an otherwise identical simulation using the default seeing database, and explore the differences

Neilsen, Eric H. [Fermilab]↗

Hydropower Infrastructure – LAkes, Reservoirs, and RIvers (HILARRI)

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2024) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2024) These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

13 HYDRO ENERGY↗

Digitizing and Enhancing Accessibility of the Fusion Safety Archives

This project focuses on the digitization and public accessibility to the Fusion Safety Archives at the Idaho National Laboratory. The first phase involves a thorough review of each document in the physical archives to determine its online availability. For documents that are available online, PDF copies and unique identifiers are collected for database integration. Documents not available online are delivered to Red Inc. for digitization. Additionally, defunct storage devices such as diskettes are sent to INL’s archival department for data retrieval where possible. The second phase of the project involves the creation of a comprehensive database to house the digital copies of the archives. The database will facilitate easy access and management of the digitized documents. Following the database creation, we plan to train a Retrieval-Augmented Generation (RAG) based AI on publicly available documents. The trained AI will be integrated into a front-facing application, allowing the public to easily access information from the Fusion Safety Archives. This project aims to preserve valuable historical data, improve accessibility, and promote transparency in fusion safety research.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗

HTESP (High-throughput electronic structure package): A package for high-throughput ab initio calculations

High-throughput ab initio calculations are the indispensable parts of data-driven discovery of new materials with desirable properties, as reflected in the establishment of several online material databases. The accumulation of extensive theoretical data through computations enables data-driven discovery by constructing machine learning and artificial intelligence models to predict novel compounds and forecast their properties. Efficient usage and extraction of data from these existing online material databases can accelerate the next stage materials discovery that targets different and more advanced properties, such as electron–phonon coupling for phonon-mediated superconductivity. However, extracting data from these databases, generating tailored input files for different ab initio calculations, performing such calculations, and analyzing new results can be demanding tasks. Here, in this work, we introduce a software package named “HTESP” (High-Throughput Electronic Structure Package) written in Python and Bash languages, which automates the entire workflow including data extraction, input file generation, calculation submission, result collection and plotting. Our HTESP will help speed up future computational materials discovery processes.

36 MATERIALS SCIENCE↗

CONFLUX: A standardized framework to calculate reactor antineutrino flux

Nuclear fission reactors are abundant sources of antineutrinos for neutrino physics experiments. The flux and spectrum of antineutrinos emitted by a reactor can indicate its activity and composition, suggesting potential applications of neutrino measurements beyond fundamental scientific studies that may be valuable to society. The utility of reactor antineutrinos for applications and fundamental science is dependent on the availability of precise predictions of these emissions. For example, in the last decade, disagreements between reactor antineutrino measurements and models have inspired revision of reactor antineutrino calculations and standard nuclear databases as well as searches for new fundamental particles not predicted by the Standard Model of particle physics. Past predictions and descriptions of the methods used to generate them are documented to varying degrees in the literature, with different modeling teams incorporating a range of methods, input data, and assumptions. The resulting difficulty in accessing or reproducing past models and reconciling results from differing approaches complicates the future study and application of reactor antineutrinos. The CONFLUX (Calculation Of Neutrino FLUX) software framework is a neutrino prediction tool built with the goal of simplifying, standardizing, and democratizing the process of reactor antineutrino flux calculations. CONFLUX includes three primary methods for calculating the antineutrino emissions of nuclear reactors or individual beta decays that incorporate common nuclear data and beta decay theory. The software is prepackaged with the current nuclear databases, including ENDF.B/VIII, JEFF-3.3, and ENSDF, and it includes the capability to predict time-dependent reactor emissions, adjust nuclear database or beta decay inputs/assumptions, and propagate related sources of uncertainty. Here, this paper describes the CONFLUX software structure, details the methods used for flux and spectrum calculations, and provides examples of potential use cases.

Zhang, Xianyi [Lawrence Livermore National Laborat↗

Mechanically induced thermal runaway severity analysis of Li-ion batteries and continuous energy release monitoring

The large-scale deployment of Li-ion batteries in stationary energy storage and electrical vehicle applications demands a strong focus on safety, particularly on the thermal runaway risk and severity evaluation. A standardized single-side mechanical indentation test protocol was developed to induce an internal short-circuit (ISC) and evaluate cells' thermal runaway severity at different state of charge (SOC). The observed hazard severity (OHS in five categories) and evaluated scores in this work have a comprehensive consideration of each cell's capacity, initial voltage, SOC, temperature and voltage change, allowing a better evaluation of the cells' thermal runaway potential. This method was applied to about 200 Li-ion batteries in order to build an extensive thermal runaway database covering various SOCs, capacities and chemistries. In this study, we monitored the transitions of stored electrochemical energy and applied mechanical energy into both thermal energy and acoustic emissions (AE). The surface temperature and mechanical failures were monitored by infrared imaging and AE to capture critical events within battery cells throughout the mechanical indentation tests. Furthermore, the initial temperature maps can predict two types of follow-up events: thermal runaway or gradual heat release via conduction. Analyzing each cell's severity, AEs, and leveraging the evolving database offer insights into predicting occurrences of thermal runaway. The test method, thermal runaway severity evaluation and prediction, and the corresponding database provide battery designers, manufacturers, and end-users a clear overview of Li-ion batteries' thermal runaway potential under mechanical abuse, advancing the safety design of Li-ion batteries.

Acoustic emission↗

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic in-context learning with conversational models for data extraction and materials property prediction

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs such as Google gemini-pro and OpenAI gpt-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies—enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall that exceed 95% with an error rate of ∼9%, highlighting the effectiveness and versatility of the toolkit. Finally, databases for 2D material thicknesses, a critical parameter for device integration, and energy bandgap values are developed using PropertyExtractor. In particular, for the thickness database, the rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of various material property databases, advancing the field.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Analysis of the impact of parallel magnetic fluctuations on linear gyrokinetic stability in NSTX-U and verification of gyro-fluid models

In this work, we use the CGYRO gyrokinetic code to analyze two L- and one H-mode discharges from the National Spherical Torus Experiment (NSTX) and NSTX-Upgrade (NSTX-U) selected due to their different mix of ion-scale driftwaves, ion temperature gradient (ITG) mode and trapped electron mode (TEM), and electromagnetic instabilities, kinetic ballooning mode (KBM), and micro-tearing mode (MTM) in the plasma core. It is found that the effect of parallel magnetic fluctuations is strongly destabilizing to the unstable KBMs compared to calculations with only perpendicular magnetic fluctuations. Two discharges have a mix of ITG/TEM and MTMs that are predicted to be dominant instability across the plasma radius. The parallel magnetic fluctuations are found to have little effect on the MTM stability but are destabilizing to ITG/TEM modes. To test the validity of the gyro-fluid linear stability codes TGLF and GFS at low aspect ratio, a database of linear growth rates has been created using the CGYRO gyrokinetic code. The database is comprised of various parameter scans around a standardized set of NSTX-U core parameters. It contains a group of electrostatic cases and an electromagnetic group that includes the effects of perpendicular and parallel magnetic fluctuations. Comparing the results from the GFS and TGLF models, we find that GFS exhibits the best agreement with the database of CGYRO linear growth rates. Comparing the model results for the electromagnetic scans shows that GFS captures the effects of parallel magnetic fluctuations accurately, while the TGLF model does not, as it lacks sufficient perpendicular energy resolution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prediction of carbon nanostructure mechanical properties and the role of defects using machine learning

Graphene-based nanostructures hold immense potential as strong and lightweight materials, however, their mechanical properties such as modulus and strength are difficult to fully exploit due to challenges in atomic-scale engineering. This study presents a database of over 2,000 pristine and defective nanoscale CNT bundles and other graphitic assemblies, inspired by microscopy, with associated stress–strain curves from reactive molecular dynamics (MD) simulations using the reactive INTERFACE force field (IFF-R). These 3D structures, containing up to 80,000 atoms, enable detailed analyses of structure-stiffness-failure relationships. By leveraging the database and physics- and chemistry-informed machine learning (ML), accurate predictions of elastic moduli and tensile strength are demonstrated at speeds 1,000 to 10,000 times faster than efficient MD simulations. Hierarchical Graph Neural Networks with Spatial Information (HS-GNNs) are introduced, which integrate chemistry knowledge. HS-GNNs as well as extreme gradient boosted trees (XGBoost) achieve forecasts of mechanical properties of arbitrary carbon nanostructures with only 3 to 6% mean relative error. The reliability equals experimental accuracy and is up to 20 times higher than other ML methods. Predictions maintain 8 to 18% accuracy for large CNT bundles, CNT junctions, and carbon fiber cross-sections outside the training distribution. The physics- and chemistry-informed HS-GNN works remarkably well for data outside the training range while XGBoost works well with limited training data inside the training range. The carbon nanostructure database is designed for integration with multimodal experimental and simulation data, scalable beyond 100 nm size, and extendable to chemically similar compounds and broader property ranges. The ML approaches have potential for applications in structural materials, nanoelectronics, and carbon-based catalysts.

Winetrout, Jordan J.↗

Validation of Fast Reactor Depletion Tools Using EBR-II Measured Data

The validation of simulation tools for calculating fuel depletion and evolution in fast reactors is vital for design, licensing, deployment, operations, and material accountancy. The Physics Analysis Database (PADB) and Analytical Laboratory (AL) database contain measured data collected from Experimental Breeder Reactor II (EBR-II) and were used to validate the most recent versions of the Argonne Reactor Computation (ARC) tool suite and ORIGEN-S for calculating isotopic compositions in irradiated fast reactor fuel. The PADB contains important modeling and operational information about the EBR-II core design, fuel cycle, and analytical results from the legacy versions of the ARC tool suite. The AL database contains the measured isotopic compositions of irradiated samples taken from core subassemblies. A new procedure was developed for the ARC tool suite to perform the EBR-II depletion simulation, as well as to perform more detailed isotopic calculations using ORIGEN-S calculations by coupling it with the ARC suite. Both the ARC and ARC-ORIGEN results were compared with the AL measured data for all relevant samples and showed good agreement for the major actinides. Good agreement with measured data was also achieved using the ARC-ORIGEN approach for several fission products.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Metrics and extrapolation of resonant magnetic perturbation thresholds for ELM suppression

This large database study of resonant magnetic perturbation (RMP) edge localized mode (ELM) suppression thresholds in the AUG, DIII-D, EAST, and KSTAR tokamaks details the key strengths and weaknesses of RMP metrics. The RMP ELM suppression database used for this work contains plasma information at the time of transition from ELMing to ELM suppressed states where a clear experimental threshold is identified. The experimental threshold distributions are compared for five metrics: (1) the island overlap width, (2) pedestal top Chirikov overlap, (3) peeling edge displacement, (4) pedestal top resonant drive, and (5) edge dominant mode overlap. The distributions, the regularity of the dependence on RMP coil currents, and the sensitivities of a given metric to equilibrium reconstruction details are compared. The overlap metric proves to be a good compromise between including the appropriate plasma response physics and maintaining a numerical robustness. This quantity does not exhibit clear power-law scalings for projection, but machine learning can assist in predicting thresholds within the existing parameter ranges and providing uncertainty quantification of those predictions. Two new first-principles models, one utilizing a threshold from the non-linear Modified Rutherford equation evaluated at the pedestal top and one utilizing the SLAYER code to calculate the linear tearing threshold from torque balance, offer possible paths to extrapolation beyond the existing database parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗