Search NASASearch

SEARCH · Search NASA

Results for “data holdings”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES

The dark energy survey supernova program: investigating beyond-ΛCDM

We report constraints on a variety of non-standard cosmological models using the full 5-yr photometrically classified type Ia supernova sample from the Dark Energy Survey (DES-SN5YR). Both Akaike Information Criterion (AIC) and Suspiciousness calculations find no strong evidence for or against any of the non-standard models we explore. When combined with external probes, the AIC and Suspiciousness agree that 11 of the 15 models are moderately preferred over Flat-|$\Lambda$|CDM suggesting additional flexibility in our cosmological models may be required beyond the cosmological constant. We also provide a detailed discussion of all cosmological assumptions that appear in the DES supernova cosmology analyses, evaluate their impact, and provide guidance on using the DES Hubble diagram to test non-standard models. An approximate cosmological model, used to perform bias corrections to the data holds the biggest potential for harbouring cosmological assumptions. We show that even if the approximate cosmological model is constructed with a matter density shifted by |$\Delta \Omega _{\rm m}\sim 0.2$| from the true matter density of a simulated data set the bias that arises is subdominant to statistical uncertainties. Nevertheless, we present and validate a methodology to reduce this bias.

79 ASTRONOMY AND ASTROPHYSICS

Improved constraints on hematite refractive index for estimating climatic effects of dust aerosols

Abstract Uncertainty in desert dust composition poses a big challenge to understanding Earth’s climate across different epochs. Of particular concern is hematite, an iron-oxide mineral dominating the solar absorption by dust particles, for which current estimates of absorption capacity vary by over two orders of magnitude. Here, we show that laboratory measurements of dust composition, absorption, and scattering provide valuable constraints on the absorption potential of hematite, substantially narrowing its range of plausible values. The success of this constraint is supported by results from an atmospheric transport model compared with station-based measurements. Additionally, we identify substantial bias in simulating hematite abundance in dust aerosols with current soil mineralogy descriptions, underscoring the necessity for improved data sources. Encouragingly, the next-generation imaging spectroscopy remote sensing data hold promise for capturing the spatial variability of hematite. These insights have implications for enhancing dust modeling, thus contributing to efforts in climate change mitigation and adaptation.

Environmental Sciences & Ecology

Data for: Miscanthus × giganteus changes soil structure and increases maximum water holding capacity

The data provided include results from a comparative study evaluating the impact of Miscanthus × giganteus (miscanthus) versus maize on soil structural properties and maximum water holding capacity (MWHC) across two Iowa sites. The dataset includes MWHC values determined using the Funnel Filter Paper Drainage (FFPD) method, as well as additional measurements of MWHC following structural disruption of the soil to isolate the effect of aggregation. It also contains three-dimensional micro-computed tomography (microCT) data used to quantify total porosity and pore size distribution (PSD) of soil aggregates at a 5 µm resolution. All data are provided as raw replicate-level measurements, organized by site, crop, and depth, along with processed summary files in table form in CSV (.csv) format to support reproducibility and downstream analysis.

Misanthus x giganteus

Super Resolution for Renewable Energy Resource Data With Wind From Reanalysis Data (Sup3rWind) and Application to Ukraine [Slides]

In this work we present a novel deep learning-based downscaling method, using generative adversarial networks (GANs), for generating high-resolution wind resource data from ECMWF Reanalysis v5 data (ERA5). We show that by training a GAN model on ERA5, as opposed to coarsened high-resolution data, we achieve results that are competitive with conventional dynamical downscaling. This GAN-based downscaling method additionally reduces computational costs over dynamical downscaling by two orders of magnitude. All GANs are trained on data sampled from CONUS, selected to provide a diverse sampling of terrain conditions, and validated on observational data along with data held out from training. This cross-validation shows low error and high correlations with observations and excellent agreement with hold out data across physical distributions. Our approach is finally used to downscale 30km hourly ERA5 to 2-km 5-minute wind data, for January 2000 through December 2023, at multiple hub heights, over Ukraine, Moldova, and part of Romania. Comparisons against observational data from Meteorological Assimilation Data Ingest System (MADIS) and multiple wind farms show the same level of performance as for CONUS validation. This 24 year data record is the first member of the "super resolution for renewable energy resource data with wind from reanalysis data" dataset (Sup3rWind).

17 WIND ENERGY

Portable Acceleration of CMS Computing Workflows with Coprocessors as a Service

Computing demands for large scientific experiments, such as the CMS experiment at the CERN LHC, will increase dramatically in the next decades. To complement the future performance increases of software running on central processing units (CPUs), explorations of coprocessor usage in data processing hold great potential and interest. Coprocessors are a class of computer processors that supplement CPUs, often improving the execution of certain functions due to architectural design choices. We explore the approach of Services for Optimized Network Inference on Coprocessors (SONIC) and study the deployment of this as-a-service approach in large-scale data processing. In the studies, we take a data processing workflow of the CMS experiment and run the main workflow on CPUs, while offloading several machine learning (ML) inference tasks onto either remote or local coprocessors, specifically graphics processing units (GPUs). With experiments performed at Google Cloud, the Purdue Tier-2 computing center, and combinations of the two, we demonstrate the acceleration of these ML algorithms individually on coprocessors and the corresponding throughput improvement for the entire workflow. This approach can be easily generalized to different types of coprocessors and deployed on local CPUs without decreasing the throughput performance. We emphasize that the SONIC approach enables high coprocessor usage and enables the portability to run workflows on different types of coprocessors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Expandable Log Analyzing Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, IL]

Expandable Log Analyzing Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, IL]

Database-Agnostic Log Analysis and Monitoring Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, US, IL; Fermilab]

MARIAH PCAP data for Validation Demonstration

This dataset holds simulated PCAP (packet capture) data from the SCEPTRE validation demonstration model as a set of pairwise communications between devices via specific protocols. All connections should be assumed to be symmetric, as this data is an aggregation of the true PCAP. A mapping is also provided associating each IP address with its true device type.

cyber-physical system

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles

Trap‐Engineering the Persistent Luminescence of Ca 3 Ga 4 O 9 :Tb 3+ via Al 3+ Substitution for Optical Data Storage

Abstract Optically stimulated luminescence (OSL) materials hold great potential for optical data storage (ODS) and anticounterfeiting applications. Nevertheless, the scarcity of suitable luminescent materials with deep‐level traps remains a significant obstacle. Herein, a host substation strategy have been employed to tune the persistent luminescence (PersL) and OSL properties of Ca 3 Ga 4 O 9 :Tb 3+ by Al 3+ substitution through trap engineering and demonstrated their potential. Specifically, the photoluminescence of the Ca 2.985 (Ga 1‐y% Al y% ) 4 O 9 :0.5%Tb 3+ of Tb 3+ is first investigated due to its different occupancies of Ca 2+ . The influence of host substitution on the crystal structure, trap depth, trap density, PersL, and OSL properties have further investigated. A series of strong PersL and OSL peaks from the Ca 2.985 (Ga 1‐y% Al y% ) 4 O 9 :0.5%Tb 3+ with bluish‐green emissions have been observed. The Ca 2.985 (Ga 1‐y% Al y% ) 4 O 9 :0.5%Tb 3+ have shown controllable photon release upon thermal and optical stimuli, enhancing their performance for ODS. Thermally stimulated luminescence suggests that vacancy and defect concentrations inside the Ca 3‐x% (Ga 1‐y% Al y% ) 4 O 9 :x%Tb 3+ can be manipulated by Tb 3+ doping and Al 3+ substitution, which ultimately leads to the formation of deep traps and a broad distribution of traps with increased deep trap concentration. The work demonstrates that trap engineering through Al 3 ⁺ substitution is an effective method for tuning PersL and OSL properties of Ca 2.985 (Ga 1‐y% Al y% ) 4 O 9 :0.5%Tb 3+ for ODS.

Abeywickrama, Thulitha M. [Department of Chemistry

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE

Visual Systems Mapping to Define and Compare Woody Biomass LCAs for Sustainable Systems

The challenge addressed in this research centres on the need to choose between several biomass sources and energy production processes, while supporting rural economies and resilience of forest systems. A key barrier to effective decision-making for strategies using biomass is the lack of standardized and transparent life cycle assessment (LCA) baselines. These baselines are critical for assessing the impacts of biomass strategies but often vary due to regional factors and chosen simplifying assumptions of the LCAs. However, omitting key variables can mean the LCA omits key feedback and balancing loops relevant to fully assessing impacts of the change or test scenario. To address these complexities, this project employs a systems engineering approach: visual systems mapping. This technique is used to define the boundaries and dynamic behaviours of LCA baselines, enhancing transparency. By examining five literature sources and their documented baseline scenarios, the systems mapping case-studies demonstrates an approach to documenting and archiving these baselines. Recommendations are that visual systems mapping should be used to document key assumptions, such as baselines, of LCAs. Further, where possible open data repositories should hold key information about LCA baselines and reproducible workflows (e.g., using open-source tools) should be used to improve transparency and comparability in LCAs. Given the consensus within the broader scientific community on the importance of replicable data practices, this research reinforces the need for standardized frameworks and systems engineering tools in LCAs. This research demonstrates a pathway to more transparent, standardized, and comparable LCAs, that may bolster decisions for biomass systems.

Davis, Maggie [ORNL] (ORCID:0000000181319328)

2015-2017 California Vehicle Survey

The 2015-2017 California Vehicle Survey of residential and commercial light-duty vehicle owners in California assessed consumer preferences for vehicles and included a targeted sample of plug-in electric vehicle (PEV) owners. Resource Systems Group conducted the survey on behalf of the California Energy Commission. In addition to economic and demographic data, the survey integrated light-duty vehicle holding and use information with vehicle choice data collected via the stated preferences survey's set of eight vehicle and fuel type choice exercises. The PEV owner survey participants provided additional data on charging behavior, electricity rates, and their main motivations for purchasing PEVs.

1Hz data

2019 California Vehicle Survey

The 2019 California Vehicle Survey of residential and commercial light-duty fleet owners in California assessed consumer preferences for vehicles and included a targeted sample of plug-in electric vehicle (PEV) owners. In addition to economic and demographic data, the survey integrated light-duty vehicle holding and use information with vehicle choice data, which was collected via a set of eight exercises on vehicle and fuel type choice. The PEV owner survey participants provided additional data on charging behavior, electricity rates, and their main motivations for purchasing PEVs.

1Hz data

Advances in Metallic Fuel Database Development and Data Qualification

The Fuels Irradiation and Physics Database (FIPD [1]) is a comprehensive repository of data and documents related to Uranium-Zirconium based metallic fuel test pins. This database stores operational conditions of these pins, calculated using a suite of Argonne National Laboratory analysis codes developed during the Integral Fast Reactor (IFR) program. Key calculated data include axial distributions of power, temperature, fluence, burnup, and isotopic densities. Additionally, the FIPD holds post-irradiation examination (PIE) data such as fission gas release, gas chemistry measurements, and axial distributions derived from profilometry, gamma scanning, and neutron radiography. Complementing these data is an extensive archive of documents related to various pins and experiments. These include raw PIE records, design details, safety analyses, and operational reports. More detail about FIPD can be found in ref. [2]. The database development is an ongoing effort covering metallic fuel experiments from the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF). The recent improvements to the database and the data QA status are summarized in this paper.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)