Search NASA⌕ Search

SEARCH · Search NASA

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

WholeTraveler Anonymized Data Phase 1

Phase 1 of the WholeTraveler Study data collection consisted of an online-only survey. This survey captured data on three categories of observable variation in the population relevant to transportation decisions. First, the survey collected traditional demographic data such as age, gender, income, and education level. Second, it collected data across personality, psychological, and preference categories. This included: 1. The "Big Five" inventory personality traits: openness to new experience, conscientiousness, extroversion, agreeableness, and neuroticism; 2. Risk and time preferences; and 3. Environmental preferences. Third, the survey collected data on historical behavior patterns including: 1. Adoption of (as well as interest in) new technologies or innovations (e.g., smartphones, PEVs, solar panels, adaptive cruise control [ACC]); 2. Car ownership history and current car ownership status; 3. Recent mode use across different time scales (e.g., previous week, previous month, previous year); and 4. Timing of major life events such as starting a family as well as overall lifecycle trajectory patterns. Data from Phase 1 and Phase 2 are linked by a unique respondent identifier. Anonymized versions of the Phase 1 and Phase 2 data are both available on Livewire.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

JUSTIFI: Software for Improving Performance Objectives via Energy Efficiency

With growing energy supply concerns and rising costs, energy efficiency is a critical component of industrial energy resilience and competitiveness by directly reducing energy operating costs. Energy efficiency projects in manufacturing also yield valuable benefits to other key metrics, such as improved quality, reduced maintenance costs, improved safety, decreased pollution, and enhanced productivity. However, it is difficult to receive approval for energy efficiency projects, so implementation rates are low, even when meeting capital project payback period criteria. The inclusion and quantification of non-energy benefits (NEBs) in the decision-making process for energy efficiency projects can improve the overall financial payback period while demonstrating a positive impact on the firm's key performance metrics and business strategy. Despite their significant financial and strategic value, NEBs are rarely factored into decision-making due to lack of tools to effectively identify and quantify them. Therefore, a comprehensive and integrative approach is needed for the rapidly evolving energy landscape. To address these challenges, through funding from U.S. Department of Energy, our new assessment methodology integrates common continuous improvement six sigma concepts, such as the DMAIC process, and a protocol of guiding questions, into energy efficiency assessments to identify NEBs. We have also developed open-source software, JUSTIFI, to guide users through this process, data collection, and quantification. It is designed to be used concurrently with DOE energy system analysis software suite, MEASUR. Our methodology and tools inform energy assessors, firm engineering, decision makers, and workforce seeking to increase energy resilience and to maximize benefits aligned with performance metrics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

WE-Validate: An Open-Source Framework For Wind Power Validation

Grid operators rely on historical weather time series at existing and planned wind power plants to make informed decisions when planning for a future power grid with very high penetration of renewable power. While synthetic wind power time series have been developed based on historical weather models, their validation with actual power production data remains complex due to variations in modeling practices and methodologies. This paper introduces the WE-Validate framework, originally designed for wind speed validation and now enhanced for wind power validation with a graphical user interface to support users with minimal programming experience. Validation of wind power with WE-Validate is based on robust metrics consisting of RMSE, centered RMSE, average bias, average percent bias, mean absolute error, mean absolute percent error, cross correlation, and calculation of ramping magnitude, rate, and duration. This paper showcases WE-Validate with validation of synthetically derived power for a wind plant in Washington state for one month in 2018. Validation of the synthetic power from two comparison data sets compared with observations shows both comparison series have strong correlation with observed across weekly and monthly aggregations while suffering from persistent negative bias. The suite of metrics within WE-Validate facilitates immediate insight into the utility of the comparison data sets through compression across multiple axes. This user-friendly, open-source tool can be extended beyond wind power, making it a valuable resource for system planners and operators in different domains.

Moncheur de Rieudotte, Malcolm P.↗

A universal language for finding mass spectrometry data patterns

Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here, in this study, we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.

Damiani, Tito [Czech Academy of Sciences (CAS), Pr↗

Resolving root causes of experiment discrepancies guided by machine learning

Abstract Scientists rely on accurate experimental data to explain nature and then harness this knowledge for applications addressing human needs. However, discrepancies between experiments of the same observable can impede scientific progress if one does not understand the underlying causes. Here, we developed a process that unravels data discrepancies by first using Bayesian machine learning to relate discrepancies to few of many, potentially biasing metadata features that encode experiment procedures. This machine learning output guides human experts to study discrepancy causes by simulating suspicious aspects of historical experiments or designing modern ones to address open questions. The study findings then lead to rejecting or correcting historical data on firm scientific bases. This process is demonstrated for the energy spectrum of neutrons emitted promptly (<1 ns) after fission of 252 Cf, a trusted nuclear physics Standard. It reduces the spread in experimental 252 Cf spectra by up to a factor of 6.

Neudecker, D. (ORCID:0000000339200627)↗

Machine learning and TDDFT software for stopping power computation

(SF-24-012) Stopping power describes the rate that a material slows radiation particles passing through it and is useful in designing many technologies. Few organizations can perform new measurements, which require significant resources and rare equipment, and all others rely on coarse approximations rendered from pre-existing data. Methods for computing stopping power in new materials, such as time-dependent density functional theory (TD-DFT), have only recently (circa-2015) become available but are too computationally costly to use frequently enough to have a pronounced impact. We have created a method that opens a pathway to computing stopping power without any need for experimental data by combining electronic structure computations and machine learning.

Ward, Logan↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Modeling the Effects of Artificial Drainage on Agriculture-dominated Watersheds using a Fully Distributed Integrated Hydrology Model: Datasets, scripts, model files

This model-data archive supports the research paper that demonstrates the integration of agricultural drainage features—specifically, narrow engineered ditches and tile drains—into a fully distributed, basin-scale integrated surface-subsurface hydrology model (ISSHM), Amanzi-ATS. The model employs innovative computational meshes aligned with agricultural ditches and incorporates the physically based Hooghoudt's drainage equation to simulate tile drainage, offering a novel strategy that enhances the accuracy of hydrological simulations.The archived dataset includes input parameters, model configurations, and select simulation outputs for the Amanzi-ATS model that successfully captured the streamflow patterns in the Portage River Watershed as validated by USGS gauge readings. Jupyter notebook for the preparation of model inputs and post-processing of outputs are also included. The model's predictive performance achieved a normalized Kling-Gupta Efficiency (KGE) of 0.81, surpassing SWAT without the necessity for site-specific calibration.The Amanzi-ATS model presented in this modeL-data archive allows for numerical experiments to explore the shifts in the flow structure under different drainage scenarios. As a tool for advancing the understanding of distributed hydrological responses and nutrient cycling, this archived model provides valuable insights for researchers, modelers, and decision-makers involved in watershed management and environmental modeling.The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include CSV and HDF5 files, which can be read through Python scripts. The input files for the ATS model, open-source integrated hydrology, and transport model, are in XML format and can be edited in any commonly used text editors.

54 ENVIRONMENTAL SCIENCES↗

IM3 Data Center Driven Grid Stress Dataset for the U.S. Western Interconnection

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from an open-source grid operations modeling framework - GO), for the Integrated Multisector, Multiscale Modeling (IM3) project, under varying levels of data center demand growth between 2025 and 2035 in the U.S. Western Interconnection. The scenarios and sensitivity experiments are combinations of different data center demand growth rates and energy, weather, population and economic pathways. Data center demand growth projections were sourced from the Electric Power Research Institute (EPRI). The data center demand growth projection names are: Low (3.71% annual data center demand growth) Moderate (5% annual data center demand growth) High (10% annual data center demand growth) Higher (15% annual data center demand growth) Energy, weather, population and economic pathways are informed by two Shared Socioeconomic Pathways (SSP3 and SSP5) and two Representative Concentration Pathways (RCP4.5 and RCP8.5) following the hotter general circulation model (GCM) forcing group from a set of perturbed thermodynamics simulations. The resulting pathway names are: rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 The main scenarios and sensitivity experiments are detailed below. Reference scenario: The projected grid stress and reliability results for the U.S. Western Interconnection from a previous study. This scenario does not consider data center demand growth explicitly. Data center scenario: Building on the reference scenario, this scenario considers various data center growth rates and how they impact the U.S. Western Interconnection. Data center loads are modeled as flat 8760-hr profiles. This scenario does not consider new generation and transmission capacities specifically designed to meet the new data center demands. The related folder is named "flat". Delayed generator retirements sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different levels of natural gas and nuclear generator retirement delays. The resulting scenario names are: (1) postponing 100% nuclear retirements; (2) postponing 100% nuclear and 25% natural gas retirements; (3) postponing only 50% natural gas retirements; (4) postponing 100% nuclear and 50% natural gas retirements; (5) postponing 100% nuclear and 75% natural gas retirements; and (6) postponing 100% nuclear and 100% natural gas retirements. The related folder names are: no_gen_retire_0_gas, no_gen_retire_25_gas, no_gen_retire_50_gas, no_gen_retire_50_gas_only, no_gen_retire_75_gas, and no_gen_retire_100_gas. Demand response through curtailment sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different participation and compensation levels of data center demand response. The resulting scenario names are: (1) 5% demand available for curtailment with 750 $/MWh compensation; (2) 5% demand available for curtailment with 500 $/MWh compensation; (3) 5% demand available for curtailment with 250 $/MWh compensation; (4) 15% demand available for curtailment with 750 $/MWh compensation; (5) 15% demand available for curtailment with 500 $/MWh compensation; and (6) 15% demand available for curtailment with 250 $/MWh compensation. The related folder names are: dr_cost_250_drup_0_drdown_5, dr_cost_250_drup_0_drdown_15, dr_cost_500_drup_0_drdown_5, dr_cost_500_drup_0_drdown_15, dr_cost_750_drup_0_drdown_5, and dr_cost_750_drup_0_drdown_15. Combination of delayed generator retirements and demand response through curtailment sensitivity experiment: The impact of combining postponing 100% nuclear and 25% natural gas retirements with 5% demand available for curtailment with 750 $/MWh compensation is simulated. The related folder is named "dr_cost_750_drup_0_drdown_5_nuc_100_gas_25". Please refer to the README file for a detailed description of the dataset including individual files and references.

Artificial Intelligence↗

Divide and conquer: using RhizoVision Explorer to aggregate data from multiple root scans using image concatenation and statistical methods

Roots are important in agricultural and natural systems for determining plant productivity and soil carbon inputs. Sometimes, the amount of roots in a sample is too much to fit into a single scanned image, so the sample is divided among several scans, and there is no standard method to aggregate the data. Here, we describe and validate two methods for standardizing measurements across multiple scans: image concatenation and statistical aggregation. We developed a Python script that identifies which images belong to the same sample and returns a single, larger concatenated image. These concatenated images and the original images were processed with RhizoVision Explorer, a free and open-source software. An R script was developed, which identifies rows of data belonging to the same sample and applies correct statistical methods to return a single data row for each sample. These two methods were compared using example images from switchgrass, poplar, and various tree and ericaceous shrub species from a northern peatland and the Arctic. Most root measurements were nearly identical between the two methods except median diameter, which cannot be accurately computed by statistical aggregation. We believe the availability of these methods will be useful to the root biology community.

59 BASIC BIOLOGICAL SCIENCES↗

Power Profile Monitoring and Tracking Evolution of System-Wide HPC Workloads

The power & energy demands of HPC machines have grown significantly. Modern exascale HPC systems require tens of megawatts of combined power for computing resources and cooling facilities at full capacity. The current energy trend is not sustainable for future HPC systems, and there is a need to work toward the energy efficiency aspect of HPC performance. Energy awareness of the HPC applications at the job level is essential for running an efficient HPC system. This work aims to develop a pipeline to provide a production-level system-wide overview of the HPC workloads' power profile while handling evolving workloads exhibiting new power trends. We developed an open-set classification model for HPC jobs based on the properties of power profiles to continuously provide a system-wide holistic view of recently completed jobs. The pipeline helps continuously monitor the job-level power usage pattern of HPC and enables us to capture the new trends in applications' power behavior. We employed a comprehensive set of techniques to generate job-level data, custom-designed feature extraction methods to extract critical features from jobs' power profiles, clustering techniques powered by generative modeling, and open-set classification for identifying job profiles into known classes or an unknown set. With extensive evaluations, we demonstrate the effectiveness of each component in our pipeline. We provide an analysis of the resulting clusters that characterize the power profile landscape of the Summit supercomputer from more than 60K jobs executed in a year. The open-set classification classifies the known data sets into known classes with high accuracy and identifies unknown data noints with over 85% accuracy.

Karimi, Ahmad Maroof↗

Dissolution zone model of the oxide structure in additively manufactured dispersion-strengthened alloys

The structural evolution of oxides in dispersion-strengthened superalloys during laser-powder bed fusion is considered in detail. Alloy chemistry and process parameter effects on oxide structure are assessed through a parameter study on the model alloy Ni-20Cr, doped with varying concentrations of Y 2 O 3 and Al. Small angle neutron scattering measurements of the dispersoid size distribution show the dispersoid size increases with higher laser power, slower scan speed, and increasing Y 2 O 3 and Al content. Complementary electron microscopy measurements reveal reactions between Y 2 O 3 and Al, even in nanoscale dispersoids, and the presence of micron-scale oxide slag inclusions in select specimens. A scaling analysis of mass and momentum transport within the melt pool, presented here, establishes that diffusional structural evolution mechanisms dominate for nanoscale dispersoids, while fluid forces and advection become significant for larger slag inclusions. These findings are developed into a theory of dispersoid structural evolution, integrating quantitative models of diffusional processes – dispersoid dissolution, nucleation, growth, coarsening – with a reduced order model of time-temperature trajectories of fluid parcels within the melt pool. Calculations of the dispersoid size in single-pass melting reveal a zone in the center of the melt track in which the oxide feedstock fully dissolves. Within this zone the final Y 2 O 3 size is independent of feedstock size and determined by nucleation and growth kinetics. If the dissolution zones of adjacent melt tracks overlap sufficiently with each other to dissolve large oxides, formed during printing or present in the powder feedstock, then the dispersoid structure throughout the build volume is homogeneous and matches that from a single pass within the dissolution zone. Gaps between adjacent dissolution zones result in oxide accumulation into larger slag inclusions. Predictions of final dispersoid size and slag formation using this dissolution zone model match the present experimental data and explain process-structure linkages speculated in the open literature.

36 MATERIALS SCIENCE↗

Source characterization of a detector for heavy and superheavy nuclei

A new focal plane detector system for the Berkeley Gas-filled Separator (BGS) was designed, constructed, and tested offline with various α-decay and conversion-electron sources. The SuperHeavy RECoils (SHREC) detector comprises sets of double-sided silicon strip detectors arranged in an open-faced cuboid geometry. Alongside the detector upgrade new digital data acquisition electronics have been commissioned offline. This setup aims to detect separated recoiling heavy and superheavy nuclei as well as their correlated radioactive decay paths with improved energy resolution and overall sensitivity.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

KBKit: A Python Toolkit for Kirkwood–Buff Theory from Molecular Dynamics

Thermodynamic properties of liquid mixtures govern processes that range from drug delivery to energy storage, yet extracting these properties from molecular simulations remains challenging. Kirkwood–Buff (KB) theory offers a rigorous route by linking microscopic pair distribution functions to macroscopic free energies, but practical use of the theory has been hindered by two obstacles: (i) the long simulations needed to obtain well-converged Kirkwood-Buff integrals (KBIs) and (ii) the specialized corrections required to translate finite-size data to the thermodynamic limit. $\texttt{KBKit}$ is an open-source Python package that removes these barriers. It automatically computes KBIs and derived thermodynamic quantities from GROMACS input files, applies state-of-the-art finite-size corrections, and provides built-in diagnostic tools to quantify statistical uncertainty. Written with modern software-engineering practices—continuous integration, extensive unit testing, and thorough documentation—$\texttt{KBKit}$ is both reliable and easy to extend. By condensing complex KBI analysis into a few intuitive commands, $\texttt{KBKit}$ enables researchers to incorporate KB theory into routine simulation workflows and accelerate the discovery of solution-phase thermodynamics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automation-Accelerated Electrolyte Design Mitigates Solubility Competition between Redox-Active Molecules and Supporting Salts

In nonaqueous redox-flow batteries (NRFBs), redox-active organic molecules (ROMs) and supporting salts compete for solvation sites, limiting achievable energy density. We combine automated high-throughput experimentation (HTE) with camera-based saturation monitoring and quantitative NMR to measure paired (ROM, salt) solubilities across single and mixed organic solvents. Using 2,1,3-benzothiadiazole (BTZ) with lithium bis(trifluoromethanesulfonyl)imide (LiTFSI) as a model system, we find that a binary m-xylene/acetonitrile mixture dissolves ≈3 M of both BTZ and LiTFSI─surpassing the previously reported 2 M ceiling for neat acetonitrile─by leveraging complementary solvation (MX is BTZ-philic and salt-phobic; ACN stabilizes LiTFSI). A random-forest model (RMSE ≈ 0.24) trained on solvent descriptors highlights log P and salt concentration as dominant predictors and predicts MX/ACN ≈0.3/0.7 (v/v) to be near-optimal. These formulations retain practical viscosity and ∼5 mS·cm –1 conductivity at high loading. In conclusion, the workflow provides a reproducible, data-centric route to NRFB electrolyte design and motivates an open, standardized dual-solute solubility resource for accelerated electrolyte discovery.

Electrolytes↗

Spectrum and extension of the inverse-Compton emission of the Crab Nebula from a combined Fermi -LAT and H.E.S.S. analysis

The Crab Nebula is a unique laboratory for studying the acceleration of electrons and positrons through their non-thermal radiation. Observations of very-high-energy γ rays from the Crab Nebula have provided important constraints for modelling its broadband emission. We present the first fully self-consistent analysis of the Crab Nebula’s γ-ray emission between 1 GeV and ∼100 TeV, that is, over five orders of magnitude in energy. Using the open-source software package GAMMAPY, we combined 11.4 yr of data from the Fermi Large Area Telescope and 80 h of High Energy Stereoscopic System (H.E.S.S.) data at the event level and provide a measurement of the spatial extension of the nebula and its energy spectrum. We find evidence for a shrinking of the nebula with increasing γ-ray energy. Furthermore, we fitted several phenomenological models to the measured data, finding that none of them can fully describe the spatial extension and the spectral energy distribution at the same time. Especially the extension measured at TeV energies appears too large when compared to the X-ray emission. Our measurements probe the structure of the magnetic field between the pulsar wind termination shock and the dust torus, and we conclude that the magnetic field strength decreases with increasing distance from the pulsar. We complement our study with a careful assessment of systematic uncertainties.

79 ASTRONOMY AND ASTROPHYSICS↗

Structural complexity of snapshots of two-dimensional Fermi-Hubbard systems

The development of quantum gas microscopy for two-dimensional optical lattices has provided an unparalleled tool to study the Fermi-Hubbard model (FHM) with ultracold atoms. Spin-resolved projective measurements, or snapshots, have played a significant role in quantifying correlation functions, theory verification, and thus the uncovering of underlying physical phenomena such as antiferromagnetism at commensurate filling on bipartite lattices and other charge and spin correlations, as well as dynamical properties at various densities. Here we employ a recent concept, the multiscale structural complexity, and show that when computed for the snapshots (of single spin species, local moments, or total density) it can provide a theory-free property, immediately accessible to experiments. Specifically, after benchmarking results for Ising and $XY$ models, we study the structural complexity of snapshots of the repulsive FHM in the two-dimensional square lattice as a function of doping and temperature. We generate projective measurements using determinant quantum Monte Carlo and compare their complexities against those from the experiment. We demonstrate that these complexities are linked to relevant physical observables such as the entropy and double occupancy. Their behaviors capture the development of correlations and relevant length scales in the system. Furthermore, we provide an open-source code in python which can be implemented into data analysis routines in experimental settings for the square lattice.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗