Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Dynamic Server-Based KML Code Generator Method for Level-of-Detail Traversal of Geospatial Data

Web-based geospatial client applications such as Google Earth and NASA World Wind must listen to data requests, access appropriate stored data, and compile a data response to the requesting client application. This process occurs repeatedly to support multiple client requests and application instances. Newer Web-based geospatial clients also provide user-interactive functionality that is dependent on fast and efficient server responses. With massively large datasets, server-client interaction can become severely impeded because the server must determine the best way to assemble data to meet the client applications request. In client applications such as Google Earth, the user interactively wanders through the data using visually guided panning and zooming actions. With these actions, the client application is continually issuing data requests to the server without knowledge of the server s data structure or extraction/assembly paradigm. A method for efficiently controlling the networked access of a Web-based geospatial browser to server-based datasets in particular, massively sized datasets has been developed. The method specifically uses the Keyhole Markup Language (KML), an Open Geospatial Consortium (OGS) standard used by Google Earth and other KML-compliant geospatial client applications. The innovation is based on establishing a dynamic cascading KML strategy that is initiated by a KML launch file provided by a data server host to a Google Earth or similar KMLcompliant geospatial client application user. Upon execution, the launch KML code issues a request for image data covering an initial geographic region. The server responds with the requested data along with subsequent dynamically generated KML code that directs the client application to make follow-on requests for higher level of detail (LOD) imagery to replace the initial imagery as the user navigates into the dataset. The approach provides an efficient data traversal path and mechanism that can be flexibly established for any dataset regardless of size or other characteristics. The method yields significant improvements in userinteractive geospatial client and data server interaction and associated network bandwidth requirements. The innovation uses a C- or PHP-code-like grammar that provides a high degree of processing flexibility. A set of language lexer and parser elements is provided that offers a complete language grammar for writing and executing language directives. A script is wrapped and passed to the geospatial data server by a client application as a component of a standard KML-compliant statement. The approach provides an efficient means for a geospatial client application to request server preprocessing of data prior to client delivery. Data is structured in a quadtree format. As the user zooms into the dataset, geographic regions are subdivided into four child regions. Conversely, as the user zooms out, four child regions collapse into a single, lower-LOD region. The approach provides an efficient data traversal path and mechanism that can be flexibly established for any dataset regardless of size or other characteristics.

Baxes, Gregory↗

NASA Power: Global Solar Insolation, Meteorological Parameter Data, and Web Services to Support Sustainable Building Design and Operations

The buildings industry is currently striving to adopt green solutions to make infrastructure more energy-efficient in order to meet the 2050 net-zero climate goals. This planning requires reliable environmental datasets that are crucial in designing, building, and maintaining our world’s-built environment, as well as other energy-related processes and investments. This webinar for the National Institute of Building Sciences provides an overview of NASA’s Prediction Of Worldwide Energy Resources (POWER) Project that informs decision-making and development for sustainable building design and operations by enabling public open discovery, efficient access, and convenient distribution of NASA’s Earth Observations and global atmospheric model datasets. POWER’s datastore is comprised of solar radiation and surface meteorology parameters, spanning nearly 40 years of hourly data, that are easily accessible via several access methods and tools to support three focus areas: 1) renewable energy deployment and management, 2) sustainable infrastructure, and 3) agroclimatology applications. POWER and NASA Earth Science both plan future data parameters, updated tools, and improved observations that could directly support U.S. and international sustainable development goals, climate strategies, and building information modeling. To this end, solar data from several NASA projects and meteorological data from NASA assimilation models have already been reformatted and disseminated to the public via a user-friendly web GIS-enabled based data portal through the POWER Project. POWER data is analysis-ready and accessible through an Application Programming Interface (API), ArcGIS Image Services, and the project’s Data Access Viewer enhanced (DAVe), an interactive online tool. The POWER DAVe also features data consistent with ASHRAE Climate Design Conditions and has developed web image services showing Building Climate Zones and their variability. Through those tools, the data can be downloaded into multiple formats that support the infrastructure community, including CSV and Energy Plus Weather (EPW). POWER’s entire data product catalog is available through Amazon Web Services (AWS) Open Data Registry (ODR) via a free and publicly accessible Simple Storage Service (S3). This webinar provides a full overview of the NASA POWER Project's data and services developed in collaboration with the sustainable infrastructure community. Examples of how the renewable energy and building communities have utilized POWER data products to make decisions and a preview of future data product expansion, including climate projections, and web services will also be provided. Additionally, use case stories from our broad community of users will be presented.

Paul W. Stackhouse↗

Assessing Spatial Representativeness of Global Flux Tower Eddy-Covariance Measurements Using Data from FLUXNET2015

Large datasets of carbon dioxide, energy, and water fluxes were measured with the eddy-covariance (EC) technique, such as FLUXNET2015. These datasets are widely used to validate remote-sensing products and benchmark models. One of the major challenges in utilizing EC-flux data is determining the spatial extent to which measurements taken at individual EC towers reflect model-grid or remote sensing pixels. To minimize the potential biases caused by the footprint-to-target area mismatch, it is important to use flux datasets with awareness of the footprint. This study analyze the spatial representativeness of global EC measurements based on the open-source FLUXNET2015 data, using the published flux footprint model (SAFE-f). The calculated annual cumulative footprint climatology (ACFC) was overlaid on land cover and vegetation index maps to create a spatial representativeness dataset of global flux towers. The dataset includes the following components: (1) the ACFC contour (ACFCC) data and areas representing 50%, 60%, 70%, and 80% ACFCC of each site, (2) the proportion of each land cover type weighted by the 80% ACFC (ACFCW), (3) the semivariogram calculated using Normalized Difference Vegetation Index (NDVI) considering the 80% ACFCW, and (4) the sensor location bias (SLB) between the 80% ACFCW and designated areas (e.g. 80% ACFCC and window sizes) proxied by NDVI. Finally, we conducted a comprehensive evaluation of the representativeness of each site from three aspects: (1) the underlying surface cover, (2) the semivariogram, and (3) the SLB between 80% ACFCW and 80% ACFCC, and categorized them into 3 levels. The goal of creating this dataset is to provide data quality guidance for international researchers to effectively utilize the FLUXNET2015 dataset in the future.

54 ENVIRONMENTAL SCIENCES↗

MINERvA s Open Data Product: A First for Neutrino Data Preservation

Access to information on neutrino nucleus interactions is critical to the success of all neutrino oscillation experiments. MINERvA's rich dataset covers a range of energies and nuclei unique amongst experiments, and as such is critical to the community in building the important shared knowledge needed to unravel the mysteries of the neutrino. In particular, its dataset provides the greatest statistical coverage in in the range of neutrino energies pertinent for DUNE until DUNE's near detector begins operation. Historically, such significant datasets in neutrino physics have been preserved primarily through their published results. While meaningful and useful, this limits the ability to explore the data to its fullest extent as new perspectives continue to form. MINERvA has undertaken a major effort to break this trend and preserve its data in a format to be as analyzable as possible from outside the collaboration. This has culminated in the officially-released MINERvA Open Data Product for the community to take advantage of and utilize. Maintaining direct access to the dataset in an analyzable form will allow new insights to continue to be extracted indefinitely. This talk will cover the contents of this product, the information included (and excluded), the tools provided to utilize the product effectively, the support MINERvA intends to provide in its use, and some lessons learned through the process.

Last, David [Rochester U.] (ORCID:0000000245147183↗

An Open-Science Approach to Address Individual Response to Simulated GCR In Genetically Diverse Populations of Mice and Humans

This project addresses the challenge of understanding and predicting individual radiation sensitivity by integrating genetics, demographics and biomarker characteristics across species (mice and humans). We hypothesize that ex vivo DNA repair response to GCR components is a central determinant of cancer risk from space radiation and can serve as a biomarker of radiation risk in combination with genetics. Automated image quantification of 53BP1+ radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET in non-immortalized primary skin fibroblasts derived from 76 mice across 15 strains (5 inbred reference strains and 10 collaborative-cross strains) exposed to X rays (0.1, 1 and 4 Gy), 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm), as well as in peripheral blood mononuclear cells (PBMCs) from 768 healthy donors (matched ethnicity, 50/50 male/female, 18-70 years old) exposed to gamma rays (0.1 and 1 Gy), 350 MeV/n 28Si, 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm). A genome-wide association study (GWAS) was performed on the mouse strains between DNA damage responses to space radiation and single nucleotide polymorphisms (SNPs). We found SNPs, which were significantly associated to the RIF phenotype, mapped to genes and pathways that are functionally linked to health hazards for deep space exploration (e.g. carcinogenesis, nervous system damage and immune dysfunction). Some of these SNPs were located within protein coding regions, potentially interfering with protein functions and providing promising genetic targets for countermeasures. We also found correlations between both spontaneous and radiation-induced DNA damage and SNPs mapped to pathways associated with cellular metabolism. GWAS is undergoing for the human data. All data have been made available via the NASA Space Biology Open-Science database (genelab.nasa.gov) and we will discuss how various genomic and transcriptomic datasets can be accessed for modeling and integrated using machine learning methods for discovering new radiation biology.

Sylvain V Costes↗

Fundamental Test of a Hovering Rotor: Comprehensive Measurements for CFD Validation

A model-scale hover test of a 4-bladed, 11.08-ft diameter rotor was recently completed inside the National Full-Scale Aerodynamics Complex 80- by 120-Foot Wind Tunnel test section. The primary objective of the test was to acquire key experimental data for a hovering rotor of sufficient quality and quantity to allow validation of state-of-the-art analysis codes. A comprehensive measurement set has been acquired, including rotor performance, blade airloads, flow transition locations, blade deflections, and wake geometry for a range of tip Mach numbers and collective settings. The present paper provides an overview of the test, including detailed descriptions of the hardware, instrumentation, and measurement systems. In addition, the specific test objectives, approach, and sample results are presented. The full test database, as well as detailed rotor geometry information, will ultimately be shared openly on a NASA-sponsored website to serve as a benchmark validation dataset.

Hover↗

Annual and sub-seasonal dynamics of a rapidly eroding permafrost coastline along the Beaufort Sea in northern Alaska

Drew Point, an unlithified ice-rich permafrost coastline along the Alaskan Beaufort Sea, is among the most rapidly eroding Arctic coastlines, with an average erosion rate of 19 m/yr from 2007 to 2019. We use 16 high-resolution remote sensing datasets (satellite, airborne, and UAV imagery) to analyze erosion mechanisms (thermal abrasion and denudation) in relation to environmental forcings along a 1.5 km stretch of coastline during the 2018 and 2019 open water seasons. In a striking contrast, 2019 exhibited the highest mean erosion rate (34.5 m) within the 2007–2019 record, while 2018 had the second lowest (11.2 m). Block failure contributed to sub-seasonal erosion rates 6 to 21 times higher than thermal denudation, with staggered block fall timing, lag responses post-storm, and non-storm block collapse influencing overall erosion magnitude and timing. To quantify wind effects, we developed wind sums, a metric combining cumulative wind speed and directional data that can be used as a proxy for integrated storm intensity capable of incorporating lagged responses that correlated strongly with erosion at sub-seasonal and annual scales. Our findings emphasize the dominant role of wind during periods of open water and air temperature during the thaw season in driving permafrost coastline erosion dynamics, while highlighting the importance of spatiotemporally high-resolution datasets for understanding Arctic coastal change dynamics.

Alaska Beaufort Sea Coast↗

CoCoMET v1.0: a unified open-source toolkit for atmospheric object tracking and analysis

Advances in performance and analysis capabilities have accelerated the development of object tracking algorithms for atmospheric research. This has resulted in a growing number of studies using Lagrangian tracking techniques to analyze the evolution of atmospheric phenomena and the underlying processes. However, the increasing complexity and variety of tracking algorithms present a steep learning curve for new users and make it difficult for existing users to compare algorithm performance. We introduce CoCoMET (Community Cloud Model Evaluation Toolkit), an open-source toolkit that addresses these issues. CoCoMET simplifies the process of running multiple tracking algorithms simultaneously and analyzing objects in both model and observational datasets by specifying parameters in a single configuration file. It standardizes input data from different sources into a consistent format and unifies the tracking output across algorithms. CoCoMET enhances the functionality of existing tracking methods by calculating additional properties such as cell growth and dissipation rates, perimeter, surface area, convexity, and irregularity. In addition, CoCoMET includes a novel method for identifying mergers and splits in 2D and 3D tracks and supports the integration of Eulerian/stationary datasets external to the tracking data for process studies. Its potential utility is demonstrated through examples of model intercomparison, model evaluation against observations, and comparisons between tracking algorithms. Designed for open-source environments, CoCoMET will continue to expand with future releases, incorporating more input data types and tracking algorithms.

54 ENVIRONMENTAL SCIENCES↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

An analysis of parameter compression and Full-Modeling techniques with Velocileptors for DESI 2024 and beyond

In anticipation of forthcoming data releases of current and future spectroscopic surveys, we present the validation tests and analysis of systematic effects within velocileptors modeling pipeline when fitting mock data from the AbacusSummit N-body simulations. We compare the constraints obtained from parameter compression methods to the direct fitting (Full-Modeling) approaches of modeling the galaxy power spectra, and show that the ShapeFit extension to the traditional template method is consistent with the Full-Modeling method within the standard ΛCDM parameter space. We show the dependence on scale cuts when fitting the different redshift bins using the ShapeFit and Full-Modeling methods. We test the ability to jointly fit data from multiple redshift bins as well as joint analysis of the pre-reconstruction power spectrum with the post-reconstruction BAO correlation function signal. We further demonstrate the behavior of the model when opening up the parameter space beyond ΛCDM and also when combining likelihoods with external datasets, namely the Planck CMB priors. Finally, we describe different parametrization options for the galaxy bias, counterterm, and stochastic parameters, and employ the halo model in order to physically motivate suitable priors that are necessary to ensure the stability of the perturbation theory.

79 ASTRONOMY AND ASTROPHYSICS↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

Land Surface Temperature Product Validation Best Practice Protocol Version 1.0 - October, 2017

The Global Climate Observing System (GCOS) has specified the need to systematically generate andvalidate Land Surface Temperature (LST) products. This document provides recommendations on goodpractices for the validation of LST products. Internationally accepted definitions of LST, emissivity andassociated quantities are provided to ensure the compatibility across products and reference data sets. Asurvey of current validation capabilities indicates that progress is being made in terms of up-scaling and insitu measurement methods, but there is insufficient standardization with respect to performing andreporting statistically robust comparisons.Four LST validation approaches are identified: (1) Ground-based validation, which involvescomparisons with LST obtained from ground-based radiance measurements; (2) Scene-based intercomparisonof current satellite LST products with a heritage LST products; (3) Radiance-based validation,which is based on radiative transfer calculations for known atmospheric profiles and land surface emissivity;(4) Time series comparisons, which are particularly useful for detecting problems that can occur during aninstrument's life, e.g. calibration drift or unrealistic outliers due to undetected clouds. Finally, the need foran open access facility for performing LST product validation as well as accessing reference LST datasets isidentified.

best practice↗

Detecting And Characterizing Archetypes of Unintended Consequences in Engineered Systems

When designing engineered systems, the potential for unintended consequences of design policies or design decisions exists despite best intentions. Conditions that might cause the formation of unintended consequences are often known only in hindsight. However, since these conditions are associated with a single event, it is difficult to uncover the general patterns of conditions leading to unintended consequences. In this research, patterns of conditions associated with unintended consequences are learned from historical data and represented in the form of archetypes. While previous work using systems theoretic modeling has identified high-level archetypes, this work leverages a self-organizing map to learn archetypes of unintended consequences from human-tagged risk factors in a large data set of lessons learned from adverse events at NASA. The sixty-six identified archetypes contain patterns of conditions such as complexity and human-machine interaction associated with the formation of unintended consequences. To validate the archetypes, a sample of the archetypes is represented using system dynamics in order to illustrate that the identified archetypes are specialized versions of known high-level archetypes of unintended consequences. While the research is based upon a specific dataset, the archetypes apply to any engineered system and the pattern of leading indicators open a new path to manage unintended consequences and mitigate the magnitude of potentially adverse outcomes.

Hannah S Walsh↗

IM3 Projected U.S. Western Interconnection Grid Stress Dataset

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from GO WEST and TEP) for Integrated Multisector, Multiscale Modeling (IM3) Phase 2 simulations across eight different scenarios for the U.S. Western Interconnection through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States from a set of Thermodynamic Global Warming (TGW) simulations. These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 GO WEST is an open-source power grid modeling framework for the U.S. Western Interconnection, which allows users to tailor the model depending on their research study and science questions. It covers 28 balancing authorities (BAs) and 12 states in U.S. Western Interconnection. GO WEST allows users to select different number of nodes and come up with a simplified network by utilizing 10,000 nodal topology of the U.S. Western Interconnection (ACTIVSg10k). Users can select different number of nodes, mathematical formulations (linear programming vs. mixed-integer linear programming), transmission line limit scaling factors, and hurdle rate scaling factors. GO WEST offers a unit commitment and economic dispatch (UC/ED) module to simulate grid operations on an hourly scale. In this sense, users can calibrate and validate their model versions by comparing model outputs to historical datasets. TEP is an open-source transmission capacity expansion model, built on the GO WEST framework. It utilizes linear programming to optimize transmission capacity addition investment on existing lines within the GO WEST framework. The TEP model only increases the thermal capacity of existing transmission lines and does not add new lines to the system, which leaves the topology preserved. In order to use TEP model, users need to create scenarios with the GO WEST framework. Please refer to README file for a detailed description of the dataset including individual files and references.

Capacity Expansion Model↗

Improved representation of agricultural land use and crop management for large-scale hydrological impact simulation in Africa using SWAT+

To date, most regional and global hydrological models either ignore the representation of cropland or consider crop cultivation in a simplistic way or in abstract terms without any management practices. Yet, the water balance of cultivated areas is strongly influenced by applied management practices (e.g. planting, irrigation, fertilization, and harvesting). The SWAT+ (Soil and Water Assessment Tool) model represents agricultural land by default in a generic way, where the start of the cropping season is driven by accumulated heat units. However, this approach does not work for tropical and subtropical regions such as sub-Saharan Africa, where crop growth dynamics are mainly controlled by rainfall rather than temperature. In this study, we present an approach on how to incorporate crop phenology using decision tables and global datasets of rainfed and irrigated croplands with the associated cropping calendar and fertilizer applications in a regional SWAT+ model for northeastern Africa. We evaluate the influence of the crop phenology representation on simulations of leaf area index (LAI) and evapotranspiration (ET) using LAI remote sensing data from Copernicus Global Land Service (CGLS) and WaPOR (Water Productivity through Open access of Remotely sensed derived data) ET data, respectively. Results show that a representation of crop phenology using global datasets leads to improved temporal patterns of LAI and ET simulations, especially for regions with a single cropping cycle. However, for regions with multiple cropping seasons, global phenology datasets need to be complemented with local data or remote sensing data to capture additional cropping seasons. In addition, the improvement of the cropping season also helps to improve soil erosion estimates, as the timing of crop cover controls erosion rates in the model. With more realistic growing seasons, soil erosion is largely reduced for most agricultural hydrologic response units (HRUs), which can be considered as a move towards substantial improvements over previous estimates. We conclude that regional and global hydrological models can benefit from improved representations of crop phenology and the associated management practices. Future work regarding the incorporation of multiple cropping seasons in global phenology data is needed to better represent cropping cycles in areas where they occur using regional to global hydrological models.

crop phenology↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗