Search NASASearch

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Visualizing Geospatial Data through ESRI Story Maps for Earth Science Education: Lessons Learned from My NASA Data

For 20 years My NASA Data (MND) has curated NASA Earth science data and provided the data to educators in engaging learner-centered resources. MND has recently featured story maps as an innovative way to engage students in NASA Earth data. A story map is a cloud-based lesson that engages the learner in interactive geospatial maps using NASA data, and other multimedia content, text, and tasks that can be seamlessly incorporated in classroom instruction. This immersive technology eliminates the need for the user to move among tricky interfaces to access and visualize Earth science data, and no special software is required to be downloaded. Each story map integrates data from different NASA satellite missions, retrieved from Distributed Active Archive Centers (DAACs). Story maps also employ data analysis tools, such as time series options and swipe tools that allow learners to view and analyze relationships between scientific variables. MND has produced 25 story map lesson plans on the topics of air quality, the urban heat island effect, Earth’s energy budget, phytoplankton distribution, hurricane formation, solar eclipses, ocean circulation patterns, sea ice extent, and volcanic eruptions. Nine of them are extended story maps and written in the 5E format, which is internationally recognized as best practice based on how children learn science. Each story map resource is developed by the MND team featuring a GIS programming specialist, a lead scientist, and educational specialist/s to ensure the context, content, and methods are scientifically and educationally sound. The MND story maps are written for middle and high school science teachers and students as they connect with the Earth Systems Science phenomena featured in the Next Generation Science Standards. Each story map includes supporting resources for smooth integration in the classroom. During Fiscal Year 2023, The My NASA Data website received over 100,000 story map engagements during. These metrics highlight the interest in story maps as an Earth Science educational resource.

Desiray Wilson

A New Look at Data Usage by Using Metadata Attributes as Indicators of Data Quality

This study reviews the key metrics (users, distributed volume, and files) in multiple ways to gain an understanding of the significance of the metadata. Characterizing the usability of data by key metadata elements, such as discipline and study area, will assist in understanding how the user needs have evolved over time. The data usage pattern based on product level provides insight into the level of data quality. In addition, the data metrics by various services, such as the Open-source Project for a Network Data Access Protocol (OPeNDAP) and subsets, address how these services have extended the usage of data. Over-all, this study presents the usage of data and metadata by metrics analyses, which may assist data centers in better supporting the needs of the users.

metadata

Natural Language Processing for Extracting Rich Disease Data Aligned To Satellite Meteorological Data

Global climate change is redefining our understanding of how diseases spread. In Sri Lanka, vector-borne diseases such as dengue fever, encephalitis, and leptospirosis historically surged during the monsoon seasons when temperatures were high enough for mosquito eggs to hatch. Unfortunately, due to rising temperatures and more erratic rainfall patterns, mosquito eggs can now hatch year-round and are increasingly unpredictable, leading to an alarmingly increasing number of hospitalizations and deaths. More data is needed to adapt our response to these diseases in an increasingly warmer world. In the contemporary landscape, a wealth of disease information is available, yet accessibility remains limited due to unstructured data formats such as PDFs. Therefore, converting unstructured disease reports into structured formats is necessary for effectively leveraging data. This paper introduces a comprehensive framework for collecting unstructured disease reports and transforming them into analyzable formats. By creating separate models tailored to each data format, we can ensure accuracy compared to general models. These straightforward models enhance accessibility and empower other researchers to use our tools. The returned structured data can then be harnessed for analysis, statistical purposes, and informing evidence-based public health interventions, thus facilitating more informed decision-making in healthcare. We deploy this framework to produce geospatial data for Sri Lanka and Brazil for many different conditions and align these data with satellite environmental data, providing for the first time a structured, aligned powerful dataset for disease modeling.

Open Source Open Science

Transportation Secure Data Center: Frequently Asked Questions for Data Owners/Contributors

The Transportation Secure Data Center is a centralized repository for detailed transportation data from travel and transit surveys and studies conducted across the nation. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. Hundreds of datasets from surveys and studies of household travel and transit passenger travel are archived in the TSDC, including surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Detailed data from travel surveys and studies are extremely valuable for research purposes. However, the fine-grained information they contain could potentially be misused to identify individual travelers, so access to these data should only be granted with safeguards in place to protect participant privacy. The TSDC was created to address this challenge and to relieve public agencies from the burden of archiving their data and responding to data requests.

33 ADVANCED PROPULSION SYSTEMS

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML

Features of Point Clouds Synthesized from Multi-View ALOS/PRISM Data and Comparisons with LiDAR Data in Forested Areas

LiDAR waveform data from airborne LiDAR scanners (ALS) e.g. the Land Vegetation and Ice Sensor (LVIS) havebeen successfully used for estimation of forest height and biomass at local scales and have become the preferredremote sensing dataset. However, regional and global applications are limited by the cost of the airborne LiDARdata acquisition and there are no available spaceborne LiDAR systems. Some researchers have demonstrated thepotential for mapping forest height using aerial or spaceborne stereo imagery with very high spatial resolutions.For stereo imageswith global coverage but coarse resolution newanalysis methods need to be used. Unlike mostresearch based on digital surface models, this study concentrated on analyzing the features of point cloud datagenerated from stereo imagery. The synthesizing of point cloud data from multi-view stereo imagery increasedthe point density of the data. The point cloud data over forested areas were analyzed and compared to small footprintLiDAR data and large-footprint LiDAR waveform data. The results showed that the synthesized point clouddata from ALOSPRISM triplets produce vertical distributions similar to LiDAR data and detected the verticalstructure of sparse and non-closed forests at 30mresolution. For dense forest canopies, the canopy could be capturedbut the ground surface could not be seen, so surface elevations from other sourceswould be needed to calculatethe height of the canopy. A canopy height map with 30 m pixels was produced by subtracting nationalelevation dataset (NED) fromthe averaged elevation of synthesized point clouds,which exhibited spatial featuresof roads, forest edges and patches. The linear regression showed that the canopy height map had a good correlationwith RH50 of LVIS data with a slope of 1.04 and R2 of 0.74 indicating that the canopy height derived fromPRISM triplets can be used to estimate forest biomass at 30 m resolution.

LiDARD

The Process of Bringing Dark Data to Light: The Rescue of the Early Nimbus Satellite Data

Myriad environmental satellite missions are currently orbiting the earth. The comprehensive monitoring by these sensors provide scientists, policymakers, and the public critical information on the earths weather and climate system. The state of the art technology of our satellite monitoring system is the legacy of the first environment satellites, the Nimbus systems launched by NASA in the mid-1960s. Such early data can extend our climate record and provide important context in longer-term climate changes. However, the data was stowed away and, over the years, largely forgotten. It was nearly lost before its value was recognized and attempts to recover the data were undertaken. This paper covers what it took the authors to recover, navigate and reprocess the data into modern formats so that it could be used as a part of the satellite climate record. The procedures to recover the Nimbus data, from both film and tape, could be used by other data rescue projects, however the algorithms presented will tend to be Nimbus specific. Data rescue projects are often both difficult and time consuming but the data they bring back to the science community makes these efforts worthwhile.

dark

Microbial Optical Data Processing: A Key Step in the Metabolic Assessment of Lunar Explorer Instrument for Space Biology Applications (LEIA) and Biosentinel’s Payload Data

The BioSensor payload platform on BioSentinel and LEIA autonomously collects optical data from microbial model organisms in liquid culture. The BioSensor is designed to monitor metabolic activity using absorbance measurements of cell density and alamarBlue, a readily available colorimetric redox indicator dye. BioSentinel, a pioneering NASA CubeSat, uses yeast to study deep space radiation. LEIA investigates radiation and lunar gravity response. The experimental setup includes 16 wells equipped with three LEDs (570, 630, and 850 nm) and their corresponding photodetectors. One well is a calibration control without biology while the rest have desiccated cultures. Autonomous rehydration initiates the experiment. Data from the BioSensor are received from the flight and ground units, enabling comparison to uncover location-based metabolic rate variations. This study presents a Python Jupyter notebook developed for efficient data processing of multiple CSV files containing date and time columns, temperature, and well illumination data. It offers a user-friendly interface while maintaining computational power, automatically recognizing and iteratively processing data files in a user-input path. A Hampel filter with a short window eliminates outlier artifacts from sensor dropout. Because absorbance is a relative measurement, conversion from raw illumination requires defining a “blank” value, so the first data points are averaged to provide the necessary denominator. A cube-root function correction mitigates undesired drift caused by air pockets during the fluidic card filling phase, maintaining optical path length consistency. Beer-Lambert's law is applied to further convert absorbance values to cell and dye form concentrations, the desired science parameters. The processed data are saved and visualized as SVG plots. Future plans include extracting specific science parameters from the processed data like growth rate and metabolic rate, and identification of features corresponding to metabolic and phenotypic shifts such as starvation, shifts from aerobic to anaerobic growth, and osmotic stresses.

Space biology

Hierarchical Data Format for Earth Observing System Data Product Developer's Guide

The "Hierarchical Data Format for Earth Observing System" talk will address the best practices for creating ESDIS data products. The work presented is done in support of Data Product Developers Guide Working Group with mission "to help data product developers make data usable for end users". During the presentation, we will use some examples of NASA data products and show how to modify them to make data more usable.

Data usability

Evaluating the risk of data loss due to particle radiation damage in a DNA data storage system

DNA data storage is a potential alternative to magnetic tape for archival storage purposes, promising substantial gains in information density. Critical to the success of DNA as a storage media is an understanding of the role of environmental factors on the longevity of the stored information. In this paper, we evaluate the effect of exposure to ionizing particle radiation, a cause of data loss in traditional magnetic media, on the longevity of data in DNA data storage pools. We develop a mass action kinetics model to estimate the rate of damage accumulation in DNA strands due to neutron interactions with both nucleotides and residual water molecules, then utilize the model to evaluate the effect several design parameters of a typical DNA data storage scheme have on expected data longevity. Finally, we experimentally validate our model by exposing dried DNA samples to different levels of neutron irradiation and analyzing the resulting error profile. Our results show that particle radiation is not a significant contributor to data loss in DNA data storage pools under typical storage conditions.

97 MATHEMATICS AND COMPUTING

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES

Data Quality Assessment Process for Real-Time Data-Driven Traffic Microsimulation of Smart Corridor

Smart corridor digital twins are often created for the development and evaluation of emerging intelligent transportation systems and Connected and Autonomous Vehicle (CAV) technologies. However, limited guidance exists for data quality assessment for digital twin development. To address this, this paper discusses the data quality assessment utilized to develop data-driven real-time microscopic simulation models, i.e., digital twins, for two separate smart corridors: the North Avenue Smart Corridor in Atlanta, GA, and the Martin Luther King Smart Corridor in Chattanooga, Tennessee. This paper provides a summary of the author’s investigations of data requirements and data characteristics for the given smart corridor digital twin development efforts. With a focus on data, this summary includes a description of the data investigation process, key data issues observed, and strategies to address observed issues. Discussion is provided to help expand the lessons from these studies to other digital twin development efforts.

Saroj, Abhilasha [ORNL] (ORCID:0000000191178063)

DESI 2024: Constraints on physics-focused aspects of dark energy using DESI DR1 BAO data

Baryon acoustic oscillation data from the first year of the Dark Energy Spectroscopic Instrument (DESI) provide near percent-level precision of cosmic distances in seven bins over the redshift range z=0.1–4.2. Here, this paper is the follow-up to the original DESI BAO cosmology paper [A. G. Adame et al. (DESI Collaboration), arXiv:2404.03002], which considered the conventional w 0 w a cold dark matter (CDM) model. We use the novel DESI data, together with other cosmic probes, to constrain the background expansion history using some well-motivated physical classes of dark energy. In particular, we explore three physics-focused behaviors of dark energy from the equation of state and energy density perspectives: the thawing class (matching many simple quintessence potentials), emergent class (where dark energy comes into being recently, as in phase transition models), and mirage class [where phenomenologically the distance to cosmic microwave background (CMB) last scattering is close to that from a cosmological constant Λ despite dark energy dynamics]. All three classes fit the data at least as well as Λ ⁢CDM, and indeed can improve on it by Δ⁢χ 2 ≈ –5 to –17 for the combination of DESI BAO with CMB and supernova data while having one more parameter. The mirage class does essentially as well as w 0 ⁢w a CDM and exhibits moderate to strong Bayesian evidence preference with respect to Λ⁢ CDM. These classes of dynamical behaviors highlight worthwhile avenues for further exploration into the nature of dark energy.

79 ASTRONOMY AND ASTROPHYSICS

Data and Scripts Associated with "Modeling Ecohydrological Responses of Vegetation to Urban Microclimates Using the E3SM Land Model"

This dataset supports the study of vegetation ecohydrological responses to urban microclimates using the land component of the Energy Exascale Earth System Model (ELM) at four urban sites in Knoxville, Tennessee, USA. It includes the model inputs, simulation outputs, and associated scripts for running ELM simulations and analyzing the resulting data. The Model_Inputs folder includes static surface data, satellite-derived phenology (i.e., leaf area index), and atmospheric forcing data used to drive ELM simulations. Detailed descriptions of these datasets are provided in Section 2.3.2 of the associated manuscript. The Model_Outputs folder contains simulation results for the baseline, treatment, and ensemble experiments. Outputs from the baseline and treatment simulations are provided as raw ELM NetCDF files. Because the raw outputs from the 4,000-member ensemble are prohibitively large, the ensemble results are provided as summarized CSV files, which also serve as the source data for Figure 5 of the associated manuscript. The Scripts folder contains three components: E3SM, the core codebase of the Energy Exascale Earth System Model (E3SM); elm-olmt, the Offline Land Model Testbed (OLMT) used to perform the simulations; and knoxville_elm, which contains the analysis scripts used to process model outputs and generate the figures and results presented in the associated manuscript. Additional information is provided in Scripts_readme.txt within the Scripts directory.

Lu, Xiaoman [ORNL] (ORCID:0000000306698780)

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING

Offshore Geologic Carbon Storage Data Collection and Data Gaps Analysis

This is a TRS documenting the Offshore Geologic Carbon Storage Data Collection. It describes the Data Collection web application and its creation as well as an accompanying Data Gaps Assessment. We present an interactive data collection and data gaps analysis to aggregate, understand, and disseminate the data that are publicly available to support offshore GCS in the United States. This data collection and data gaps analysis can be leveraged by stakeholders to understand where GCS may be viable offshore, create GCS project analogs, and address challenges to GCS in offshore environments.

58 GEOSCIENCES

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds Developed lands Areas >0.8 km (0.5 miles) from developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall