Search NASA⌕ Search

SEARCH · Search NASA

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Open Science Approach to Analyze Climate-Crop Relationships in the US Leveraging GES DISC and Galaxy Workflows

Understanding the intricate relationship between climate variability and agricultural production is crucial for ensuring food security. This study investigates the impact of climate parameters, such as temperature, precipitation, and soil moisture, on major US crop yields. Adopting an open science approach, the study analyzes the impact of climate on agricultural production in the United States. The Galaxy workflow engine serves as the primary tool for integrating climate data from the Goddard Earth Sciences Data and Information Services Center (GES DISC), retrieved via the Giovanni system, with yield statistics from the United States Department of Agriculture’s National Agricultural Statistics Service (USDA NASS). Extensions for reading, preprocessing, and analyzing external data have been developed, enabling the creation of workflows within the Galaxy platform. The development of a reproducible workflow allows for the calculation of seasonal climate averages, which are then assessed for their correlation with crop yields. This methodology ensures the replicability of the research, promoting transparency and collaboration in the scientific community. Correlational and regression analyses have been applied to different sub-zones and crops. The findings from this research offer valuable insights into the relationship between climate parameters and crop yields. These insights contribute to a deeper understanding of climate-crop relationships, providing a solid foundation for informed decision-making in the agricultural sector. The high correlation values indicate a significant relationship between climate parameters and crop yields, underscoring the importance of considering climate factors in agricultural planning and policymaking. This research also exemplifies the power of open science in advancing our understanding of complex environmental and agricultural phenomena. By leveraging open data and services, it provides a robust and replicable framework for future studies in this critical field.

Open science↗

Open-Source Science-led Development of the Atmosphere Observing System (AOS) Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David Giles↗

Recovery, Restoration and Archiving of Previously Lost Data and Metadata from the Apollo Lunar Surface Experiments Package (ALSEP)

The Apollo Lunar Surface Experiments Package (ALSEP) is the name used to collectively represent the geophysical instruments deployed on the lunar surface by the astronauts on Apollo 12, 14, 15, 16, and 17. These instruments were active from the times of their deployment (November 1969 – December 1972) to September 1977. During that time, fourteen types of experiments were conducted, and their data were transmitted to Earth. The experiment PIs processed them. At the conclusion of the experiments, some of these data were submitted to the NASA Space Science Data Coordinated Archive (NSSDCA) for archiving, while others were not. The raw instrument data received from the Moon prior to March 1976 were not archived, either. The unarchived data, resided on open-reel magnetic tapes, became lost in the decades since, along with much of the metadata (the information necessary/useful in properly processing/analyzing the data). This article retraces the history of the ALSEP data archiving efforts in the 1970s, the subsequent loss of the data tapes, and the search, recovery, and restoration of the lost data by contemporary researchers in the 21st century. In 2006, NSSDCA began reformatting some of the ALSEP data archived in the 1970s to conform with the current Planetary Data System (PDS). In 2010, 440 of the previously lost magnetic tapes containing the raw ALSEP data were recovered. From these tapes, the data were extracted, re-packaged for individual experiments, and, for those with sufficient metadata, processed into higher order data readily usable by researchers. All of these data products have been recently archived with either PDS or NSSDCA. These newly restored data fill a number of gaps in the previously existing archive of the ALSEP data. In addition, tens of thousands of pages of Apollo era documents have been optically scanned and compiled into an online searchable catalog. This article also describes the content, organization, and usage of the restored raw ALSEP data and metadata.

S Nagihara↗

ENDFtk: A robust tool for reading and writing ENDF-formatted nuclear data

ENDFtk is a recently developed C++ and Python interface to interact with ENDF-6 formatted nuclear data files. It provides a robust and complete interface, allowing the reading and writing of all formats currently part of the ENDF-6 formats manual, as well as some non-ENDF formats used by the NJOY processing code. It provides an interface that mimics the names in the ENDF-6 formats manual as well as an equivalent interface using human-readable attribute names. It is robust and powerful enogh for nuclear data experts to develop complex applications, while also simple enough to be used non-experts to retrieve and manipulate evaluated nuclear data. ENDFtk offers the ability to easily interrogate and manipulate data either in large-scale code projects or in simple Python scripts. Here, in this paper, a brief overview of the interface is given, as well as more substantial examples demonstrating plotting simple data, interacting with more complex data, and writing new data to files. ENDFtk is open source and available for download via GitHub (https://github.com/njoy/ENDFtk).

97 MATHEMATICS AND COMPUTING↗

Improvements in lake volume predictions using Landsat data

A cumulative error in the water balance budget for Lake Okeechobee produces a one million acre-foot discrepancy in the predicted water volume over a 4-year period. The major source of error appears to be complex shoreline marshes that comprise 20 percent of the lake surface. The water balance budget model presently treats these marshes as open water. Using Landsat data, the vegetation in the lake's littoral zone was classified multispectrally to provide a data base for determining water budget information. First, the acreage of a given plant species in the littoral zone was obtained with satellite data. Second, the surface area occupied by plants (which therefore could not be considered open water) was used to adjust the vegetation acreage giving an effective water surface. Based on this information, more detailed representations of evapotranspiration and total water surface (and hence total lake volume) could be provided to the water balance budget computation.

Gervin, J. C.↗

Pretraining Billion-Scale Geospatial Foundational Models on Frontier

As AI workloads increase in scope, generalization capability becomes challenging for small task-specific models and their demand for large amounts of labeled training samples increases. On the contrary, Foundation Models (FMs) are trained with internet-scale unlabeled data via self-supervised learning and have been shown to adapt to various tasks with minimal fine-tuning. Although large FMs have demonstrated significant impact in natural language processing and computer vision, efforts toward FMs for geospatial applications have been restricted to smaller size models, as pretraining larger models requires very large computing resources equipped with state-of-the-art hardware accelerators. Current satellite constellations collect 100+TBs of data a day, resulting in images that are billions of pixels and multimodal in nature. Such geospatial data poses unique challenges opening up new opportunities to develop FMs. We investigate billion scale FMs and HPC training profiles for geospatial applications by pretraining on publicly available data. We studied from end-to-end the performance and impact in the solution by scaling the model size. Our larger 3B parameter size model achieves up to 30% improvement in top1 scene classification accuracy when comparing a 100M parameter model. Moreover, we detail performance experiments on the Frontier supercomputer, America's first exascale system, where we study different model and data parallel approaches using PyTorch's Fully Sharded Data Parallel library. Specifically, we study variants of the Vision Transformer architecture (ViT), conducting performance analysis for ViT models with size up to 15B parameters. By discussing throughput and performance bottlenecks under different parallelism configurations, we offer insights on how to leverage such leadership-class HPC resources when developing large models for geospatial imagery applications.

Tsaris, Aristeidis (aris)↗

Tropical Cyclone Intensity Estimation Using Deep Convolutional Neural Networks

Estimating tropical cyclone intensity by just using satellite image is a challenging problem. With successful application of the Dvorak technique for more than 30 years along with some modifications and improvements, it is still used worldwide for tropical cyclone intensity estimation. A number of semi-automated techniques have been derived using the original Dvorak technique. However, these techniques suffer from subjective bias as evident from the most recent estimations on October 10, 2017 at 1500 UTC for Tropical Storm Ophelia: The Dvorak intensity estimates ranged from T2.3/33 kt (Tropical Cyclone Number 2.3/33 knots) from UW-CIMSS (University of Wisconsin-Madison - Cooperative Institute for Meteorological Satellite Studies) to T3.0/45 kt from TAFB (the National Hurricane Center's Tropical Analysis and Forecast Branch) to T4.0/65 kt from SAB (NOAA/NESDIS Satellite Analysis Branch). In this particular case, two human experts at TAFB and SAB differed by 20 knots in their Dvorak analyses, and the automated version at the University of Wisconsin was 12 knots lower than either of them. The National Hurricane Center (NHC) estimates about 10-20 percent uncertainty in its post analysis when only satellite based estimates are available. The success of the Dvorak technique proves that spatial patterns in infrared (IR) imagery strongly relate to tropical cyclone intensity. This study aims to utilize deep learning, the current state of the art in pattern recognition and image recognition, to address the need for an automated and objective tropical cyclone intensity estimation. Deep learning is a multi-layer neural network consisting of several layers of simple computational units. It learns discriminative features without relying on a human expert to identify which features are important. Our study mainly focuses on convolutional neural network (CNN), a deep learning algorithm, to develop an objective tropical cyclone intensity estimation. CNN is a supervised learning algorithm requiring a large number of training data. Since the archives of intensity data and tropical cyclone centric satellite images is openly available for use, the training data is easily created by combining the two. Results, case studies, prototypes, and advantages of this approach will be discussed.

tropical cyclone intensity↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth independence and autonomy of mission operations. Here we present an overview of AI/ML architecture to support deep space mission goals, developed with leaders in the field. First, we focus on the fundamental biological research that supports our understanding of physiological responses to spaceflight, and we describe current efforts to support AI/ML research including data standardization and data engineering through maximally open and FAIR (findable, accessible, interoperable, reusable) databases and the generation of AI-ready datasets for reuse and analysis. We also discuss remote data management frameworks for research data as well as environmental and health data that are generated during deep space missions. We highlight several research projects that leverage data standardization and management for fundamental biological discovery to uncover the complex effects of space travel on living systems. Next, we provide an overview of cutting-edge AI/ML approaches that can be integrated to support remote monitoring and analysis during deep space missions, including generative models and large language models to learn the underlying biomedical patterns and predict outcomes or answer questions during off world medical scenarios. We also describe current AI/ML methods to support this research and monitoring through automated cloud-based labs which enable limited human intervention and closed-loop experimentation in remote settings. These labs could support mission autonomy by analyzing environmental data streams, and would be facilitated through in situ analytics capabilities to avoid sending large raw data files through low bandwidth communications. Finally, in the context of deep space missions with limited communications or access to medical advice from Earth, we describe a solution for integrated, real-time mission biomonitoring across hierarchical levels from continuous environmental monitoring, to wearables and point-of-care devices, to molecular and physiological monitoring. We introduce a precision space health system that will ensure that the future of space health is predictive, preventative, participatory and personalized.

artificial intelligence↗

Using Big Data Technologies with Earth Science Data in HDF5: HDF5 Scalable Solutions

HDF5 (Hierarchical Data Format 5) is open-source, high-performance software that consists of an abstract data model, library, and fileformat used for storing and managing extremely large and/or complex data collections. NASA Earth Observing System (EOS) Data and Information Systems use HDF5 as an archival format to store remote sensing data from EOS satellites. HDF5 is also used to store other types of Geoscience and Strophysical data, e.g., seismic data and data from Low-Frequency Array (LOFAR) radio telescopes. Data stored in HDF5 has reached tens of petabytes and is growing at an accelerated rate.With the growing amout of HDF5 Earth Science data to analyze and process, scientists need to adopt big data technologies including new storage paradigms such as cloud and object storage. To run models and perform data analysis they also need to utilizied efficient and diverse ways to access data, from high-performance computing's (HPC) Message Passing Interface (MPI) I/O and deep memory hierarchies (DMH) to non-HPC frameworks such as Apache Hadoop, Spark, and Drill. The HDF Group continually works to enable usage of big data technologies in HDF software.

Knox, Larry↗

IMAGE Software Suite

The IMAGE Mission is generating a truely unique set of magnetospheric measurement through a first-of-its-kind complement of remote, global observations. These data are being distributed in the Universal Data Format (UDF), which consists of data, calibration, and documentation. This is an open dataset, available to all by request to the National Space Science Data Center (NSSDC) at NASA Goddard Space Flight Center. Browse data, which consists of summary observations, is also available through the NSSDC in the Common Data Format (CDF) and graphic representations of the browse data. Access to the browse data can be achieved through the NSSDC CDAWeb services or by use of NSSDC provided software tools. This presentation documents the software tools, being provided by the IMAGE team, for use in viewing and analyzing the UDF telemetry data. Like the IMAGE data, these tools are openly available. What these tools can do, how they can be obtained, and how they are expected to evolve will be discussed.

Gallagher, Dennis L.↗

High-Resolution Meteorology with Climate Change Impacts from Global Climate Model Data Using Generative Machine Learning

As renewable energy generation increases, the impacts of weather and climate on energy generation and demand become critical to the reliability of the energy system. However, these impacts are often overlooked. Global climate models (GCMs) can be used to understand possible changes to our climate, but their coarse resolution makes them difficult to use in energy system modelling. Here we present open-source generative machine learning methods that produce meteorological data at a nominal spatial resolution of 4 km at an hourly frequency based on inputs from 100 km daily-average GCM data. These methods run 40 times faster than traditional downscaling methods and produce data that have high-resolution spatial and temporal attributes similar to historical datasets. We demonstrate that these methods can be used to downscale projected changes in wind, solar and temperature variables across multiple GCMs including projections for more frequent low-wind and high-temperature events in the Eastern United States.

climate change↗

2024 OES-Environmental 2024 State of the Science Report, Chapter 8: Marine Renewable Energy Data and Information Systems

As the marine renewable energy (MRE) sector grows, large amounts of environmental and technical data and information are being collected. When these data and information are openly available, they can be used to guide research and development, inform responsible siting and consenting of projects, and increase stakeholder understanding through transparency. For example, quality environmental data collected during the siting, consenting, construction, operation, and decommissioning of MRE projects can all play key roles in better characterizing baseline conditions, developing effective monitoring and mitigation strategies, and retiring environmental risks through data transferability (see Chapter 6). Ensuring that these data and information are easily discoverable and accessible will help the MRE sector make informed decisions and coexist in an increasingly busy ocean environment.

16 TIDAL AND WAVE POWER↗

An experimental study of the closure behavior of short and long cracks

Experiments were conducted on 2024-T3 aluminum specimens. The laser-based interferometric strain/displacement optical system was used to measure crack opening displacement. The displacement data were measured near the crack tip, at the crack mouth, and at other locations along the crack line for both short and long cracks. The dependence of opening load ratio on measurement locations was observed. The ratio decreased systematically as the measurement locations moved away from the crack tip. However, the ratio stayed roughly the same for the near tip surface measurements at various crack lengths.

Su, X.↗

Smokescreen: A Python package for data vector blinding and encryption in cosmological analyses

Smokescreen is an open-source Python library for data-vector concealment (blinding) in cosmological analyses. Data-vector blinding works by applying cosmology-dependent shifts to the observed data vector, moving it away from the true cosmological signal without affecting its statistical properties, so that analysts cannot infer the true result until the analysis is frozen and the blinding is lifted. The package computes these shifts using Firecrown likelihoods applied to data vectors stored in the SACC format, ensuring that the theoretical model used for blinding is identical to that used for inference whilst remaining agnostic to the specific observable being blinded. To prevent accidental unblinding, the original SACC file, containing the true cosmology, is encrypted. Although developed for the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), Smokescreen is applicable to any experiment using Firecrown likelihoods and the SACC data format.

Loureiro, Arthur [Stockholm U., OKC; Imperial Coll↗

Use of Open Networks and Delay-Tolerant Protocol to Decrease WAN Latency of EOS near Real-Time Data

Since 1999, NASA's Earth Observing System Data Operations System (EDOS) project at Goddard Space Flight Center (GSFC) has provided high-rate data capture, level zero processing, and product distribution services for a majority of NASA's EOS (Earth Observing System) high-rate missions, including Terra, Aqua, Aura, ICESat, EO-1, SMAP, and OCO-2. EDOS high-rate science and engineering (150-300 Mbps) data-driven capture systems are deployed at 7 worldwide ground stations which are connected via both private (closed) and public (open) wide area networks (WANs) to the centralized EDOS Level Zero Processing Facility (LZPF) located at GSFC, where the data is processed and Level 0 products are distributed to users worldwide. All data transferred over the open networks to GSFC traverse an IPSec tunnel, providing the same level of security as a VPN connection. EDOS produces both time-based and near real-time products (session-based). Near real-time data products are produced from a single ground station contact; time-based products are produced from multiple ground station contacts. EDOS is the primary supplier of EOS Level 0 data to the NASA near real-time user community known as the Land, Atmosphere Near real-time Capability for EOS (LANCE). For the past few years, EDOS has streamlined its systems to reduce WAN latency for near real-time data delivery, including implementing Quality of Service (QoS), expanding closed network bandwidth, adding open network connections with more bandwidth, and implementing a delay-tolerant protocol to mitigate long round-trip times to remote ground stations.

Delay-Tolerant Protocol↗

Data management in NOAA

The NOAA archives contain 150 terabytes of data in digital form, most of which are the high volume GOES satellite image data. There are 630 data bases containing 2,350 environmental variables. There are 375 million film records and 90 million paper records in addition to the digital data base. The current data accession rate is 10 percent per year and the number of users are increasing at a 10 percent annual rate. NOAA publishes 5,000 publications and distributes over one million copies to almost 41,000 paying customers. Each year, over six million records are key entered from manuscript documents and about 13,000 computer tapes and 40,000 satellite hardcopy images are entered into the archive. Early digital data were stored on punched cards and open reel computer tapes. In the late seventies, an advanced helical scan technology (AMPEX TBM) was implemented. Now, punched cards have disappeared, the TBM system was abandoned, most data stored on open reel tapes have been migrated to 3480 cartridges, many specialized data sets were distributed on CD ROM's, special archives are being copied to 12 inch optical WORM disks, 5 1/4 inch magneto-optical disks were employed for workstation applications, and 8 mm EXABYTE tapes are planned for major data collection programs. The rapid expansion of new data sets, some of which constitute large volumes of data, coupled with the need for vastly improved access mechanisms, portability, and improved longevity are factors which will influence NOAA's future systems approaches for data management.

Callicott, William M.↗