Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Integrating NASA Satellite Data Into USDA World Agricultural Outlook Board Decision Making Environment To Improve Agricultural Estimates

The USDA World Agricultural Outlook Board (WAOB) is responsible for monitoring weather and climate impacts on domestic and foreign crop development. One of WAOB's primary goals is to determine the net cumulative effect of weather and climate anomalies on final crop yields. To this end, a broad array of information is consulted. The resulting agricultural weather assessments are published in the Weekly Weather and Crop Bulletin, to keep farmers, policy makers, and commercial agricultural interests informed of weather and climate impacts on agriculture. The goal of the current project is to improve WAOB estimates by integrating NASA satellite precipitation and soil moisture observations into WAOB's decision making environment. Precipitation (Level 3 gridded) is from the TRMM Multi-satellite Precipitation Analysis (TMPA). Soil moisture (Level 2 swath and Level 3 gridded) is generated by the Land Parameter Retrieval Model (LPRM) and operationally produced by the NASA Goddard Earth Sciences Data and Information Services Center (GBS DISC). A root zone soil moisture (RZSM) product is also generated, via assimilation of the Level 3 LPRM data by a land surface model (part of a related project). Data services to be available for these products include GeoTIFF, GDS (GrADS Data Server), WMS (Web Map Service), WCS (Web Coverage Service), and NASA Giovanni. Project benchmarking is based on retrospective analyses of WAOB analog year comparisons. The latter are between a given year and historical years with similar weather patterns and estimated crop yields. An analog index (AI) was developed to introduce a more rigorous, statistical approach for identifying analog years. Results thus far show that crop yield estimates derived from TMPA precipitation data are closer to measured yields than are estimates derived from surface-based precipitation measurements. Work is continuing to include LPRM surface soil moisture data and model-assimilated RZSM.

Teng, William↗

Wavefront Sensing for WFIRST with a Linear Optical Model

In this paper we develop methods to use a linear optical model to capture the field dependence of wavefront aberrations in a nonlinear optimization-based phase retrieval algorithm for image-based wavefront sensing. The linear optical model is generated from a ray trace model of the system and allows the system state to be described in terms of mechanical alignment parameters rather than wavefront coefficients. This approach allows joint optimization over images taken at different field points and does not require separate convergence of phase retrieval at individual field points. Because the algorithm exploits field diversity, multiple defocused images per field point are not required for robustness. Furthermore, because it is possible to simultaneously fit images of many stars over the field, it is not necessary to use a fixed defocus to achieve adequate signal-to-noise ratio despite having images with high dynamic range. This allows high performance wavefront sensing using in-focus science data. We applied this technique in a simulation model based on the Wide Field Infrared Survey Telescope (WFIRST) Intermediate Design Reference Mission (IDRM) imager using a linear optical model with 25 field points. We demonstrate sub-thousandth-wave wavefront sensing accuracy in the presence of noise and moderate undersampling for both monochromatic and polychromatic images using 25 high-SNR target stars. Using these high-quality wavefront sensing results, we are able to generate upsampled point-spread functions (PSFs) and use them to determine PSF ellipticity to high accuracy in order to reduce the systematic impact of aberrations on the accuracy of galactic ellipticity determination for weak-lensing science.

Jurling, Alden S.↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

A Generalized approach to the operationalization of Software Quality Models

Comprehensive measures of quality are a research imperative, yet the development of software quality models is a wicked problem. Definitive solutions do not exist and quality is subjective at its most abstract. Definitional measures of quality are contingent on a domain, and even within a domain, the choice of representative characteristics to decompose quality is subjective. Thus, the operationalization of quality models brings even more challenges. A promising approach to quality modeling is the use of hierarchies to represent characteristics, where lower levels of the hierarchy represent concepts closer to real-world observations. Building upon prior hierarchical modeling approaches, we developed the Platform for Investigative software Quality Understanding and Evaluation (PIQUE). PIQUE surmounts several quality modeling challenges because it allows modelers to instantiate abstract hierarchical models in any domain by leveraging organizational tools tailored to their specific contexts. Here, we introduce PIQUE; exemplify its utility with two practical use cases; address challenges associated with parameterizing a PIQUE model; and describe algorithmic techniques that tackle normalization, aggregation, and interpolation of measurements.

Data aggregation↗

Electra: A Modular-Based Expansion of NASA's Supercomputing Capability

NASA has increasingly relied on high-performance computing (HPC) re- sources for computational modeling, simulation, and data analysis to meet the science and engineering goals of its missions in space exploration, aeronautics, and Earth and space science. The NASA Advanced Supercomputing (NAS) Division at Ames Research Center in Silicon Valley, Calif., hosts NASA’s premier supercomputing resources, integral to achieving and enhancing the success of the agency’s missions. NAS provides a balanced environment, funded under the High-End Computing Capability (HECC) project, comprised of world-class supercomputers, including its flagship distributed-memory cluster, Pleiades; high-speed networking; and massive data storage facilities, along with multi-disciplinary support teams for user support, code porting and optimization, and large-scale data analysis and scientific visualization. However, as scientists have increased the fidelity of their simulations and engineers are conducting larger parameter-space studies, the requirements for supercomputing resources have been growing by leaps and bounds. With the facility housing the HECC systems reaching its power and cooling capacity, NAS undertook a prototype project to investigate an alternative approach for housing supercomputers. Modular supercomputing, or container-based computing, is an innovative concept for expanding NASA’s HPC capabilities. With modular supercomputing, additional containers—similar to portable storage pods—can be connected together as needed to accommodate the agency’s ever-increasing demand for computing resources. In addition, taking advantage of the local weather permits the use of cooling technologies that would additionally save energy and reduce annual water usage. The first stage of NASA’s Modular Supercomputing Facility (MSF) prototype, which resulted in a 1,000 square-foot module on a concrete pad with room for 16 compute racks, was completed in Fall 2016 and an SGI (now HPE) computer system, named Electra, was deployed there in early 2017. Cooling is performed via an evaporative system built into the module, and preliminary experience shows a Power Usage Effectiveness (PUE) measurement of 1.03. Electra achieved over a petaflop on the LINPACK benchmark, sufficient to rank number 96 on the November 2016 TOP500 list [14]. The system consists of 1,152 InfiniBand-connected Intel Xeon Broadwell-based nodes. Its users access their files on a facility-wide file system shared by all HECC compute assets via Mellanox MetroX InfiniBand extenders, which connect the Electra fabric to Lustre routers in the primary facility over fiber-optic links about 900 feet long. The MSF prototype has exceeded expectations and is serving as a blueprint for future expansions. In the remainder of this chapter, we detail how modular data center technology can be used to expand an existing compute resource. We begin by describing NASA’s requirements for supercomputing and how resources were provided prior to the integration of the Electra module-based system.

Biswas, Rupak↗

The Earth in Living Color - NASA’s Surface Biology and Geology Designated Observable

The Surface Biology and Geology (SBG) Designated Observable will transform our understanding of the global land surface, inland and coastal aquatic ecosystems through visible-to-shortwave infra-red imaging (VSWIR) spectroscopy and thermal infra-red (TIR) imaging. SBG is one of four high-priority observables recommended in the 2017 NASA Earth Science Decadal Survey t o address science questions on vegetation and aquatic ecosystem health, snow-cover dynamics, volcanic activity, and minerology. With a planned launch readiness date of 2028, SBG is currently in Pre-Phase A, with Level 1 requirements being developed for a two-spacecraft architecture, including an additional constellation pathfinder. The recommended architecture emerged from an extensive study (2018-2021) that engaged the research and applications community to consider the science questions and measurement objectives of the Decadal Survey. A Science and Applications Traceability Matrix was used as a basis for scoring candidate architectures, with inputs from four working groups that covered algorithms, applications, calibration and validation, and modeling. Two pathfinder studies, Modeling End-to-End Traceability in support of SBG (MEET-SBG) and Space-based Imaging Spectroscopy and Thermal pathfindER (SISTER) are providing pre-launch modeling tools and data for algorithm development to support science value trades. The architecture consists of one spacecraft hosting a wide-swath VSWIR imaging spectrometer providing 30-m ground-sample distance (GSD), a spectral range of 380-2500 nm (at 10 nm resolution), 16-day revisit with 400 signal-to-noise for VNIR and 250 for SWIR (at 25% reflectance). A separate spacecraft will host a wide swath thermal imager, with five to seven bands placed between 4-12 μm), with 60-m (GSD), 3- day revisit, and 0.2K noise-equivalent differential temperature (NeDT). A VNIR compact camera will be hosted on the TIR spacecraft to enable coincident TIR and VNIR observations. A constellation pathfinder will evaluate options for enabling VSWIR mission continuity using Small Sats or data buys. Partnerships with international space agencies contribute technology as well as improvements to temporal revisit. SBG, when launched, will be the first dedicated mission collecting the full spectra of the Earth’s ‘living color’ and will play a critical role in NASA’s Earth System Observatory.

David S Schimel↗

ncompare: A Python Package for Comparing netCDF Structures

Earth science researchers and data engineers have a common problem: they often need to compare data files to see what is different between them. A lot of time is spent developing code to test differences. When it comes to comparing multidimensional data file formats like netCDFs (Network Common Data Form), this is particularly challenging and time-consuming, since there is frequently a need to evaluate the differences between dimension sizes, variable structures, and variable attributes, especially for regression testing. Since netCDFs are widely used in Earth science — with climate models, oceanographic or atmospheric reanalyses, and observational data — improved means of evaluating netCDF files can help enable a wide range of applications. We have developed a reusable open source approach through `ncompare`, which is a Python package for comparing netCDF structures [[https://github.com/nasa/ncompare]]. The `ncompare` tool compares the structure of two Network Common Data Form (NetCDF) files at the command line. It facilitates rapid comparisons by generating a formatted display of the matching and non-matching groups, variables, and associated metadata between two NetCDF datasets. The user has the option to colorize the terminal output for ease of viewing, and `ncompare` can optionally save comparison reports in text, comma-separated value (CSV), and/or Microsoft Excel formats. Despite the availability of tools (such as ncmpidiff or nccmp) that compare the values of variables, there was not previously a readily available, Python-based tool for rapid visual comparisons of group and variable structures, attributes, and chunking. `ncompare` was developed at NASA’s Atmospheric Science Data Center (ASDC) and is a collaboration with NASA Openscapes [[https://nasa-openscapes.github.io]] mentors across 11 of NASA’s data centers. Openscapes’ overarching vision is to support scientific researchers using NASA Earthdata as they migrate their workflows to the cloud. Relevant links: - https://github.com/nasa/ncompare - https://github.com/pyOpenSci/software-submission/issues/146 - https://nasa-openscapes.github.io

Daniel Kaufman↗

Mars Global Reference Atmospheric Model (Mars-GRAM 2005) Applications for Mars Science Laboratory Mission Site Selection Processes

The new Mars-GRAM auxiliary profile capability, using data from TES observations, mesoscale model output, or other sources, allows a potentially higher fidelity representation of the atmosphere, and a more accurate way of estimating inherent uncertainty in atmospheric density and winds. Figure 3 indicates that, with nominal value rpscale=1, Mars-GRAM perturbations would tend to overestimate observed or mesoscale-modeled variability. To better represent TES and mesoscale model density perturbations, rpscale values as low as about 0.4 could be used. Some trajectory model implementations of Mars-GRAM allow the user to dynamically change rpscale and rwscale values with altitude. Figure 4 shows that an mscale value of about 1.2 would better replicate wind standard deviations from MRAMS or MMM5 simulations at the Gale, Terby, or Melas sites. By adjusting the rpscale and rwscale values in Mars-GRAM based on figures such as Figure 3 and 4, we can provide more accurate end-to-end simulations for EDL at the candidate MSL landing sites.

Justh, H. L.↗

Systems and Methods for Advanced Rapid Imaging and Analysis for Earthquakes

Many embodiments provide a hybrid data processing system (HySDS) of an end-to-end geodetic imaging data system enabling near-real-time science, assessment, response, and rapid recovery. The HySDS may be an operation data processing system that integrates data from many different geodetic data sources and/or sensors, including interferometric synthetic aperture radar (InSAR), GPS, pixel tracking, seismology, and/or modeling, and processes the data to generate actionable high quality science data products. The HySDS may provide for an automated imaging and analysis capabilities that is able to handle the imminent increases in raw data from new and existing geodetic monitoring sensor systems.

Owen, Susan Ethel↗

What Preparatory Science is Needed in Coronal Structure and Activity

Solar Orbiter and Solar Probe Plus will launch in six short years! Before then, we need to accomplish a great deal of science in order to be able to maximize the return of these missions. Preparatory science is especially important for exploratory missions such as SO and SPP, because they truly will be going "where no mission has gone before". Such preparatory science may include all types of research: theory, modeling, data exploitation, and supporting observations. This meeting provides an opportunity for the community to define and begin this critical preparatory work. In this talk I will provide an overview of our state of knowledge in coronal structure and activity, describe what I believe are the most promising opportunities for advances by SO and SPP, and lead a discussion on what programs need to be implemented now in order to achieve these science advances by the time SO and SPP launch.

Antiochos, S. K.↗

Semantic Web Data Discovery of Earth Science Data at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC)

Mirador is a web interface for searching Earth Science data archived at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). Mirador provides keyword-based search and guided navigation for providing efficient search and access to Earth Science data. Mirador employs the power of Google's universal search technology for fast metadata keyword searches, augmented by additional capabilities such as event searches (e.g., hurricanes), searches based on location gazetteer, and data services like format converters and data sub-setters. The objective of guided data navigation is to present users with multiple guided navigation in Mirador is an ontology based on the Global Change Master directory (GCMD) Directory Interchange Format (DIF). Current implementation includes the project ontology covering various instruments and model data. Additional capabilities in the pipeline include Earth Science parameter and applications ontologies.

Hegde, Mahabaleshwara↗

Elastic Bayesian Model Calibration

Functional data are ubiquitous in scientific modeling. For instance, quantities of interest are modeled as functions of time, space, energy, density, etc. Uncertainty quantification methods for computer models with functional response have resulted in tools for emulation, sensitivity analysis, and calibration that are widely used. However, many of these tools do not perform well when the computer model’s parameters control both the amplitude variation of the functional output and its alignment (or phase variation). This paper introduces a framework for Bayesian model calibration when the model responses are misaligned functional data. The approach generates two types of data out of the misaligned functional responses: (1) aligned functions so that the amplitude variation is isolated and (2) warping functions that isolate the phase variation. These two types of data are created for the computer simulation data (both of which may be emulated) and the experimental data. The calibration approach uses both types so that it seeks to match both the amplitude and phase of the experimental data. The framework is careful to respect constraints that arise, especially when modeling phase variation, and is framed in a way that it can be done with readily available calibration software. In conclusion, we demonstrate the techniques on two simulated data examples and on two dynamic material science problems: a strength model calibration using flyer plate experiments and an equation of state model calibration using experiments performed on the Sandia National Laboratories’ Z-machine.

97 MATHEMATICS AND COMPUTING↗

Reanalysis of Rodent Data from Spacelab Life Sciences-1

The space bioscience field has long been plagued by the challenge of spaceflight with effects of radiation and microgravity. Having multiple and repeated spaceflight experiments for model organisms to solve these space stressors is costly and time consuming. Therefore, reusing and reanalyzing legacy experiments is one way that scientists can draw new conclusions in a timely manner and without using too many resources. Moreover, advances in general biological knowledge allows legacy experiments to be placed into more complete context.Here we aim to analyze all data and metadata taken from rats flown on the SLS-1 mission to create a comprehensive biological model that can be supplemented with current data to allow new discoveries in how space flown organisms adapt to the space environment. Our approach begins with the identification of all the data and metadata, including graphs and tables, for SLS-1 in NASA archives and other sources. Then, each piece of data and metadata will be digitized, reformatted and analyzed. Lastly, a previously developed astronaut model will be used to create the data framework and a comprehensive biological rodent model. The datasets we are using is from the 1991 SpaceLab Life Science 1 (SLS-1) NASA Mission. This was the first designated spacelab mission flown. All 29 rodents were tested for nine days in two different habitats: Research Animal Holding Facility (RAHF) and Animal Enclosure Module (AEM). The rodents were prepared for a live return and compared to a ground control. A total of 30 rodent experiments were accepted as flight studies on the mission. By digitization and reorganizing SLS-1 rat data we will both directly generate new insights and indirectly enable other scientists to by providing the data and metadata in a digitized form.

Space Biology↗

Reanalysis of Rodent Data from Spacelab Life Science-1

The space bioscience field has long been plagued by the challenge of spaceflight with effects of radiation and microgravity. Having multiple and repeated spaceflight experiments for model organisms to solve these space stressors is costly and time consuming. Therefore, reusing and reanalyzing legacy experiments is one way that scientists can draw new conclusions in a timely manner and without using too many resources. Moreover, advances in general biological knowledge allows legacy experiments to be placed into more complete context.Here we aim to analyze all data and metadata taken from rats flown on the SLS-1 mission to create a comprehensive biological model that can be supplemented with current data to allow new discoveries in how space flown organisms adapt to the space environment. Our approach begins with the identification of all the data and metadata, including graphs and tables, for SLS-1 in NASA archives and other sources. Then, each piece of data and metadata will be digitized, reformatted and analyzed. Lastly, a previously developed astronaut model will be used to create the data framework and a comprehensive biological rodent model. The datasets we are using is from the 1991 SpaceLab Life Science 1 (SLS-1) NASA Mission. This was the first designated spacelab mission flown. All 29 rodents were tested for nine days in two different habitats: Research Animal Holding Facility (RAHF) and Animal Enclosure Module (AEM). The rodents were prepared for a live return and compared to a ground control. A total of 30 rodent experiments were accepted as flight studies on the mission. By digitization and reorganizing SLS-1 rat data we will both directly generate new insights and indirectly enable other scientists to by providing the data and metadata in a digitized form.

Space Biology↗

SOHIP Abel Transform and Onion Peeling Model Module

This software provides tools for analyzing and modeling physical systems using mathematical transforms and layered models. It includes (1) functions for performing the Abel transform, which is used to relate measurements of bending angles to properties such as refractive index and radius in a medium. The code can compute bending angles from input profiles and also reconstruct these profiles from observed data; (2) the functions for modeling systems with multiple layers using an onion-peeling approach, allowing users to simulate and analyze the behavior of layered materials or structures. These capabilities are useful for researchers and engineers working in fields such as optics, atmospheric science, and materials analysis, enabling them to interpret and model data from experiments or simulations relates to refraction in spherical symmetric medium.

Xu, Shuang [Lawrence Livermore National Laboratory↗

Advancing Open Science in Atmospheric Research: Integrating Data Usability and Machine Learning

In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences." "In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences.

Jennifer Wei↗

The AE-8 trapped electron model environment

The machine sensible version of the AE-8 electron model environment was completed in December 1983. It has been sent to users on the model environment distribution list and is made available to new users by the National Space Science Data Center (NSSDC). AE-8 is the last in a series of terrestrial trapped radiation models that includes eight proton and eight electron versions. With the exception of AE-8, all these models were documented in formal reports as well as being available in a machine sensible form. The purpose of this report is to complete the documentation, finally, for AE-8 so that users can understand its construction and see the comparison of the model with the new data used, as well as with the AE-4 model.

Vette, James I.↗