Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Enhancing Drinking Water Quality Modeling: Leveraging Physics Informed Neural Networks for Learning with Imperfect Reaction Models and Partial Data

Chemical kinetics models, typically formulated as systems of ordinary or partial differential equations, are valuable tools for simulating drinking water quality. However, these models often face inaccuracies due to discrepancies between the laboratory and the real-world conditions, as well as limitations in experimental analytical methods, hindering the accurate representation of the true underlying chemical mechanisms. In this study, we propose a Physics Informed Neural Network (PINN), using the eXtreme Theory of Functional Connections, to improve the prediction of chemical concentrations over time. The PINN method accounts for imperfect chemical models and incorporates partial data to improve predictions. Focusing on reactions describing water disinfection residual and disinfectant byproduct formation, which are crucial for public health and regulatory compliance, we demonstrate that the PINN model is able to accurately predict the concentrations of chemical species across various pH values. Notably, the model extends its accuracy to predict concentrations of chemical species not originally included in its training data. The developed method can be extended to a variety of chemical systems, offering a wide array of potential applications.

13 HYDRO ENERGY↗

Modeling Longitudinal Data Containing Non-Normal Within Subject Errors

The mission of the National Aeronautics and Space Administration’s (NASA) human research program is to advance safe human spaceflight. This involves conducting experiments, collecting data, and analyzing data. The data are longitudinal and result from a relatively few number of subjects; typically 10 – 20. A longitudinal study refers to an investigation where participant outcomes and possibly treatments are collected at multiple follow-up times. Standard statistical designs such as mean regression with random effects and mixed–effects regression are inadequate for such data because the population is typically not approximately normally distributed. Hence, more advanced data analysis methods are necessary. This research focuses on four such methods for longitudinal data analysis: the recently proposed linear quantile mixed models (lqmm) by Geraci and Bottai (2013), quantile regression, multilevel mixed–effects linear regression, and robust regression. This research also provides computational algorithms for longitudinal data that scientists can directly use for human spaceflight and other longitudinal data applications, then presents statistical evidence that verifies which method is best for specific situations. This advances the study of longitudinal data in a broad range of applications including applications in the sciences, technology, engineering and mathematics fields.

Feiveson, Alan↗

Technical Report Series on Global Modeling and Data Assimilation, Volume 43: Initial Evaluation of the Climate - MERRA-2

The years since the introduction of MERRA have seen numerous advances in the GEOS-5 Data Assimilation System as well as a substantial decrease in the number of observations that can be assimilated into the MERRA system. To allow continued data processing into the future, and to take advantage of several important innovations that could improve system performance, a decision was made to produce MERRA-2, an updated retrospective analysis of the full modern satellite era. One of the many advances in MERRA-2 is a constraint on the global dry mass balance; this allows the global changes in water by the analysis increment to be near zero, thereby minimizing abrupt global interannual variations due to changes in the observing system. In addition, MERRA-2 includes the assimilation of interactive aerosols into the system, a feature of the Earth system absent from previous reanalyses. Also, in an effort to improve land surface hydrology, observations-corrected precipitation forcing is used instead of model-generated precipitation. Overall, MERRA-2 takes advantage of numerous updates to the global modeling and data assimilation system. In this document, we summarize an initial evaluation of the climate in MERRA-2, from the surface to the stratosphere and from the tropics to the poles. Strengths and weaknesses of the MERRA-2 climate are accordingly emphasized.

MERRA-2↗

Elliptical storm cell modeling of digital radar data

A model for spatial distributions of reflectivity in storm cells was fitted to digital radar data. The data were taken with a modified WSR-57 weather radar with 2.6-km resolution. The data consisted of modified B-scan records on magnetic tape of storm cells tracked at 0 deg elevation for several hours. The MIT L-band radar with 0.8-km resolution produced cross-section data on several cells at 1/2 deg elevation intervals. The model developed uses ellipses for contours of constant effective-reflectivity factor Z with constant orientation and eccentricity within a horizontal cell cross section at a given time and elevation. The centers of the ellipses are assumed to be uniformly spaced on a straight line, with areas linearly related to log Z. All cross sections are similar at different heights (except for cell tops, bottoms, and splitting cells), especially for the highest reflectivities; wind shear causes some translation and rotation between levels. Goodness-of-fit measures and parameters of interest for 204 ellipses are considered.

Altman, F. J.↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization↗

Modeling the data systems role of the scientist (for the NEEDS Command and Control Task)

Research was conducted into the command and control activities of the scientists for five space missions: International Ultraviolet Explorer, Solar Maximum Mission, International Sun-Earth Explorer, High-Energy Astronomy Observatory 1, and Atmospheric Explorer 5. A basis for developing a generalized description of the scientists' activities was obtained. Because of this characteristic, it was decided that a series of flowcharts would be used. This set of flowcharts constitutes a model of the scientists' activities within the total data system. The model was developed through three levels of detail. The first is general and provides a conceptual framework for discussing the system. The second identifies major functions and should provide a fundamental understanding of the scientists' command and control activities. The third level expands the major functions into a more detailed description.

Hei, D. J., Jr.↗

Prediction and Warning of Transported Turbulence in Long-Haul Aircraft Operations

An aviation flight planning system is used for predicting and warning for intersection of flight paths with transported meteorological disturbances, such as transported turbulence and related phenomena. Sensed data and transmitted data provide real time and forecast data related to meteorological conditions. Data modelling transported meteorological disturbances are applied to the received transmitted data and the sensed data to use the data modelling transported meteorological disturbances to correlate the sensed data and received transmitted data. The correlation is used to identify transported meteorological disturbances source characteristics, and identify predicted transported meteorological disturbances trajectories from source to intersection with flight path in space and time. The correlated data are provided to a visualization system that projects coordinates of a point of interest (POI) in a selected point of view (POV) to displays the flight track and the predicted transported meteorological disturbances warnings for the flight crew.

Shipley, Scott T.↗

Dynamic in-context learning with conversational models for data extraction and materials property prediction

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs such as Google gemini-pro and OpenAI gpt-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies—enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall that exceed 95% with an error rate of ∼9%, highlighting the effectiveness and versatility of the toolkit. Finally, databases for 2D material thicknesses, a critical parameter for device integration, and energy bandgap values are developed using PropertyExtractor. In particular, for the thickness database, the rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of various material property databases, advancing the field.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks and challenges associated with deep space missions and experiments (cis-Lunar, Mars transit/surface) require new knowledge discovery and development of novel ecosystems. Supporting distant and long-duration missions and experiments requires biological data (from yeast, microbes, fruit flies, C. elegans, plants, crops, rodents, humans) be findable, accessible, interoperable, reusable (FAIR), and maximally open-access. As data-intensive, bioinformatic, meta-analytical, and computer-assisted approaches continue to be a centerpiece of modern research, the NASA Biological and Physical Sciences division is expanding its Open Science capabilities beyond NASA GeneLab. The NASA Ames Life Sciences Data Archive (ALSDA) is a repository which is responsible for collecting and access to space biological imagery and video, alongside tabular and environmental data. In this presentation, we will discuss strategies dealing with archiving, curating, and accessibility of images from very distinct imaging modalities (e.g., micro-computed tomography, magnetic resonance imaging, photographic images of plants, fluorescence microscopy, behavioral videos, etc.). There are two main challenges: 1. Open-source data storage and 2. Metadata related to the imagery-video. Both have been solved by leveraging two existing open-source systems. For data storage, ALSDA is utilizing components through the Open Microscopy Environment (OME), which can read most imaging proprietary formats and display on a web interface complex multidimensional images (Z stack, multi-channel, temporal, spectral). Most technical metadata from imaging modalities are captured seamlessly. For metadata capturing experimental details, ALSDA (like GeneLab) uses the ISA-Tab specification which relies on the ISA data model to order and classify metadata. The ISA data model uses a tree structure with three files to capture the metadata: The top layer is the Investigations file, the second layer is the Study file(s), and the last layer is the Assay file(s). We believe such an approach may be useful for other types of image research data from other investigators in the AGU community.

imaging↗

Intercomparison of Pulsed Lidar Data with Flight Level CW Lidar Data and Modeled Backscatter from Measured Aerosol Microphysics Near Japan and Hawaii

Aerosol backscatter coefficient data were examined from two nights near Japan and Hawaii undertaken during NASA's Global Backscatter Experiment (GLOBE) in May-June 1990. During each of these two nights the aircraft traversed different altitudes within a region of the atmosphere defined by the same set of latitude and longitude coordinates. This provided an ideal opportunity to allow flight level focused continuous wave (CW) lidar backscatter measured at 9.11-micron wavelength and modeled aerosol backscatter from two aerosol optical counters to be compared with pulsed lidar aerosol backscatter data at 1.06- and 9.25-micron wavelengths. The best agreement between all sensors was found in the altitude region below 7 km, where backscatter values were moderately high at all three wavelengths. Above this altitude the pulsed lidar backscatter data at 1.06- and 9.25-micron wavelengths were higher than the flight level data obtained from the CW lidar or derived from the optical counters, suggesting sample volume effects were responsible for this. Aerosol microphysics analysis of data near Japan revealed a strong sea-salt aerosol plume extending upward from the marine boundary layer. On the basis of sample volume differences, it was found that large particles were of different composition compared with the small particles for low backscatter conditions.

Cutten, D. R.↗

Uniform Data Access Using GXD

This paper gives an overview of GXD, a framework facilitating publication and use of data from diverse data sources. GXD defines an object-oriented data model designed to represent a wide range of things including data, its metadata, resources and query results. GXD also defines a data transport language. a dialect of XML, for representing instances of the data model. This language allows for a wide range of data source implementations by supporting both the direct incorporation of data and the specification of data by various rules. The GXD software library, proto-typed in Java, includes client and server runtimes. The server runtime facilitates the generation of entities containing data encoded in the GXD transport language. The GXD client runtime interprets these entities (potentially from many data sources) to create an illusion of a globally interconnected data space, one that is independent of data source location and implementation.

Vanderbilt, Peter↗

Consequences of Different Air-Sea Feedbacks on Ocean Using MITgcm and MERRA-2 Forcing: Implications for Coupled Data Assimilation Systems

Ocean surface flux estimates from atmospheric and oceanic reanalyses contain errors that compensate for inaccuracies in the respective atmosphere and ocean models used to generate these reanalyses. A conundrum for climate studies is the discrepancy between surface fluxes that minimize model-data differences for an atmosphere-only model vs surface fluxes that minimize model-data differences for an ocean model. As a first step towards a consistent coupled ocean-atmosphere data-assimilation (DA) system, we compare surface net heat flux from a state-of-the-art atmospheric reanalysis, the Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2), to net heat flux from a state-of-the-art ocean state estimate, the Estimating the Circulation and Climate of the Ocean Version 4 (ECCO-v4). The possible impacts of the MERRA-2 and ECCO-v4 air-sea net heat flux difference in a coupled DA system were assessed using a set of experiments designed to imitate different “flavors” of a coupled DA system in an ocean-only setup. This was done by forcing the ECCO-v4 underlying ocean model - the Massachusetts Institute of Technology general circulation model (MITgcm) - with different sets of MERRA-2 fields and utilizing different forcing methods. By doing so we were able to turn off different air-sea feedbacks which, in a coupled DA setup, are partially muted by the constraining observations. The set of experiments, therefore, represents a range of active feedbacks in different “flavors” of coupled data-assimilation systems. For the period 1992–2011, MERRA-2 net heat flux has a global mean difference of -4.9 Wm(exp -2) relative to ECCO-v4. When MERRA-2 surface fields are used to force MITgcm, imbalances in the energy and the hydrological cycles of MERRA-2, which are directly related to the fact that MERRA-2 was created without an interactive ocean, propagate to the ocean. The experiment in which MITgcm is forced with MERRA-2 fluxes (MERRA-2-flux experiment) results in a 2.5°C global mean Sea Surface Temperature (SST) cooling, a 1m reduction in global mean sea level, and other drastic changes in the large scale ocean circulation relative to those resulting when the MITgcm is forced with the optimized ECCO-v4 net heat flux (the ECCO-v4 experiment itself). When MITgcm is forced with MERRA-2 state variables (MERRA-2-state experiment), the SST is somewhat restored to the observed SST, but the errors are shifted to the water cycle, resulting in a global mean sea level increase of 2.7 m. To further explore the pros and cons of these two approaches, we introduce a new intermediate forcing method in which the ocean is forced with turbulent fluxes but has a long wave feedback. This method, unlike MERRA-2 state, preserves the MERRA-2 water and salinity cycles, and it reduces the SST error compared to the MERRA-2-flux experiment, but the SST is not as good as that in the MERRA-2-state experiment. Our results have implications for ocean-model forcing recipes and clearly reveal the undesirable consequences of limiting the feedbacks in either these types of experiments or in coupled DA.

Ehud Strobach↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Explosive east coast cyclogenesis - Numerical experimentation and model-based diagnostics

Numerical experimentation of explosive east-coast cyclogenesis is performed using the Florida State University Global Spectral Model (FSUGSM). The three cases examined here are the Presidents' Day storm of February 18-19, 1979 and the North Atlantic and Pacific bombs of January 18-20, 1979 which formed off the east coasts of the United States and Japan, respectively. The use of a global model provides a framework for studying the phenomena on the 3-5 day time scale. The forecast verifications of the numerical experiments indicate that the FSUGSM was able to adequately predict the phase, intensity, and synoptic-scale structure. These results justify the use of model data for diagnostic studies of the bomb. The model data are used to quantify the role of the adiabatic and diabatic forcing in the explosive cyclogenetic process, using surface pressure tendency to gage development.

Manobianco, John↗

Overview of NASA MSFC IEC Federated Engineering Collaboration Capability

The MSFC IEC federated engineering framework is currently developing a single collaborative engineering framework across independent NASA centers. The federated approach allows NASA centers the ability to maintain diversity and uniqueness, while providing interoperability. These systems are integrated together in a federated framework without compromising individual center capabilities. MSFC IEC's Federation Framework will have a direct affect on how engineering data is managed across the Agency. The approach is directly attributed in response to the Columbia Accident Investigation Board (CAB) finding F7.4-11 which states the Space Shuttle Program has a wealth of data sucked away in multiple databases without a convenient way to integrate and use the data for management, engineering, or safety decisions. IEC s federated capability is further supported by OneNASA recommendation 6 that identifies the need to enhance cross-Agency collaboration by putting in place common engineering and collaborative tools and databases, processes, and knowledge-sharing structures. MSFC's IEC Federated Framework is loosely connected to other engineering applications that can provide users with the integration needed to achieve an Agency view of the entire product definition and development process, while allowing work to be distributed across NASA Centers and contractors. The IEC DDMS federation framework eliminates the need to develop a single, enterprise-wide data model, where the goal of having a common data model shared between NASA centers and contractors is very difficult to achieve.

Moushon, Brian↗

Confronting Models with Data: The GEWEX Cloud Systems Study

The GEWEX Cloud System Study (GCSS; GEWEX is the Global Energy and Water Cycle Experiment) was organized to promote development of improved parameterizations of cloud systems for use in climate and numerical weather prediction models, with an emphasis on the climate applications. The strategy of GCSS is to use two distinct kinds of models to analyze and understand observations of the behavior of several different types of clouds systems. Cloud-system-resolving models (CSRMs) have high enough spatial and temporal resolutions to represent individual cloud elements, but cover a wide enough range of space and time scales to permit statistical analysis of simulated cloud systems. Results from CSRMs are compared with detailed observations, representing specific cases based on field experiments, and also with statistical composites obtained from satellite and meteorological analyses. Single-column models (SCMs) are the surgically extracted column physics of atmospheric general circulation models. SCMs are used to test cloud parameterizations in an un-coupled mode, by comparison with field data and statistical composites. In the original GCSS strategy, data is collected in various field programs and provided to the CSRM Community, which uses the data to "certify" the CSRMs as reliable tools for the simulation of particular cloud regimes, and then uses the CSRMs to develop parameterizations, which are provided to the GCM Community. We report here the results of a re-thinking of the scientific strategy of GCSS, which takes into account the practical issues that arise in confronting models with data. The main elements of the proposed new strategy are a more active role for the large-scale modeling community, and an explicit recognition of the importance of data integration.

Randall, David↗

Dynamics of the Antarctic Circumpolar Current. Evidence for Topographic Effects from Altimeter Data and Numerical Model Output

Geosat altimeter data and numerical model output are used to examine the circulation and dynamics of the Antarctic Circumpolar Current (ACC). The mean sea surface height across the ACC has been reconstructed from height variability measured by the altimeter, without assuming prior knowledge of the geoid. The results indicate locations for the Subantarctic and Polar Fronts which are consistent with in situ observations and indicate that the fronts are substantially steered by bathymetry. Detailed examination of spatial and temporal variability indicates a spatial decorrelation scale of 85 km and a temporal e-folding scale of 34 days. Empirical Orthogonal Function analysis suggests that the scales of motion are relatively short, occuring on 1000 km length-scales rather than basin or global scales. The momentum balance of the ACC has been investigated using output from the high resolution primitive equation model in combination with altimeter data. In the Semtner-Chervin quarter-degree general circulation model topographic form stress is the dominant process balancing the surface wind forcing. In stream coordinates, the dominant effect transporting momentum across the ACC is bibarmonic friction. Potential vorticity is considered on Montgomery streamlines in the model output and along surface streamlines in model and altimeter data. (AN)

Gille, Sarah T.↗

A cloud cover model based on satellite data

A model for worldwide cloud cover using a satellite data set containing infrared radiation measurements is proposed. The satellite data set containing day IR, night IR and incoming and absorbed solar radiation measurements on a 2.5 degree latitude-longitude grid covering a 45 month period was converted to estimates of cloud cover. The global area was then classified into homogeneous cloud cover regions for each of the four seasons. It is noted that the developed maps can be of use to the practicing climatologist who can obtain a considerable amount of cloud cover information without recourse to large volumes of data.

Somerville, P. N.↗