Search NASA⌕ Search

SEARCH · Search NASA

Results for “Common data models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Physics-informed machine learning for building performance simulation-A review of a nascent field

Building performance simulation (BPS) is critical for understanding building dynamics and behavior, analyzing the performance of the built environment, optimizing energy efficiency, improving demand flexibility, and enhancing building resilience. However, conducting BPS is not trivial. Traditional BPS relies on accurate building energy models, which are primarily physics-based and heavily dependent on detailed building information, expert knowledge, and case-by-case model calibrations, significantly limiting their scalability. With the development of sensing technology and the increased availability of data, there is growing attention and interest in data-driven BPS. However, purely data-driven models often suffer from limited generalization ability and a lack of physical consistency, resulting in poor performance in real-world applications. To address these limitations, recent studies have begun integrating physics priors into data-driven models, a methodology known as physics-informed machine learning (PIML). PIML is an emerging field where its definitions, methodologies, evaluation criteria, application scenarios, and future directions remain open. To bridge those gaps, this study systematically reviews the state-of-the-art PIML for BPS, offering a comprehensive definition of PIML and comparing it to traditional BPS approaches regarding data requirements, modeling effort, performance, and computational cost. We also summarize the commonly used methodologies, validation approaches, application domains, available data sources, open-source packages, and testbeds. In addition, this study provides a general guideline for selecting appropriate PIML models based on BPS applications. Finally, this study identifies key challenges and outlines future research directions, providing a solid foundation and valuable insights to advance R&D of PIML in BPS.

Jiang, Zixin↗

Spatial Predictive Modeling and Remote Sensing of Land Use Change in the Chesapeake Bay Watershed

This project was focused on modeling the processes by which increasing demand for developed land uses, brought about by changes in the regional economy and the socio-demographics of the region, are translated into a changing spatial pattern of land use. Our study focused on a portion of the Chesapeake Bay Watershed where the spatial patterns of sprawl represent a set of conditions generally prevalent in much of the U.S. Working in the region permitted us access to (i) a time-series of multi-scale and multi-temporal (including historical) satellite imagery and (ii) an established network of collaborating partners and agencies willing to share resources and to utilize developed techniques and model results. In addition, a unique parcel-level tax assessment database and linked parcel boundary maps exists for two counties in the Maryland portion of this region that made it possible to establish a historical cross-section time-series database of parcel level development decisions. Scenario analyses of future land use dynamics provided critical quantitative insight into the impact of alternative land management and policy decisions. These also have been specifically aimed at addressing growth control policies aimed at curbing exurban (sprawl) development. Our initial technical approach included three components: (i) spatial econometric modeling of the development decision, (ii) remote sensing of suburban change and residential land use density, including comparisons of past change from Landsat analyses and more traditional sources, and (iii) linkages between the two through variable initialization and supplementation of parcel level data. To these we added a fourth component, (iv) cellular automata modeling of urbanization, which proved to be a valuable addition to the project. This project has generated both remote sensing and spatially explicit socio-economic data to estimate and calibrate the parameters for two different types of land use change models and has undertaken analyses of these models. One (the CA model) is driven largely by observations on past patterns of land use change, while the other (the EC model) is driven by mechanisms of the land use change decision at the parcel level. Our project may be the first serious attempt at developing both types of models for the same area, using as much common data as possible. We have identified the strengths and weaknesses of the two approaches and plan to continue to revise each model in the light of new data and new lessons learned through continued collaboration. Questions, approaches, findings, publication and presentation lists concerning the research are also presented.

Goetz, Scott J.↗

Harmonized Emissions Component (HEMCO) 3.0 as a Versatile Emissions Component for Atmospheric Models: Application in the GEOS-Chem, NASA GEOS, WRF-GC, CESM2, NOAA GEFS-Aerosol, and NOAA UFS Models

Emissions are a central component of atmospheric chemistry models. The Harmonized Emissions Component (HEMCO) is a software component for computing emissions from a user-selected ensemble of emission inventories and algorithms. It allows users to re-grid, combine, overwrite, subset, and scale emissions from different inventories through a configuration file and with no change to the model source code. The configuration file also maps emissions to model species with appropriate units. HEMCO can operate in offline stand-alone mode, but more importantly it provides an online facility for models to compute emissions at runtime. HEMCO complies with the Earth System Modeling Framework (ESMF) for portability across models. We present a new version here, HEMCO 3.0, that features an improved three-layer architecture to facilitate implementation into any atmospheric model and improved capability for calculating emissions at any model resolution including multiscale and unstructured grids. The three-layer architecture of HEMCO 3.0 includes (1) the Data Input Layer that reads the configuration file and accesses the HEMCO library of emission inventories and other environmental data, (2) the HEMCO Core that computes emissions on the user-selected HEMCO grid, and (3) the Model Interface Layer that re-grids (if needed) and serves the data to the atmospheric model and also serves model data to the HEMCO Core for computing emissions dependent on model state (such as from dust or vegetation). The HEMCO Core is common to the implementation in all models, while the Data Input Layer and the Model Interface Layer are adaptable to the model environment. Default versions of the Data Input Layer and Model Interface Layer enable straightforward implementation of HEMCO in any simple model architecture, and options are available to disable features such as re-gridding that may be done by independent couplers in more complex architectures. The HEMCO library of emission inventories and algorithms is continuously enriched through user contributions so that new inventories can be immediately shared across models. HEMCO can also serve as a general data broker for models to process input data not only for emissions but for any gridded environmental datasets. We describe existing implementations of HEMCO 3.0 in (1) the GEOS-Chem “Classic” chemical transport model with shared-memory infrastructure, (2) the high-performance GEOS-Chem (GCHP) model with distributed-memory architecture, (3) the NASA GEOS Earth System Model (GEOS ESM), (4) the Weather Research and Forecasting model with GEOS-Chem (WRF-GC), (5) the Community Earth System Model Version 2 (CESM2), and (6) the NOAA Global Ensemble Forecast System – Aerosols (GEFS-Aerosols), as well as the planned implementation in the NOAA Unified Forecast System (UFS). Implementation of HEMCO in CESM2 contributes to the Multi-Scale Infrastructure for Chemistry and Aerosols (MUSICA) by providing a common emissions infrastructure to support different simulations of atmospheric chemistry across scales.

Haipeng Lin↗

A Rapid Approach to Modeling Species-Habitat Relationships

A growing number of species require conservation or management efforts. Success of these activities requires knowledge of the species' occurrence pattern. Species-habitat models developed from GIS data sources are commonly used to predict species occurrence but commonly used data sources are often developed for purposes other than predicting species occurrence and are of inappropriate scale and the techniques used to extract predictor variables are often time consuming and cannot be repeated easily and thus cannot efficiently reflect changing conditions. We used digital orthophotographs and a grid cell classification scheme to develop an efficient technique to extract predictor variables. We combined our classification scheme with a priori hypothesis development using expert knowledge and a previously published habitat suitability index and used an objective model selection procedure to choose candidate models. We were able to classify a large area (57,000 ha) in a fraction of the time that would be required to map vegetation and were able to test models at varying scales using a windowing process. Interpretation of the selected models confirmed existing knowledge of factors important to Florida scrub-jay habitat occupancy. The potential uses and advantages of using a grid cell classification scheme in conjunction with expert knowledge or an habitat suitability index (HSI) and an objective model selection procedure are discussed.

Carter, Geoffrey M.↗

Maven: a multimodal foundation model for supernova science

Abstract A common setting in astronomy is the availability of a small number of high-quality observations, and larger amounts of either lower-quality observations or synthetic data from simplified models. Time-domain astrophysics is a canonical example of this imbalance, with the number of supernovae observed photometrically outpacing the number observed spectroscopically by multiple orders of magnitude. At the same time, no data-driven models exist to understand these photometric and spectroscopic observables in a common context. Contrastive learning objectives, which have grown in popularity for aligning distinct data modalities in a shared embedding space, provide a potential solution to extract information from these modalities. We present Maven, the first foundation model for supernova science. To construct Maven, we first pre-train our model to align photometry and spectroscopy from 0.5 M synthetic supernovae using a contrastive objective. We then fine-tune the model on 4702 observed supernovae from the Zwicky transient facility. Maven reaches state-of-the-art performance on both classification and redshift estimation, despite the embeddings not being explicitly optimized for these tasks. Through ablation studies, we show that pre-training with synthetic data improves overall performance. In the upcoming era of the Vera C. Rubin observatory, Maven will serve as a valuable tool for leveraging large, unlabeled and multimodal time-domain datasets.

Zhang, Gemma (ORCID:0000000280198082)↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Data Format Standardization of Space Weather Model Output at the Community Coordinated Modeling Center

The disparate nature of space weather model output provides many challenges with regards to the portability and reuse of not only the data itself, but also any tools that are developed for analysis and visualization. We are developing and implementing a comprehensive data format standardization methodology that allows heterogeneous model output data to be stored uniformly in any common science data format. We will discuss our approach to identifying core meta-data elements that can be used to supplement raw model output data, thus creating self-descriptive files. The meta-data should also contain information describing the simulation grid. This will ultimately assists in the development of efficient data access tools capable of extracting data at any given point and time. We will also discuss our experiences standardizing the output of two global magnetospheric models, and how we plan to apply similar procedures when standardizing the output of the solar, heliospheric, and ionospheric models that are also currently hosted at the Community Coordinated Modeling Center.

Maddox, M.↗

The Land Surface Data Toolkit (LDT v7.2) - A Data Fusion Environment for Land Data Assimilation Systems

The effective applications of land surface models (LSMs) and hydrologic models pose a varied set of data input and processing needs, ranging from ensuring consistency checks to more derived data processing and analytics. This article describes the development of the Land surface Data Toolkit (LDT), which is an integrated framework designed specifically for processing input data to execute LSMs and hydrological models. LDT not only serves as a preprocessor to the NASA Land Information System (LIS), which is an integrated framework designed for multi-model LSM simulations and data assimilation (DA) integrations, but also as a land-surface-based observation and DA input processor. It offers a variety of user options and inputs to processing datasets for use within LIS and stand-alone models. The LDT design facilitates the use of common data formats and conventions. LDT is also capable of processing LSM initial conditions and meteorological boundary conditions and ensuring data quality for inputs to LSMs and DA routines. The machine learning layer in LDT facilitates the use of modern data science algorithms for developing data-driven predictive models. Through the use of an object-oriented framework design, LDT provides extensible features for the continued development of support for different types of observational datasets and data analytics algorithms to aid land surface modeling and data assimilation.

droughts and floods↗

Creation of lumped parameter thermal model by the use of finite elements

In the finite difference technique, the thermal network is represented by an analogous electrical network. The development of this network model, which is used to describe a physical system, often requires tedious and mental data preparation and checkout by the analyst which can be greatly reduced through the use of the computer programs to develop automatically the mathematical model and associated input data and graphically display the analytical model to facilitate model verification. Three separate programs are involved which are linked through common mass storage files and data card formats. These programs are SPAR, CINGEN and GEOMPLT, and are used to (1) develop thermal models for the MITAS II thermal analyzer program; (2) produce geometry plots of the thermal network; and (3) produce temperature distribution and time history plots.

Source record↗

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)↗

The phase of the crosspolarized signal generated by millimeter wave propagation through rain

Proposed schemes for cancelling rain-induced crosstalk in dual-polarized communications systems depend upon the phase relationships between the wanted and unwanted signals. This report investigates the phase relationship of the rain-generated crosspolarized signal relative to the copolarized signal. Theoretical results obtained from a commonly accepted propagation model are presented. Experimental data from the Communications Technology Satellite beacon and from the Comstar beacon are presented and the correlation between theory and data is discussed. An inexpensive semi-adaptive cancellation system is proposed and its performance expectations are presented. The implications of phase variations on a cancellation system are also discussed.

Overstreet, W. P.↗

The dust around R Coronae Borealis type stars

Measurements taken by the International Ultraviolet Explorer spacecraft of the stars RY Sgr and R CrB have been analyzed using Mie theory. The extinction data, which show a 2400-2500 A peak, are consistent with a distribution of 5-60 nm glassy or amorphous carbon particles obscuring the stellar flux. The data are also fairly consistent with a cloud ejection model. Since the extinction data lack the commonly observed peak at 2170 A, it is proposed that this difference is due to the conditions present when the dust condenses. Interstellar carbon grains appear to originate in normal carbon stars which are carbon and hydrogen rich. In contrast, the grains around R CrB type stars seem to condense from a carbon-rich and hydrogen-poor vapor.

Hecht, J. H.↗

Atmospheric Tides Middle Atmosphere Program (ATMAP): Report of the November/december 1981, and May 1982, Observational Campaigns

Atmospheric tides, oscillations in meteorological fields occurring at subharmonics of a solar or lunar day, comprise a major component of middle atmosphere global dynamics. The nature of atmospheric tides requires investigations and coordination on a global, and hence international, scale. The purpose of ATMAP is to create an interaction among observationalists, data analysts, theoreticians and modellers working towards common goals. The focus for this interaction is a series of global observational campaigns involving ground based remote sensing methods. The calendar of activities for ATMAP is presented. The results of solstice campaigns I and II are assembled, a preliminary interpretation of the results is presented.

Jeffrey M Forbes↗

A combined optical/X-ray study of the Galaxy cluster Abell 2256

The dynamics of Abell 2256 is investigated by combining X-ray observations of the intracluster gas with optical observations of the galaxy distribution and kinematics. Magnitudes and positions are presented for 172 galaxies and new redshifts for 75. Abell 2256 is similar to the Coma Cluster in its X-ray luminosity, mass, and galaxy density. Both the X-ray surface brightness and the galaxy surface density distributions exhibit an elliptical morphology. The radial galaxy distribution is steeper than the density profile of the X-ray-emitting gas, yet the galaxy velocity dispersion is higher than the equivalent value for the gas. Under the simplest assumptions that the galaxy velocity distribution is isotropic and the gas is isothermal, the galaxies and gas cannot be in hydrostatic equilibrium in a common gravitational potential. Models consistent with available data have mass-to-light ratios which increase with radius and galaxy orbits that are anisotropic with a radial bias.

Fabricant, Daniel G.↗

Interoperability of Heliophysics Virtual Observatories

If you'd like to find interrelated heliophysics (also known as space and solar physics) data for a research project that spans, for example, magnetic field data and charged particle data from multiple satellites located near a given place and at approximately the same time, how easy is this to do? There are probably hundreds of data sets scattered in archives around the world that might be relevant. Is there an optimal way to search these archives and find what you want? There are a number of virtual observatories (VOs) now in existence that maintain knowledge of the data available in subdisciplines of heliophysics. The data may be widely scattered among various data centers, but the VOs have knowledge of what is available and how to get to it. The problem is that research projects might require data from a number of subdisciplines. Is there a way to search multiple VOs at once and obtain what is needed quickly? To do this requires a common way of describing the data such that a search using a common term will find all data that relate to the common term. This common language is contained within a data model developed for all of heliophysics and known as the SPASE (Space Physics Archive Search and Extract) Data Model. NASA has funded the main part of the development of SPASE but other groups have put resources into it as well. How well is this working? We will review the use of SPASE and how well the goal of locating and retrieving data within the heliophysics community is being achieved. Can the VOs truly be made interoperable despite being developed by so many diverse groups?

Thieman, J.↗

Assessment of Aeroacoustic Simulations of the High-Lift Common Research Model

This paper presents further validation of PowerFLOWR aeroacoustic simulations of the High-Lift Common Research Model through comparisons with experimental data from a recently completed wind tunnel test. Preliminary time- averaged surface pressure and microphone array data from the experiment are in reasonably good agreement with the simulations, and the slat is shown to be a dominant noise source on this model. The simulations did not predict slat tones that were very prominent in the experiment, but they did capture the broadband component of slat noise in the low-frequency range up to 1 kHz at full scale. Future tests are planned to demonstrate slat noise reduction technology, and simulations are being used to guide this development.

Lockard, David P.↗

Geophysical Observations Toolkit For Evaluating Coral Health (GOTECH) Fall 2021 Final Report

The NASA Langley Research Center (LaRC) Data Science Team (DST), under the Office of the Chief Information Officer (OCIO), is investigating the capacity of the Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observation (CALIPSO) satellite to infer the vitality of coral reefs. This report describes the Fall 2021 period of performance for the Geophysical Observations Toolkit for Evaluating Coral Health (GOTECH) project. During this effort, two student teams at Georgia Tech developed machine-learning models to predict the vitality of coral reefs in targeted geographic regions based on backscatter data from the CALIPSO satellite. To train these models, students fused data to form a common operating picture of how coral reefs have grown and decayed worldwide. This report describes the student assignment, background, and results of the semester's research.

Machine Learning↗

Evaluation of Aerosol Data Assimilation and Forecasts in the NASA GEOS Model during the ASIA-AQ Campaign

Fine particulate matter (PM2.5) poses significant risks to human health and the environment by penetrating the lungs and causing respiratory and cardiovascular diseases, making it crucial to understand its sources and behavior for effective air quality management. The Goddard Earth Observing System (GEOS) Forward Processing (FP) system model, operated by the Global Modeling and Assimilation Office (GMAO) at NASA's Goddard Space Flight Center, provides real-time weather and aerosol analyses and forecasts. In addition to meteorological data assimilation, the GEOS-FP system also assimilates aerosol using Moderate Resolution Imaging Spectroradiometer (MODIS) Aerosol Optical Depth (AOD) and Aerosol Robotic Network (AERONET) AOD data. In this study, the aerosol data assimilation and forecasts performance of the GEOS-FP model were evaluated for predicting PM2.5 in Korea using observations from the Airborne and Satellite Investigation of Asian Air Quality (ASIA-AQ) campaign. The ASIA-AQ campaign, an international collaborative field study initiative, aims to enhance understanding of local air quality issues and address common challenges in interpreting satellite data and air quality modeling. Conducted in South Korea from February 15 to March 13, 2024, during the high PM2.5 concentration winter season, this campaign provided extensive airborne and ground observations for intensive analysis of PM2.5 model simulations. We demonstrate how the assimilation runs and the forecasting performance of PM2.5 at 24-hour and 48-hour intervals vary. Additionally, we analyzed the differences and characteristics of PM2.5 composition in cases of long-range transport and local emissions. Using ASIA-AQ airborne data, we also examined the vertical profile of fine particulate matter. Through the intensive observations of this campaign, the GEOS model was assessed over South Korea using both in situ and airborne measurements to establish a baseline and identify priorities for future development.

Seunghee Lee↗