Search NASASearch

SEARCH · Search NASA

Results for “Data modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Reduction of Marine Magnetic Data for Modeling the Main Field of the Earth

The marine data set archived at the National Geophysical Data Center (NGDC) consists of shipborne surveys conducted by various institutes worldwide. This data set spans four decades (1953, 1958, 1960-1987), and contains almost 13 million total intensity observations. These are often less than 1 km apart. These typically measure seafloor spreading anomalies with amplitudes of several hundred nanotesla (nT) which, since they originate in the crust, interfere with main field modeling. The source for these short wavelength features are confined within the magnetic crust (i.e., sources above the Curie isotherm). The main field, on the other hand, is of much longer wavelengths and originates within the earth's core. It is desirable to extract the long wavelength information from the marine data set for use in modeling the main field. This can be accomplished by averaging the data along the track. In addition, those data which are measured during periods of magnetic disturbance can be identified and eliminated. Thus, it should be possible to create a data set which has worldwide data distribution, spans several decades, is not contaminated with short wavelengths of the crustal field or with magnetic storm noise, and which is limited enough in size to be manageable for the main field modeling. The along track filtering described above has proved to be an effective means of condensing large numbers of shipborne magnetic data into a manageable and meaningful data set for main field modeling. Its simplicity and ability to adequately handle varying spatial and sampling constraints has outweighed consideration of more sophisticated approaches. This filtering technique also provides the benefits of smoothing out short wavelength crustal anomalies, discarding data recorded during magnetically noisy periods, and assigning reasonable error estimates to be used in the least square modeling. A useful data set now exists which spans 1953-1987.

Baldwin, R. T.

Collaborative Research: Enabling multi-scale studies of magnetic reconnection with interpretable data-driven models

The development of accurate reduced descriptions and improved closures for magnetic reconnection is an important and a long‐standing challenge in plasma physics. The four‐fluid approach, and associated closures, that were investigated have the potential to improve the accuracy of plasma fluid models, capturing physical effects which would otherwise require a kinetic description. If successful, this approach could have an important impact for the modeling of laboratory and space plasmas. The major goals of this project were to develop new machine learning (ML) tools based on sparse and symbolic regression techniques, and to extract interpretable and generalizable reduced models (e.g., in the form of partial differential equations - PDEs) from data generated by first principles plasma simulations. Preserving interpretability of such data‐driven models is key to addressing the long‐standing theoretical and numerical challenges. Prior proof‐of‐principle studies have demonstrated the enormous potential of this approach, by recovering the well‐established hierarchy of plasma equations (from Vlasov to MHD) from data produced by particle‐in‐cell (PIC) simulations. Our goal in this project was to extend and apply these new tools to construct better kinetic closures for magnetic reconnection; to derive better models of particle injection and acceleration by this fundamental plasma process; and to use this understanding to accelerate the development of multi‐scale plasma algorithms. While our immediate focus was on the problem of magnetic reconnection, the tools that were will developed are general and applicable to other areas of plasma physics, and more broadly to many‐body phenomena. We anticipate that the development of these multi‐scale models will have a significant impact across different areas of plasma science, from fusion to space and astrophysical plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Learning Nonlinear Reduced Models from Data with Operator Inference

This review discusses Operator Inference, a nonintrusive reduced modeling approach that incorporates physical governing equations by defining a structured polynomial form for the reduced model, and then learns the corresponding reduced operators from simulated training data. The polynomial model form of Operator Inference is sufficiently expressive to cover a wide range of nonlinear dynamics found in fluid mechanics and other fields of science and engineering, while still providing efficient reduced model computations. The learning steps of Operator Inference are rooted in classical projection-based model reduction; thus, some of the rich theory of model reduction can be applied to models learned with Operator Inference. This connection to projection-based model reduction theory offers a pathway toward deriving error estimates and gaining insights to improve predictions. Furthermore, through formulations of Operator Inference that preserve Hamiltonian and other structures, important physical properties such as energy conservation can be guaranteed in the predictions of the reduced model beyond the training horizon. This review illustrates key computational steps of Operator Inference through a large-scale combustion example.

Mechanics

Long-term archiving and data access: modelling and standardization

This paper reports on the multiple difficulties inherent in the long-term archiving of digital data, and in particular on the different possible causes of definitive data loss. It defines the basic principles which must be respected when creating long-term archives. Such principles concern both the archival systems and the data. The archival systems should have two primary qualities: independence of architecture with respect to technological evolution, and generic-ness, i.e., the capability of ensuring identical service for heterogeneous data. These characteristics are implicit in the Reference Model for Archival Services, currently being designed within an ISO-CCSDS framework. A system prototype has been developed at the French Space Agency (CNES) in conformance with these principles, and its main characteristics will be discussed in this paper. Moreover, the data archived should be capable of abstract representation regardless of the technology used, and should, to the extent that it is possible, be organized, structured and described with the help of existing standards. The immediate advantage of standardization is illustrated by several concrete examples. Both the positive facets and the limitations of this approach are analyzed. The advantages of developing an object-oriented data model within this contxt are then examined.

Hoc, Claude

Comparing Ice Jam Hindcasting Models with Tree Scar Data

Hindcasting models can use historic ice jam observations and hydroclimatic data to identify conditions that form ice jams.However, historic ice jam records are often sparse or incomplete. New sources of historic ice jam data could improve hindcasting models, leading to better ice jam forecasting and flood warning systems. Because ice jams damage riparian trees, marker rings associated with historic scars include information about ice jam frequency and severity. This study examined marker rings from 56 trees along the Muskegon River to supplement the historic ice jam data on this system. The study team compared tree ring data to results from hindcasting models, which were independently validated with newspaper reports on 1,500 separate days. Logistic regression converted the marker ring data into annual ice jam probabilities. Ice jam dates from the dendrochronology data were too noisy to train or validate a hindcasting model. However, the marker ring data did confirm that ice jams on the Muskegon are nonstationary. Ice jams are significantly more likely now than they were before 1966.The marker rings also helped to distinguish between false-negatives and nondetects in the hindcasting model.

Stanford Gibson

Aeroservoelastic Uncertainty Model Identification from Flight Data

Uncertainty modeling is a critical element in the estimation of robust stability margins for stability boundary prediction and robust flight control system development. There has been a serious deficiency to date in aeroservoelastic data analysis with attention to uncertainty modeling. Uncertainty can be estimated from flight data using both parametric and nonparametric identification techniques. The model validation problem addressed in this paper is to identify aeroservoelastic models with associated uncertainty structures from a limited amount of controlled excitation inputs over an extensive flight envelope. The challenge to this problem is to update analytical models from flight data estimates while also deriving non-conservative uncertainty descriptions consistent with the flight data. Multisine control surface command inputs and control system feedbacks are used as signals in a wavelet-based modal parameter estimation procedure for model updates. Transfer function estimates are incorporated in a robust minimax estimation scheme to get input-output parameters and error bounds consistent with the data and model structure. Uncertainty estimates derived from the data in this manner provide an appropriate and relevant representation for model development and robust stability analysis. This model-plus-uncertainty identification procedure is applied to aeroservoelastic flight data from the NASA Dryden Flight Research Center F-18 Systems Research Aircraft.

Brenner, Martin J.

Using data with models - Ill-posed problems

An attempt is made to show that the combination of data with models leads to a sequence of conventionally ill-posed problems which can be treated in a systematic and practical way by control-theory techniques. The illustrative approach is to exploit the very great mathematical simplifications which result when the methods of finite dimensional vector spaces are used, thus reducing a lot of difficult mathematics to classical least squares. The immediate motivation for this study comes from ocean acoustic tomography.

Wunsch, C.

Technical report series on global modeling and data assimilation. Volume 1: Documentation of the Goddard Earth Observing System (GEOS) General Circulation Model, version 1

This technical report documents Version 1 of the Goddard Earth Observing System (GEOS) General Circulation Model (GCM). The GEOS-1 GCM is being used by NASA's Data Assimilation Office (DAO) to produce multiyear data sets for climate research. This report provides a documentation of the model components used in the GEOS-1 GCM, a complete description of model diagnostics available, and a User's Guide to facilitate GEOS-1 GCM experiments.

Suarez, Max J.

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling

Super high compression of line drawing data

Models which can be used to accurately represent the type of line drawings which occur in teleconferencing and transmission for remote classrooms and which permit considerable data compression were described. The objective was to encode these pictures in binary sequences of shortest length but such that the pictures can be reconstructed without loss of important structure. It was shown that exploitation of reasonably simple structure permits compressions in the range of 30-100 to 1. When dealing with highly stylized material such as electronic or logic circuit schematics, it is unnecessary to reproduce configurations exactly. Rather, the symbols and configurations must be understood and be reproduced, but one can use fixed font symbols for resistors, diodes, capacitors, etc. Compression of pictures of natural phenomena such as can be realized by taking a similar approach, or essentially zero error reproducibility can be achieved but at a lower level of compression.

Cooper, D. B.

A Robust Schema for Storing and Managing Machine Learning Data and Models

- Machine Learning (ML) has enabled models that can improve efficiency and decrease computational cost - ML models are crucial in enabling Integrated Computational Materials Engineering (ICME) - Large data sets require robust means of storing ML data and models

Brandon L. Hearley

Development of Data-Driven Models for Performance Prediction and Chemical Dosing of a Full-Scale Controlled Phosphorus Precipitation Reactor

This study evaluated the use of data-driven models to improve control of a struvite precipitation reactor that removes phosphorus from wastewater while producing a fertilizer product. The researchers developed predictive models for influent orthophosphate concentration, effluent orthophosphate concentration, and phosphorus removal using operational data from a full-scale MagPrex™ reactor at a water resource recovery facility in Denver, Colorado. Model predictions were used to recommend magnesium chloride dosing adjustments needed to achieve a target effluent phosphorus concentration. Several machine learning approaches were tested, with ridge regression providing the best predictions for influent orthophosphate concentration and phosphorus removal, and XGBoost providing the best predictions for effluent orthophosphate concentration. Simulation results indicated that the decision-support approach could correctly identify dosing adjustments in most cases and reduce chemical use. Full-scale implementation achieved lower accuracy due to changing operating conditions and limited historical data in some operating ranges. Here, the results demonstrate the potential of data-driven tools to support phosphorus recovery process control while also identifying practical limitations that affect deployment in full-scale systems.

42 ENGINEERING

A Comparison of the Age-Spectra from Data Assimilation Models

We use kinematic and diabatic back trajectory calculations, driven by winds from a general circulation model (GCM) and two different data assimilation systems (DAS), to compute the age spectrum at three latitudes in the lower stratosphere. The age-spectra are compared to chemical transport model (CTM) calculations, and the mean ages from all of these studies are compared to observations. The age spectra computed using the GCM winds show a reasonably well-isolated tropics in good agreement with observations; however, the age spectra determined from the DAS differ from the GCM spectra. For the diabatic trajectory calculations, the age spectrum is too broad as a result of too much exchange between the tropics and mid-latitudes. The age spectrum determined using the kinematic trajectory calculation is less broad and lacks an age offset; both of these features are due to excessive vertical dispersion of parcels. The tropical and mid-latitude mean age difference between the diabatically and kinematically determined age-spectra is about one year, the former being older. The CTM calculation of the age spectrum using the DAS winds shows the same dispersive characteristics of the kinematic trajectory calculation. These results suggest that the current DAS products will not give realistic trace gas distributions for long integrations; they also help explain why the mean ages determined in a number of previous DAS driven CTM's are too young compared with observations. Finally, we note trajectory-generated age spectra show significant age anomalies correlated with the seasonal cycles, and these anomalies can be linked to year-to-year variations in the tropical heating rate. These anomalies are suppressed in the CTM spectra suggesting that the CTM transport is too diffusive.

Schoeberl, Mark R.

Eight Year Climatologies from Observational (AIRS) and Model (MERRA) Data

We examine climatologies derived from eight years of temperature, water vapor, cloud, and trace gas observations made by the Atmospheric Infrared Sounder (AIRS) instrument flying on the Aqua satellite and compare them to similar climatologies constructed with data from a global assimilation model, the Modern Era Retrospective-Analysis for Research and Applications (MERRA). We use the AIRS climatologies to examine anomalies and trends in the AIRS data record. Since sampling can be an issue for infrared satellites in low earth orbit, we also use the MERRA data to examine the AIRS sampling biases. By sampling the MERRA data at the AIRS space-time locations both with and without the AIRS quality control we estimate the sampling bias of the AIRS climatology and the atmospheric conditions where AIRS has a lower sampling rate. While the AIRS temperature and water vapor sampling biases are small at low latitudes, they can be more than a few degrees in temperature or 10 percent in water vapor at higher latitudes. The largest sampling biases are over desert. The AIRS and MERRA data are available from the Goddard Earth Sciences Data and Information Services Center (GES DISC). The AIRS climatologies we used are available for analysis with the GIOVANNI data exploration tool. (see, http://disc.gsfc.nasa.gov).

Hearty, Thomas

Analysis of Forest Foliage Using a Multivariate Mixture Model

Data with wet chemical measurements and near infrared spectra of ground leaf samples were analyzed to test a multivariate regression technique for estimating component spectra which is based on a linear mixture model for absorbance. The resulting unmixed spectra for carbohydrates, lignin, and protein resemble the spectra of extracted plant starches, cellulose, lignin, and protein. The unmixed protein spectrum has prominent absorption spectra at wavelengths which have been associated with nitrogen bonds.

Hlavka, C. A.

Analyzing System on A Chip Single Event Upset Responses using Single Event Upset Data, Classical Reliability Models, and Space Environment Data

We are investigating the application of classical reliability performance metrics combined with standard single event upset (SEU) analysis data. We expect to relate SEU behavior to system performance requirements. Our proposed methodology will provide better prediction of SEU responses in harsh radiation environments with confidence metrics. single event upset (SEU), single event effect (SEE), field programmable gate array devises (FPGAs)

single event effect (SEE)

Analysis of Skylab/Apollo Telescope Mount S-056 observations based on a force-free magnetic field model

Data obtained from the S-056 X-ray experiment on Skylab/ATM have been analyzed based on the assumption that the magnetic fields in the chromosphere and lower corona are force-free. Underlying the analysis is the hypothesis that the observed X-ray filaments coincide with magnetic field lines. The photographic recording of the filaments can then be compared with the projection along the line of sight of the computed magnetic field lines of the model. Ground-based observations of the longitudinal magnetic field component complement the X-ray data and are used in the theoretical interpretation.

Meyer, R. X.

Analysis of Sting Balance Calibration Data Using Optimized Regression Models

Calibration data of a wind tunnel sting balance was processed using a candidate math model search algorithm that recommends an optimized regression model for the data analysis. During the calibration the normal force and the moment at the balance moment center were selected as independent calibration variables. The sting balance itself had two moment gages. Therefore, after analyzing the connection between calibration loads and gage outputs, it was decided to choose the difference and the sum of the gage outputs as the two responses that best describe the behavior of the balance. The math model search algorithm was applied to these two responses. An optimized regression model was obtained for each response. Classical strain gage balance load transformations and the equations of the deflection of a cantilever beam under load are used to show that the search algorithm s two optimized regression models are supported by a theoretical analysis of the relationship between the applied calibration loads and the measured gage outputs. The analysis of the sting balance calibration data set is a rare example of a situation when terms of a regression model of a balance can directly be derived from first principles of physics. In addition, it is interesting to note that the search algorithm recommended the correct regression model term combinations using only a set of statistical quality metrics that were applied to the experimental data during the algorithm s term selection process.

Ulbrich, N.