Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗

Continuous integration data-driven platform of industrial-scale subsurface storage for real-time analytics

This project helped address the growing need for efficient and scalable models to support geological carbon and energy storage, which are crucial for achieving net-zero emissions. Traditionally accurate high-fidelity numerical models have been used to simulate relevant storage processes under a handful of processes, however such models are computationally demanding, making uncertainty quantification impractical. Consequently, we first developed a machine learning framework, based on Graph Neural Operators (GNOs), to improving the accuracy of model predictions for a fixed computational budget. We then developed an Ensemble of Improved Neural Operators (ENO), which uses bagging and Monte Carlo dropout techniques, to further improve prediction accuracy. Lastly, we developed the way to explain progressive transfer learning methods to reduce the amount of training data and computational cost of training (i.e., reduce trainable parameters) when using our models for multiple storage sites. Our numerical investigation, which used real-world case studies, demonstrated that our framework can significantly improve the safety and efficiency of geological storage operations, with potential applications in other domains such as geothermal reservoirs and climate modeling.

54 ENVIRONMENTAL SCIENCES↗

EMPDF : inferring the Milky Way mass with data-driven distribution function in phase space

We introduce the emPDF (empirical distribution function), a novel dynamical modelling method that infers the gravitational potential from kinematic tracers with optimal statistical efficiency under the minimal assumption of steady state. emPDF determines the best-fitting potential by maximizing the similarity between instantaneous kinematics and the time-averaged phase-space distribution function (DF), which is empirically constructed from observation upon the theoretical foundation of oPDF (Han et al. 2016). This approach eliminates the need for presumed functional forms of DFs or orbit libraries required by conventional DF- or orbit-based methods. emPDF stands out for its flexibility, efficiency, and capability in handling observational effects, making it preferable to the popular Jeans equation or other minimal assumption methods, especially for the Milky Way (MW) outer halo where tracers often have limited sample size and poor data quality. We apply emPDF to infer the MW mass profile using Gaia DR3 data of satellite galaxies and globular clusters, obtaining enclosed masses of M (,r) = 26±8, 46±8, 90±13⁠, and 149±40 x 10 10 M ⊙ at r = 30, 50, 100⁠, and 200 kpc, respectively. These are consistent with the updated constraints from simulation-informed DF fitting (Li et al. 2020). While the simulation-informed DF offers superior precision owing to the additional information extracted from simulations, emPDF is independent of such supplementary knowledge and applicable to general tracer populations. emPDF is currently implemented for tracers with complete 6D kinematics within spherical potentials, but it can potentially be extended to address more general problems.

Astrophysics of Galaxies (astro-ph.GA)↗

A data-driven approach to real-time vertical position estimation for NSTX-U vertical stability control

In this paper, a database of 77 996 plasma equilibrium reconstructions from 727 discharges during the initial operation of the NSTX-U spherical tokamak is analyzed to develop a statistically robust model of the plasma vertical position for real-time control. A variety of regression models are developed and tested, ranging in complexity from linear models to deep neural networks, and including input signals ranging from the four pairs of flux loops used historically on NSTX-U up to the full set of 389 real-time signals available to the plasma control system. A linear model based on 140 real-time magnetics signals is found to offer excellent accuracy, with a coefficient of determination R 2 = 0.906. The robustness of this model to limited training data, new operating scenarios, and signal errors is tested, and a procedure is demonstrated to tune the model parameters to optimize its robustness. A time-dependent plasma equilibrium solver, TokaMaker, is used to simulate vertical stability control in NSTX-U, demonstrating that it should be possible to iteratively tune the parameters of a linear vertical position model to stabilize both positive and negative triangularity plasmas in future experiments.

magnetic diagnostics↗

A data-driven method to estimate the antiproton background in the Mu2e experiment

The Mu2e experiment at Fermilab will search for the Charged Lepton Flavour Violating (CLFV) process of coherent, neutrinoless µ− → e − conversion in the field of an aluminum nucleus. The expected signal is a monochromatic electron with the energy of 104.97 MeV, slightly below the muon rest mass. Observation of a CLFV process would provide unambiguous evidence for Beyond the Standard Model (BSM) physics. Mu2e is sensitive to a wide range of BSM models and has the capability to distinguish between them, guiding us towards the most accurate models. The key features of the Mu2e experiment are: (1) a high intensity pulsed negative muon beam with about 1010 stopped µ −/s, and (2) a sophisticated superconducting solenoid system with a gradient magnetic field to form and guide the intense muon beam to the target. The Mu2e physics data taking is expected to begin in 2027. For Run I, the expected 5σ discovery sensitivity is Rµe = 1.2 × 10−15, with a total expected background of 0.11 ± 0.03 events. In the absence of a signal, the expected upper limit is Rµe < 6.2 × 10−16 at 90% CL. The success of this experiment hinges on the accurate estimation of the background from various SM processes that could provide signal-like electrons. One of the background processes is antiprotons annihilating in the stopping target to produce signal like electrons through π0 → γγ decays followed by γ conversions, and π− → µ−ν¯ decays followed by µ− decay. It is a relatively small background with large uncertainty (100%) due to the lack of antiproton production cross section information for the Mu2e proton beam energy of 8 GeV. We have developed a novel methodology to estimate the antiproton background in-situ. This forms the main theme of the thesis. We observed that at Mu2e energies, antiproton annihilation in the stopping target is the only source of events with multiple, simultaneous particle trajectories. From Geant4 simulations, only about 0.2% of the simulated antiproton annihilation events have a signal-like electron. Meanwhile, ∼ 5% of events have multiple reconstructible particle tracks per event. Therefore, we have devised a methodology to reconstruct the multi-track events and estimate the antiproton background by exploiting the large ratio of the production rates of the two final states.

Chithirasreemadam, Namitha [Pisa U.] (ORCID:000000↗

Towards Improving luminosity using optics tuning and data-driven methods

The results of Run 24 experiments at Relativistic Heavy Ion Collider (RHIC) for improving luminosity using optics tuning are presented in this study. In the first experiment, MADx matching was used to output magnet strengths corresponding to specific s star movements around Interaction Region 8 (IR8). The corresponding Zero Degree Calorimeter (ZDC) signal was measured in place of luminosity, and Bayesian Optimization aids search of optimal movements. It was found that values retrieved from matching were inaccurate, resulting in negative feedback loops. The second experiment focused on calculating accurate s star movements. The matching method was replaced with a linear sensitivity matrix, directly relating optics to power supply, and its null space was used to fit constraints such as hysteresis effects. At the experiment, beam losses were observed at collimators around boundary of IR8, which were fixed for the third experiment. Dynamic mode decomposition was also introduced to improve quality of turn-by-turn (TBT) data as well as accuracy and consistency of optics measurements at IR8. These improvements will be tested in the experiment of next RHIC run for luminosity optimization.

Accelerator Physics↗

Creating Data-Driven Vector Visualizations of Satellite Orbit Tracks Using NASA GIBS and Worldview

NASA Earth Observing System (EOS) currently operates dozens of remote sensing satellites, many of which can be viewed directly in NASA’s open-source Worldview application. Much of this satellite imagery can be viewed in near-real time as it is processed and served by NASA’s Global Imagery Browse Service (GIBS). To better educate users on the time and location of imagery, GIBS serves orbit track specific layers for each satellite. Worldview has historically served these layers as raster images but recent updates have enabled the application to now serve these layers using vector tiles. With the release of Worldview v3.0, orbit track layers can be displayed using mapbox vector tiles (MVT). This visualization format allows users to not only view and change the color of orbit track layers, as they could do previously with rasters, but also inspect individual vector points and filter layers by specific parameters such as time. The data contained within a MVT is further enhanced in Worldview with the combination of a JSON description file served from GIBS used to describe the MVT data. This presentation will provide an overview of the process of consuming orbit track vector tiles and data files from GIBS using a pipeline to configure, build and ultimately display the orbit tracks in Worldview. Furthermore, the presentation aims to describe how others can leverage our open-source code to display and enhance vector layers in their own applications.

Rice, Zachary↗

Physics-coupled data-driven design of high-temperature alloys

We present a materials design loop, which streamlines physics-coupled machine learning (ML) surrogate models to discover new alloy chemistries with improved properties. The efficacy is demonstrated by discovering a high-temperature alumina-forming austenitic (AFA) stainless steel with enhanced creep, followed by experimental validation. The ML models have been trained using a well-curated, highly consistent experimental dataset augmented with synthetic microstructural features from a computational thermodynamic approach. We have populated a large number of hypothetical AFA alloys to explore the high-dimensional composition space and have predicted their creep properties by providing the same synthetic input features obtained from the trained ML models. Uncertainties from the ML training were taken as thresholds for truncating predicted results to identify alloys with improved or deteriorated creep. Individual elemental compositions have been determined via probability density distribution analysis from the group of alloys at the top and bottom of the predicted creep values for further virtual and experimental validations. In conclusion, we anticipate that this workflow can be applied to screen desired conditions, such as chemistry and processing parameters, in high-dimensional space through physics-guided data analytics.

Alloy design↗

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]↗

Experimental and data-driven characterization of window-induced air leakage in residential buildings

Windows contributes up to 40% of envelope heat losses and around 9% of total building energy consumption due to air leakage. In the U.S., 48 million homes still use single-pane windows. Although the U.S. has an estimated 1.4 billion windows in its building stock and about 24 million windows are installed annually, only around 29 million individual window replacements (∼2%) occur each year. To address this gap, this study generates empirical evidence by (1) evaluating the contribution of windows to whole-building air leakage in 20 residential buildings using blower door tests before and after window replacement and (2) assessing whether building and window characteristics influence the measured change. Most simulation studies assume that replacing windows not only lowers the U-factor but also reduces air leakage by 10–20%. However, this assumption lacks empirical validation, highlighting the need for experimental analysis of air leakage specifically associated with windows. Using blower door tests in accordance with ASTM E779–19, the results indicated an average reduction in air infiltration of 6.1% within the range of 0.5–19.30% across all buildings and no significant correlations were found between air leakage improvements and any building/window characteristics. This research aims to help homeowners, and energy modelers to provide empirical data on importance of upgrading to more energy-efficient windows, supporting energy-efficient building standards.

Air leakage↗

Exploring data-driven modeling of boundary layer transition

Prediction of laminar-turbulent transition in boundary layer flows is an important component of predicting the aerodynamic performance of a number of aerospace configurations. According to the CFD Vision 2030 [1], transition modeling represents acriticalarea in CFD simulation capability that will remain a pacing item for the foreseeable future. The fact thattransition can take placevia either one of a myriad possible paths adds to the challenges inreliable transition predictions, despite a limited knowledge of the relevant input parameters. In the low disturbance environments typical of flight applications, transition is often initiated by small amplitude disturbances in the form of linear instability waves of the laminar boundary layer. These disturbances amplify linearly at first and eventually undergo a sequence of nonlinear interactions that result in transition to turbulence. Because the nonlinear phase is rather rapid, the amplification of boundary layer instabilities is governed by the linearstability theory over a majority of the distance leading up to the onset of transition. Semi-empirical transition correlations based on the linear stability theory have been successful in explaining the observed trends in transition location within a broad class of flows. However, the application of stability theory is highly non-robust and often requires a significant domain expertise. Recent work at the NASA Langley Research Center has beenaimed at bridging the gap between physics based transition analyses such as those based on linear stability theory and practical applications that require transition prediction by users that may not be well versed in transition physics. The applications of deep learning have been at the center of these efforts. This presentation will focus on the progress achieved thus far, highlighting the applications of neural networks to selectedtransition scenarios across a range of Mach numbers and flow configuration, as well as the lessons learnedand remaining challengeswithrespect to the selection of training data and neural networks architectures, hyperparameter tuning, and the physical insights distilled from the otherwise black-box models.

M. R. Malik↗

A Data-Driven Exploration of the Impact of Renewable Energy on Inter-Area Oscillations in the U.S. Eastern Interconnection

As increasing amounts of renewable energy (RE) resources are incorporated into the bulk-power grid, power system oscillations are expected to change. This work investigates how RE generation impacts the frequency and damping ratio (DR) of two dominant inter-area modes in the U.S. Eastern Interconnection (EI) using regularly updated estimates collected over a 12-month period. Quantile regression is used to derive the correlation between operating conditions and mode properties, and a bootstrap method is used to quantify the uncertainty associated with the correlation estimates. Results show that with an increase in system load, the frequency of a mode decreases and DR increases. Evidence that increasing RE generation results in an increase in frequency and decline in DR was found for one of the two modes studied. This work shows that increasing RE levels will impact the properties of inter-area oscillations in the EI, but it does not indicate the presence of immediate threats to grid stability. The outlined approach can be used to periodically assess changing mode properties as RE levels continue to grow and flag stability concerns before they become serious reliability threats.

Inter-area oscillation, mode meters, quantile regr↗

Data‐Driven Engineering of Thermostable Collagen‐Mimetic Peptoid Triple Helices

Collagen-mimetic peptides (CMPs) are engineered molecules designed to replicate the triple-helical structure of natural collagen. A repeating x–y-Gly sequence is the defining motif of CMPs and is critical to their triple-helical structure and stability. Substitutions to the residues occupying the x and y positions present a means to modulate the CMP structure and properties. Peptoid residues—N-substituted glycine derivatives—present an attractive potential substitution due to their thermal stability, proteolytic resistance, biocompatibility, and diverse palette of non-natural side chains, but also tend to introduce a high degree of backbone flexibility that can diminish the stability of the triple helix. In this work, we report a computational active learning cycle comprising molecular dynamics simulation, Gaussian process regression, and Bayesian optimization to computationally identify a number of promising peptoid substitutions predicted to stabilize the desired quaternary structure through side chain interactions and produce stable peptoid-based collagen-like triple helices. To experimentally test the computational predictions, a top candidate identified by the screen was synthesized and imaged using scanning electron microscopy to resolve fibril-like bundles consistent with collagen-like triple helices. This work predicts a number of CMP peptoid substitutions capable of forming stable triple-helical structures, presents a generalizable design strategy for engineering desired peptoid structures, and opens new avenues for the design of peptoid-based biomimetic materials.

active learning↗

Data-Driven Insights into the Structural Essence of Plasticity in High-Entropy Alloys

The heterogeneous mechanical response of a crystalline alloy with multiple principal elements was investigated using molecular dynamics simulations. The local configuration of the alloy in its quiescent state was characterized by the variables derived from the gyration tensor and the atomic electronegativity. A multivariate analysis identified the geometric and chemical factors that influenced the atomic packing variations. Further, upon straining, the non-affine displacement exhibited spatial heterogeneity. A statistical correlation was established between the local yield events and the specific features of the local configuration. Our findings, validated by the performance metrics analysis, provided a structural criterion for the instability mechanisms in high-entropy alloys (HEAs) and enhanced the understanding of their plasticity.

36 MATERIALS SCIENCE↗

Data-Driven Surrogate Modeling with Microstructure-Sensitivity of Viscoplastic Creep in Grade 91 Steel

Abstract To support the development of advanced steel alloys tailored to withstand extreme conditions, it is imperative to account for the mechanical performance of components, while considering the influence of local microstructure on the macroscopic response. To this end, this study focuses on the development of microstructure-sensitive constitutive models for the mechanical response of Grade 91 steel exposed to extreme thermo-mechanical environments. Polynomial chaos expansion (PCE) surrogates are used to emulate high-fidelity polycrystal simulations of the viscoplastic response of Grade 91 steel as a function of the microstructure fingerprint (e.g., dislocations and precipitates). To cover a wide temperature–stress domain, two separate PCE surrogates—one that captures softening and the other that captures hardening behavior—are combined using another (sparse) Gaussian process regression model. The resulting constitutive creep surrogate model is integrated within the MOOSE finite element framework to simulate the intricate effects of microstructure, in particular MX-phase precipitates, on a component with a graded microstructure. Surrogate sensitivity analysis is applied to quantify the relevant impact of spatially varying microstructure on the creep response in a test-case involving a Grade 91 alloy with a prototypical weld.

36 MATERIALS SCIENCE↗

Data-driven equation-free dynamics applied to many-protein complexes: The microtubule tip relaxation

Microtubules (MTs) constitute the largest components of the eukaryotic cytoskeleton and play crucial roles in various cellular processes, including mitosis and intracellular transport. The property allowing MTs to cater to such diverse roles is attributed to dynamic instability, which is coupled to the hydrolysis of GTP (guanosine-5'-triphosphate) to GDP (guanosine-5'-diphosphate) within the β-tubulin monomers. Understanding the equilibrium dynamics and the structural features of both GDP- and GTP-complexed MT tips, especially at an all-atom level, remains challenging for both experimental and computational methods because of their dynamic nature and the prohibitive computational demands of simulating large, many-protein systems. This study employs the “equation-free” multiscale computational method to accelerate the relaxation of all-atom simulations of MT tips toward their putative equilibrium conformation. Using large MT lattice systems (14 protofilaments × 8 heterodimers) comprising ~21-38 million atoms, we applied this multiscale approach to leapfrog through time and nearly double the computational efficiency in realizing relaxed all-atom conformations of GDP- and GTP-complexed MT tips. Commencing from an initial 4 μs unbiased all-atom simulation, we interleave coarse projective “equation-free” jumps with short bursts of all-atom molecular dynamics simulation to realize an additional effective simulation time of 1.875 μs. Our 5.875 μs of effective simulation trajectories for each system expose the subtle yet essential differences in the structures of MT tips as a function of whether β-tubulin monomer is complexed with GDP or GTP, as well as the lateral interactions within the MT tip, offering a refined understanding of features underlying MT dynamic instability. Furthermore, the approach presents a robust and generalizable framework for future explorations of large biomolecular systems at atomic resolution.

Wu, Jiangbo [University of Chicago, IL (United Sta↗