Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

ARM FY2026 Radar Plan

The U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) User Facility maintains a suite of advanced atmospheric radar systems that serve as critical tools in ARM’s mission to provide continuous, high-quality observations for advancing the understanding and modeling of atmospheric processes. These radar systems enable detailed characterization of clouds, precipitation, and dynamic structures in the atmosphere, supporting a broad range of scientific applications. The number of deployed systems exceeds what current staffing levels can fully support for continuous 24/7/365 operation. As such, it is essential to have a clearly defined and community-informed plan that prioritizes radar operations and communicates ARM’s strategy for sustaining and evolving these observational assets. This FY2026 Radar Plan outlines ARM’s approach to managing its radar portfolio—balancing scientific impact, operational feasibility, and long-term sustainability. It reflects ARM’s continued commitment to delivering calibrated, well-documented radar data products that enable process-level studies and support the development and evaluation of weather and climate models. Through this plan, ARM aims to ensure transparency in decision-making, alignment with user needs, and support for innovative science across the facility’s fixed and mobile observatories. Given uncertainties around the Fiscal Year (FY) 2026 budget, this plan was developed to assume business as usual and will be updated as budgets and plans may change. It should be noted that, given the limited timeframe involved, this plan will be more succinct than previous plans.

47 OTHER INSTRUMENTATION↗

Validation of a Global Geospace Model With a Systems Science Approach Based on Canonical Correlation Analysis

A systems science approach based on canonical correlation analysis (CCA) is applied as a new, behavioral way to validate global geospace models. The biggest novelty of the technique is that it validates models at a system level, whereby a side‐by‐side comparison is performed of CCA applied to a 30‐day observational and the corresponding simulation data sets comprising quiet, moderate and active times. The simulation used the Multiscale Atmosphere‐Geospace Environment (MAGE) model. It is shown that (a) CCA must be combined with sensitivity analysis to be effective, (b) the MAGE model generally reproduces the observed behavior (more so for quieter time intervals), quantified by the intercorrelations between different variables and (c) the technique identifies the SuperMAG SML index as a quantity for which refinements of the model are needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Refining Planetary Boundary Layer Height Retrievals From Micropulse‐Lidar at Multiple ARM Sites Around the World

Abstract Knowledge of the planetary boundary layer height (PBLH) is crucial for various applications in atmospheric and environmental sciences. Lidar measurements are frequently used to monitor the evolution of the PBLH, providing more frequent observations than traditional radiosonde‐based methods. However, lidar‐derived PBLH estimates have substantial uncertainties, contingent upon the retrieval algorithm used. In addressing this, we applied the Different Thermo‐Dynamic Stabilities (DTDS) algorithm to establish a PBLH data set at five separate Department of Energy's Atmospheric Radiation Measurement sites across the globe. Both the PBLH methodology and the products are subject to rigorous assessments in terms of their uncertainties and constraints, juxtaposing them with other products. The DTDS‐derived product consistently aligns with radiosonde PBLH estimates, with correlation coefficients exceeding 0.77 across all sites. This study delves into a detailed examination of the strengths and limitations of PBLH data sets with respect to both radiosonde‐derived and other lidar‐based estimates of the PBLH by exploring their respective errors and uncertainties. It is found that varying techniques and definitions can lead to diverse PBLH retrievals due to the inherent intricacy and variability of the boundary layer. Our DTDS‐derived PBLH data set outperforms existing products derived from ceilometer data, offering a more precise representation of the PBLH. This extensive data set paves the way for advanced studies and an improved understanding of boundary‐layer dynamics, with valuable applications in weather forecasting, climate modeling, and environmental studies.

54 ENVIRONMENTAL SCIENCES↗

Characterizing Wet Season Precipitation in the Central Amazon Using a Mesoscale Convective System Tracking Algorithm

To comprehensively characterize convective precipitation in the central Amazon region, we utilize the Python FLEXible object TRacKeR (PyFLEXTRKR) to track mesoscale convective systems (MCSs) observed through satellite measurements and simulated by the Weather Research and Forecasting model at a convection-permitting resolution. This study spans a 2-month period during the wet seasons of 2014 and 2015. We observe a strong correlation between the MCS track density and accumulated precipitation in the Amazon basin. Key factors contributing to precipitation, such as MCS properties (number, size, rainfall intensity, and movement), are thoroughly examined. Our analysis reveals that while the overall model produces fewer MCSs with smaller mean sizes compared to observations, it tends to overpredict total precipitation due to excessive rainfall intensity for heavy rainfall events (≥10 mm hr –1 ). These biases in simulated MCS properties could vary with the constraints on the convective background environment. Moreover, while the wet bias from heavy (convective) rainfall outweighs the dry bias in light (stratiform) rainfall, the latter can be crucial, particularly when MCS cloud cover is significantly underestimated. A case study for 1 April 2014 highlights the influence of environmental conditions on the MCS lifecycle and identifies an unrealistic model representation in both stratiform and convective precipitation features.

54 ENVIRONMENTAL SCIENCES↗

HPC-Enabled Optimization of High Temperature Heat Exchangers (CRADA Final Report)

This project was a collaborative effort between Lawrence Livermore National Security, LLC (LLNS) as manager and operator of Lawrence Livermore National Laboratory (LLNL) and Materials Sciences, LLC, to develop a technology for design and optimization of heat exchangers using powerful desktop and laptop computers. The project was originally designated as a 12-month project, and consisted of 3# major tasks and the following 8# major deliverables: 1) CFD models of 3D heat exchangers based on existing and new geometry. 2) Validation against experimental data provided by MSC and published in the literature. 3) CFD models of 3D unit cells based on TPMS. 4) Surrogate models capable of delivering the gradients of the homogenized properties with respect to the parametrization. 5) 3D design methodology using TO algorithms. 6) Conventional reference and topology optimized designs. 7) 3D optimized designs stored in a 3D printer build format. 8) Verification of the improved performance. All of the deliverables for this project were successfully completed with two no-cost time extensions.

13 HYDRO ENERGY↗

Learning Constitutive Relations From Soil Moisture Data via Physically Constrained Neural Networks

Abstract The constitutive relations of the Richardson‐Richards equation encode the macroscopic properties of soil water retention and conductivity. These soil hydraulic functions are commonly represented by models with a handful of parameters. The limited degrees of freedom of such soil hydraulic models constrain our ability to extract soil hydraulic properties from soil moisture data via inverse modeling. We present a new free‐form approach to learning the constitutive relations using physically constrained neural networks. We implemented the inverse modeling framework in a differentiable modeling framework, JAX, to ensure scalability and extensibility. For efficient gradient computations, we implemented implicit differentiation through a nonlinear solver for the Richardson‐Richards equation. We tested the framework against synthetic noisy data and demonstrated its robustness against varying magnitudes of noise and degrees of freedom of the neural networks. We applied the framework to soil moisture data from an upward infiltration experiment and demonstrated that the neural network‐based approach was better fitted to the experimental data than a parametric model and that the framework can learn the constitutive relations.

54 ENVIRONMENTAL SCIENCES↗

How robust are estimates of key parameters in standard viral dynamic models?

Mathematical models of viral infection have been developed, fitted to data, and provide insight into disease pathogenesis for multiple agents that cause chronic infection, including HIV, hepatitis C, and B virus. However, for agents that cause acute infections or during the acute stage of agents that cause chronic infections, viral load data are often collected after symptoms develop, usually around or after the peak viral load. Consequently, we frequently lack data in the initial phase of viral growth, i.e., when pre-symptomatic transmission events occur. Missing data may make estimating the time of infection, the infectious period, and parameters in viral dynamic models, such as the cell infection rate, difficult. However, having extra information, such as the average time to peak viral load, may improve the robustness of the estimation. Here, we evaluated the robustness of estimates of key model parameters when viral load data prior to the viral load peak is missing, when we know the values of some parameters and/or the time from infection to peak viral load. Although estimates of the time of infection are sensitive to the quality and amount of available data, particularly pre-peak, other parameters important in understanding disease pathogenesis, such as the loss rate of infected cells, are less sensitive. Viral infectivity and the viral production rate are key parameters affecting the robustness of data fits. Fixing their values to literature values can help estimate the remaining model parameters when pre-peak data is missing or limited. We find a lack of data in the pre-peak growth phase underestimates the time to peak viral load by several days, leading to a shorter predicted growth phase. On the other hand, knowing the time of infection (e.g., from epidemiological data) and fixing it results in good estimates of dynamical parameters even in the absence of early data. While we provide ways to approximate model parameters in the absence of early viral load data, our results also suggest that these data, when available, are needed to estimate model parameters more precisely.

59 BASIC BIOLOGICAL SCIENCES↗

Airborne LiDAR to Improve Canopy Fuels Mapping for Wildfire Modeling

Increasing conflict between wildfire and the built environment has increased the need for more up-to-date and finer resolution canopy fuels data to improve wildfire modeling and associated risk forecasts. The US Forest Service and US Department of the Interior’s LANDFIRE product, which provides 30-m resolution canopy fuels data for the entire US, is one of the most widely used sources of fuels data. However, the last complete mapping effort for LANDFIRE is based on 2016 conditions, and subsequent updates reflect disturbances 1-2 years behind the release year. Airborne systems equipped with Light Detection and Ranging (LiDAR) sensors can be deployed to actively sense canopy structure and estimate canopy fuels data (cover, height, base height, bulk density) at finer resolutions. Canopy base height (CBH) and canopy bulk density (CBD) are difficult to measure both in the field and in LiDAR point clouds. Still, they are important for accurately modeling crown fires, which are often intense and difficult to contain. Additionally, point cloud datasets are large, and calculations require efficient utilization of computational resources. To address these challenges, we are working on an approach that uses openly available National Ecological Observatory Network (NEON) airborne LiDAR data, with calculations processed in the R programming language and parallelized through the lidR package. CBH and CBD are often derived from tree height, diameter at breast height, and species-specific allometries using the Fire and Fuels Extension of the Forest Vegetation Simulator (FFE-FVS). We aim to test if airborne LiDAR can estimate CBH and CBD without the use of empirical equations. Reliable estimates of canopy fuels data directly from airborne LiDAR could streamline quick, fine-resolution updates for use in wildfire behavior models.

54 ENVIRONMENTAL SCIENCES↗

Multi-Scale Modeling of the Evolution of Structure and Properties in Materials for Nuclear Energy Applications [Slides]

Nuclear energy is an important component of an overall strategy to address climate change. Idaho National Laboratory (INL) is the U.S. Department of Energy’s primary facility for research and development in nuclear science and technology for energy generation, supporting the improvement and life extension of the existing reactor fleet and the development and licensing of new reactor designs. Computational modeling is an important component of these activities, particularly in the area of materials for nuclear applications, where experimental data can be very challenging and expensive to acquire, and where data is especially scarce for new reactor designs. INL has used multi-scale modeling – linking atomistic, mesoscale, and engineering scales – to improve the ability to predict the performance of materials for nuclear energy applications. In this talk, I will give an overview of the approach and tools used, and several examples of application, including performance of nuclear fuels, understanding radiation-driven formation of nanoscale void and gas bubble superlattices, and powder densification through electric field assisted sintering.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Neural network representations of multiphase Equations of State

Abstract Equations of State model relations between thermodynamic variables and are ubiquitous in scientific modelling, appearing in modern day applications ranging from Astrophysics to Climate Science. The three desired properties of a general Equation of State model are adherence to the Laws of Thermodynamics, incorporation of phase transitions, and multiscale accuracy. Analytic models that adhere to all three are hard to develop and cumbersome to work with, often resulting in sacrificing one of these elements for the sake of efficiency. In this work, two deep-learning methods are proposed that provably satisfy the first and second conditions on a large-enough region of thermodynamic variable space. The first is based on learning the generating function (thermodynamic potential) while the second is based on structure-preserving, symplectic neural networks, respectively allowing modifications near or on phase transition regions. They can be used either “from scratch” to learn a full Equation of State, or in conjunction with a pre-existing consistent model, functioning as a modification that better adheres to experimental data. We formulate the theory and provide several computational examples to justify both approaches, highlighting their advantages and shortcomings.

Science & Technology - Other Topics↗

Data and scripts associated with “When do Riverine Systems 'Feel the Burn'? Simulating How Burn Extent and Severity Modulate Hydrologic Controls on Biogeochemical Export” (v2)

This data package is associated with the publication “When do Riverine Systems 'Feel the Burn'? Simulating How Burn Extent and Severity Modulate Hydrologic Controls on Biogeochemical Export” published in Water Resources Research (Wampler et al. 2025; preprint: https://doi.org/10.22541/essoar.174438106.63564767/v1). This study used the Soil and Water Assessment Tool (SWAT), a processed based model to explore the impacts of area burned and burn severity on streamflow, nitrate, and dissolved organic carbon (DOC) in two test basins: a semi-arid, mixed land use basin and a humid, primarily forested basin. We developed 1800 wildfire scenarios that we ran in each basin: 20 different burn extents (5 to 100% by 5%), 3 different burn severities (low, moderate, and high), and 30 different post-fire precipitation scenarios. We also ran an additional 30 scenarios associated with no wildfire for the 30 post-fire precipitation scenarios. For each scenario we were interested in the change in runoff ratio (streamflow) and average concentration and annual loads (nitrate and DOC) across the wildfire scenarios. This data package contains the data and scripts required to build SWAT models for the two test basins, create and run the wildfire scenarios, and generate the data summaries and figures used in the associated manuscript. This data package was originally published in March 2025. It was updated in January 2026 (v2; new and modified files) to include the final files after the manuscript went through reviews. See the change history section below for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

54 ENVIRONMENTAL SCIENCES↗

Aspen Open Jets: unlocking LHC data for foundation models in particle physics

Foundation models are deep learning models pre-trained on large amounts of data which are capable of generalizing to multiple datasets and/or downstream tasks. This work demonstrates how data collected by the CMS experiment at the Large Hadron Collider can be useful in pre-training foundation models for HEP. Specifically, we introduce the AspenOpenJets (AOJs) dataset, consisting of approximately 178 M high p T jets derived from CMS 2016 Open Data. We show how pre-training the OmniJet-α foundation model on AOJs improves performance on generative tasks with significant domain shift: generating boosted top and QCD jets from the simulated JetClass dataset. In addition to demonstrating the power of pre-training of a jet-based foundation model on actual proton–proton collision data, we provide the ML-ready derived AOJs dataset for further public use.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

On the transferability of residence time distributions in two 10-km long river sections with similar hydromorphic units

Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.

54 ENVIRONMENTAL SCIENCES↗

Blending Behavioral Science and Physics-Based Models Inform Equitable Decarbonization Pathways in the US Housing Stock: Preprint

A just energy transition is an imperative of the Biden-Harris Administration, emphasizing the equitable distribution of benefits through energy-efficient and decarbonizing household technologies. Understanding the factors that increase a household's willingness to adopt these technologies helps policymakers implement more targeted and effective approaches. Our research addresses this by blending 550,000 housing stock types and energy simulation data with a nationally representative survey on residential technology adoption and decision-making (n=10,000). We identify energy equity gaps across tenure and income, highlighting disparities in energy burdens and insecurity. Findings show that households with prior modification experience are more willing to renovate, suggesting that small-scale retrofit programs could foster greater willingness. Energy secure but burdened homeowners are least willing to modify, highlighting the need to consider energy bill perceptions in policy design. Nearly half of US households that are energy burdened also face energy insecurity, with a significant gap in assistance for low-income households. We emphasize the need to understand household perceptions to improve policy. This research underscores the importance of understanding household behaviors to improve policy effectiveness, offering actionable insights for policymakers to promote equitable housing upgrades and advance a decarbonized future.

behavioral science↗

Wind Turbine Sound Setbacks and Supply Curves: Ordinances and Extrapolated Trends, 110 Hub Height, 130 Rotor Diameter

This dataset provides a comprehensive set of wind turbine sound setbacks from every residential structure in the contiguous United States (CONUS). A sound setback is defined as the minimum required distance between a residential structure and a hypothetical turbine installation site to ensure that modeled sound levels received at the residence do not exceed local sound ordinances, which are commonly expressed in A-weighted decibels (dBA). Therefore, sound setbacks are a local spatial assessment combining multiple factors, including the sound pressure curve as a function of the observer location (distance and direction) relative to the turbine, local sound regulations, and the geographical distribution of residential structures. The dataset is organized into multiple scenario-based products, detailed as follows: 1. Existing and extrapolated sound setbacks. An existing scenario characterizes sound setbacks only in states or counties that have implemented sound regulations as of 2022. The extrapolated scenarios extend a constant sound threshold to counties that lack explicit sound regulations, with thresholds ranging from 35 to 60 dBA, in 5-dBA increments reflecting the variation observed in current sound ordinances. 2. Sound setbacks in directional and worst scenarios. The directional scenario accounts for the distance and orientation of residential structures relative to a hypothetical turbine location, utilizing the turbine's sound emissions in that specific direction. In contrast, the worst scenario takes loudest sound level at each distance step from the turbine, irrespective of directional considerations, which aligns with current industry practice. 3. Supply curves for Open and Reference Access scenarios. This dataset includes supply curves generated by the reV model, which integrates each of the above sound setbacks into both Open and Reference siting scenarios. In addition, two Open and Reference baselines scenarios were included which do not consider sound setbacks for comparative analysis. All sound setback data are stored in TIF files, with partial maps of the data provided in PNG format. The values in the sound setback raster range from 0 to 1, representing the fraction of developable land within a 90 meter by 90 meter pixel due to sound ordinances. A value of 0 indicates areas where wind energy development is prohibited, while a value of 1 signifies areas fully permissible. The wind turbine parameters used in the sound modeling are based on the land-based turbine from International Energy Agency (IEA), featuring a rated electrical power of 3.4 MW, a rotor diameter of 130 meters, and a hub height of 110 meters. The atmospheric conditions, including wind speed/direction, turbulence, air temperature, relative humidity, and air pressure, that drive the sound generation are obtained from the WIND Toolkit dataset.

Array↗

G-Mapper: Learning a Cover in the Mapper Construction

The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a “nice” Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in 𝑘-means by iteratively applying the Anderson–Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. In conclusion, experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.

G-means clustering↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗