Search NASA⌕ Search

SEARCH · Search NASA

Results for “DATA BASES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING↗

Satellite Embedding-Based Population Imputation for Areas with Missing Building Footprint Data: A Computer Vision-Based Approach

High-resolution population modeling is important for supporting effective decision-making across diverse sectors. LandScan Mosaic generates population estimates at the level of individual buildings and aggregates them to 3 arc-second grids, and this approach performs well in regions where building footprint data are comprehensive and reliable. However, large portions of the globe still suffer from incomplete, sparse, or entirely missing building stock datasets, creating a structural limitation for strictly building-based population models. To address this research gap, this study proposes a computer vision-based framework that employs Google Earth Engine satellite embeddings and UNet, which allows us to directly impute grid-level population estimates in building-data-deficient areas. Applied to Taiwan as a case study, the framework achieved strong predictive performance with R$^{2}$ of 0.89, RMSE of 18.70, and MAE of 8.41, outperforming traditional machine learning approaches. Notably, the proposed framework effectively addressed building false-positive errors inherent in Global Human Settlement Layer (GHSL) data, correctly identifying uninhabited areas that were erroneously classified as populated. The framework also offers significant advantages for global population mapping, particularly in terms of scalability and temporal consistency, thereby extending the coverage and accuracy of high-resolution population products in data-scarce regions worldwide. Urban planners, decision makers, and related stakeholders can obtain granular population distributions to support more accurate and targeted infrastructure investment, service delivery, resource allocation, and risk assessment decisions.

97 MATHEMATICS AND COMPUTING↗

SITCOMTN-162: Testing the implementation of Metadetection and Cell-Based Coadds on Abell 360 LSSTComCam data

The purpose of this technote is to test the technical quality of LSSTComCam commissioning data, specifically the Rubin_SV_38_7 field, by utilizing cell-based coadds and Metadetection by measuring the tangential and cross weak lensing shear profiles of the massive cluster Abell 360 (called A360 throughout the technote). The process entails generating the cell-based coadds for Metadetection to run on, identifying and removing cluster member galaxies, applying quality cuts, calibrating the shear measurements, and validation. Cell-based coadds and Metadetection are both currently in the process of being implemented within the LSST Science Pipelines at the time of this technote. There is substantial technical value in attempting a difficult measurement prior to full implementation. Measuring the tangential shear around A360 will showcase the current abilities of these algorithms, as well as highlight where work is still needed. As seen from the resulting shear profile of A360, the cell-based coadds and Metadetection are able to work in tandem to produce a shear catalog and resulting reduced shear profile. This technote is one part of a series studying A360 in order to both stress test the commissioning camera and demonstrate the technical capabilities of the Vera Rubin Observatory. We study the quality of the PSF modeling and impact it can have on cluster WL in [Combet et al., 2025], implementation of cell-based coadds and subsequent use for Metadetect [Sheldon et al., 2023] in this technote, photometric calibration in (in prep), source selection and photometric redshifts in [Adari et al., 2025], use of Anacal [Li et al., 2024] to produce a cluster shear profile in [Li et al., 2025], and background subtraction in this field and Fornax in [Zhou et al., 2025].

79 ASTRONOMY AND ASTROPHYSICS↗

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING↗

Comparison of Global Aboveground Biomass Estimates From Satellite Observations and Dynamic Global Vegetation Models

The global forest carbon stocks represent the amount of carbon stored in woody vegetation and are important for quantifying the ability of the global forests to sequester atmospheric CO 2 and to provide ecosystem services (e.g., timber) under climate change. The forest ecosystem carbon pool estimates are highly variable and poorly quantified in areas lacking forest inventory estimates. Here, we compare and analyze aboveground biomass (AGB) estimates from five satellite-based global data sets and nine dynamic global vegetation models (DVGMs). We find that across the data sets, mean AGB exhibits the largest variability around the tropical area. In addition, AGB shows a similar latitudinal trend but large variability among the data sets. Satellite-based AGB estimates are lower than those simulated by DVGMs. The divergence among the satellite-based AGB estimates can be driven by the methodology, input satellite products, and the forested areas used to estimate AGB. The modeled NPP, autotrophic respiration, and carbon allocation mostly drive the variability of AGB simulated by DGVMs. The future availability of a high-quality global forest area map is anticipated to improve AGB estimate accuracy and to reduce the discrepancies among different satellite- and model-based AGB estimates. Furthermore, we suggest the carbon-modeling community reexamine the methodology used to estimate AGB and forested areas for a more robust global forest carbon stock estimation.

54 ENVIRONMENTAL SCIENCES↗

From Data to Knowledge: A Graph-Based Reliability Approach to Assess System Health

With the goal of maximizing plant reliability and availability, complex systems such as nuclear power plants continuously monitor and record the performance and the health status of many components, assets, and systems. Such data may take the form of online monitoring data, condition reports, and maintenance reports and it carries the potential to provide system engineers with insights into anomalous behaviors or degradation trends as well as the possible causes behind them and to predict their direct consequences. The analysis of such data poses however few challenges. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly tackles these challenges, and it focuses on the integration of all these data elements in order to assist plant system engineers in analyzing component, assets, and systems performances and optimize maintenance activities. This is performed by 1) extracting knowledge from textual data via technical language processing methods, and 2) quantifying system, asset, and component health from numeric condition-based data. We rely on model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Numeric and textual data elements are then associated with an MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 MATHEMATICS AND COMPUTING↗

National Cancer Institute (NCI) Exposomic Linkage Protocol

The purpose of this work is to create point-level linkages of residential history data, which is provided by the Surveillance, Epidemiology, and End Results (SEER) program, to air pollution exposure data so that we can develop longitudinal measures of exposure and investigate their effects on cancer incidence, treatment response, and survival. The Louisiana, New Jersey, Kentucky, and Iowa registries were previously linked to LexisNexis residential history data through the National Cancer Institute (NCI). We will enhance the utility of the existing residential location data by geocoding addresses based on data from between 1995 and 2024 and spatially linking the locations to air pollution, indoor radon, and the US Environmental Protection Agency’s (EPA) Risk-Screening Environmental Indicators (RSEI) exposure data (Figure 1).

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Data-Driven Modeling and Control of Systems with Plasma-Surface Interactions (Final Technical Report)

This final technical report summarizes the activities and accomplishments in the period from February 2023 thru January 2026. The objective of the proposed research is to investigate the physical mechanisms and processes underlying the formation of structures and patterns in systems with plasma-surface interactions. In the past decades, there have been extensive studies on the interaction of glow discharges, dielectric barrier discharges, and arc discharges with confining or intervening surfaces. The advancement of the understanding of these phenomena is not only of fundamental scientific interest and relevance to the knowledge of the plasma state, but also with profound implications in various technological applications. The research will integrate theoretical, computational, and experimental work within an innovative framework of data assimilation, i.e., optimally combining model predictions with measurements. The scientific merit of this research has three aspects. Firstly, it extends the studies of plasma-surface interactions to systems with insulator surfaces and multi-layer systems, while existing studies are predominantly on electrode surfaces. Secondly, it expects to develop a novel data-driven modeling approach based on data assimilation to enhance the predictive and control capabilities, which could make transformative contributions to basic plasma research. Thirdly, it will shed new light on outstanding problems related to formation of patterns interfacing plasmas. This project also aims to launch an education and outreach initiative at Texas A&M University-Kingsville, a non-R1, minority-serving institution in South Texas. The initiative is structured as a four-tier pyramid. Tier one will be a webinar series for culture and capacity building to inform broader audience in the region about the research fields of plasma science and engineering. Tier two will be the creation and offering of an upper-level undergraduate course on introductory plasma physics, which will help with the recruitment for the upper tiers. On tier three, we will engage and mentor senior design students to conduct work toward the research goal of this project. There will also be a certificate program on general plasma science for undergrad and graduate students, part of which will be lab training at Princeton University. Tier four will be the supervision and mentoring of Ph.D. students. Therefore, this project will systematically expand the talent pipeline, broaden participation from communities historically and geographically underrepresented in DOE SC research portfolio, significantly improve the research and education capacity at the PI’s institution, and contribute to developing a diverse workforce in plasma science and engineering.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Short-period post-common envelope binaries with Balmer emission from SDSS and LAMOST based on ZTF photometric data

ABSTRACT We present here 55 short-period post-common envelope binaries (PCEBs) containing a hot white dwarf (WD) and a low-mass main sequence (MS). Based on the photometric data from Zwicky Transient Facility survey data Release 19 (ZTF DR19), the light curves are analysed for about 200 WDMS binaries with emission line(s) identified from the Sloan Digital Sky Survey (SDSS) or the Large Sky Area Multi-Object Fibre Spectroscopic Telescope (LAMOST) spectra, in which 55 WDMS binaries are found to exhibit variability in their luminosities with a short period and are thus short-period binaries (i.e. PCEBs). In addition, it is found that the orbital periods of these PCEBs locate in a range from 2.2643 to 81.1526 h. However, only six short-period PCEBs are newly discovered and the orbital periods of 19 PCEBs are improved in this work. Meanwhile, it is found that three objects are newly discovered eclipsing PCEBs, and a object (i.e. SDSS J1541) might be the short-period PCEB with a late M-type star or a brown dwarf companion based on the analysis of its spectral energy distribution. At last, the mechanism(s) being responsible for the emission features in the spectra of these PCEBs are discussed, the emission features arising in their optical spectra might be caused by the stellar activity or an irradiated component owing to a hot WD companion because most of them contain a WD with an effective temperature higher than $\sim$10 000 K.

Li, Lifang↗

Carbon Storage Site Mapping Inquiry Tool (MapIT)

To date, 48 projects, consisting of 139 wells, are currently under review with the Environmental Protection Agency’s (EPA) Underground Injection Control (UIC) Program for Class VI – wells used for geologic sequestration of carbon dioxide. The number of applications submitted is expected to increase in coming years with the increase of the 45Q tax credit available to projects that initiate construction prior to 2033. The amount of data collected to submit a Class VI permit is vast, and often disparate, coming from state, federal, and commercial entities, as well as field-specific data collected within an area of interest. When preparing for site selection and permitting, the initial aggregation of relevant public data can be time intensive. The Carbon Storage Site Mapping Inquiry tool (MapIT) was created to support and accelerate the discovery and accessibility of open-source data and information available across the USA. Data was aggregated and organized based on data types described within the EPA UIC Class VI permit documentation. The online tool enables users to explore hundreds of geospatial data layers and connect to additional external resources, leveraging API and REST services where possible to ensure updates to data in real time. MapIT enables users to explore state and federal data related to geologic, geophysical, structural, hydrologic, and contextual information. In addition to displaying spatial data and linking to external resources, MapIT leverages custom widgets to ensure that internal data and external data are discoverable and accessible. The widgets connect users to resources such as the USGS publications and the USGS Earthquake Catalog based on a user-defined location. This talk will describe data aggregation workflows, data types, data preparation, and tool development for MapIT. The Carbon Storage Site Mapping Inquiry Tool and underlying database are valuable, intuitive resources that empower government, academic, commercial and industry stakeholders to explore, analyze, and acquire carbon storage related data.

Morkner, Paige↗

Web-Based Tools for Data-Informed Remedy Optimization: Software Theory and User Guide

This report documents the development and application of two web-based decision-support tools for pump-and-treat (P&T) groundwater remediation systems: PTOLEMY (Pump-and-Treat Optimized Location Evaluation to Maximize Yields) and OPTIMA (Optimization for Pump-and-Treat Implementation, Management, & Assessment). These tools enhance remedy design and management by leveraging advanced computational methods – specifically deep learning and multi-objective optimization – within a user-friendly platform. By integrating data-driven models with established hydrogeological knowledge, PTOLEMY and OPTIMA enable more efficient evaluation of well placement and operational strategies, helping site managers balance multiple remediation objectives under complex conditions. Both tools are implemented as modules within the SOCRATES (Suite Of Comprehensive Rapid Analysis Tools for Environmental Sites) web platform, which provides data access, visualization, and analytics to support remedy optimization across sites in the U.S. Department of Energy Office of Environmental Management complex. PTOLEMY is a rapid screening module designed to identify promising locations for new extraction wells. It employs a multi-channel three-dimensional convolutional neural network (MC3D-CNN) trained on high-fidelity simulation data to predict the relative performance (in terms of contaminant mass recovery) of potential well sites. Through an interactive web interface, PTOLEMY visualizes the probability of high performance across a site, highlighting areas where an extraction well is likely to yield above-threshold contaminant removal over a multi-year period. PTOLEMY’s map-based displays and exportable results support transparent communication of screening analyses. By focusing attention on the most favorable candidate locations, the tool augments traditional engineering judgment and physics-based modeling, providing a data informed basis for subsequent detailed evaluations. OPTIMA is a multi objective optimization module designed to find wellfield layouts and operating schedules that meet various cleanup goals. It quickly evaluates thousands of candidate setups – combinations of well locations, timing, and rates – and returns a small set of best trade-off options for comparison. At its core, OPTIMA uses a U-Net-based surrogate model – a deep-learning emulator of a groundwater flow and transport simulator – to dramatically accelerate scenario evaluations. Coupling this fast surrogate with the NSGA-II (Non-dominated Sorting Genetic Algorithm II) evolutionary algorithm, OPTIMA explores a wide decision space of well locations and schedules to identify Pareto-optimal solutions that trade off key objectives (e.g., minimizing cleanup time, maximizing contaminant mass removal, and minimizing plume extent). The tool outputs a family of optimal configurations and visualizes their trade-offs (Pareto frontiers of cleanup metrics and maps of optimized well placements). Site managers can use these results to understand the range of viable strategies and to select candidate designs for more detailed verification. OPTIMA is currently under active development and not yet fully released; this guide provides early documentation to support planning and gather user feedback.

54 ENVIRONMENTAL SCIENCES↗

MegaWatt Mayhem: Grid Operator Challenges Center Loads

This report provides a summary of the challenges faced by United States electricity grid operators in accommodating and anticipating the rapid deployment of large loads, particularly data centers, based on academic literature and industry working groups. The report highlights the unique requirements and operational characteristics of data centers, which differ significantly from traditional industrial loads. Key issues addressed utility planning considerations, with emphasis on the implications for grid operators, impacts to normal operations for grid operators, reliability considerations during periods of grid stress, and resilience considerations for the changing operational paradigms based on data centers. Real-world examples are used to highlight these challenges and the changes that grid operators must address. The findings underscore the necessity for coordinated efforts and innovative solutions from both grid operators and regulatory bodies to ensure the stable integration of large loads into the grid. This report is the first in a series that will explore the challenges of data center deployments based on several key power system perspectives.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Development of the ARM Lagrangian Large-Scale Forcing Data (ARMLAGTRAJ) Value-Added Product Based on the lagtraj Framework

The Atmospheric Radiation Measurement (ARM) large-scale forcing data developed based on the constrained variational analysis (VARANAL) value-added product (VAP) (Zhang and Lin 1997, Zhang et al. 2001, Xie et al. 2004, Tang et al. 2019) has been widely used for single-column models (SCMs), cloud-resolving models (CRMs), and large-eddy simulation models (LESs) to understand and improve physical processes in models. Recently, the U.S. Department of Energy (DOE) ARM user facility conducted several major field campaigns using ship-based moving observational platforms. For example, the Marine ARM GPCI Investigation of Clouds (MAGIC) field campaign focused on the role of subtropical marine-boundary layer (MBL) clouds, and the Multidisciplinary Drifting Observatory for the Study of Arctic Climate (MOSAiC) field campaign aimed to improve understanding of the coupled climate systems in the Arctic. Observations from moving platforms are critical to provide a comprehensive characterization of coupled-system processes associated with all stages of the cloud and/or sea-ice life cycle. Traditional ARM large-scale forcing data have been developed at fixed locations. They need to be extended to include these moving platforms to address data needs for ship-based field campaigns or to support LES modeling in a Lagrangian framework. With these considerations in mind, we develop ARM-type Lagrangian large-scale forcing data sets based on the lagtraj framework (Boeing et al. 2020) with notable enhancements in generating forcings that are more suitable for ARM field campaigns. The lagtraj is a novel tool that generates forcings for LES and SCM simulation in both Lagrangian and Eulerian perspective. This technical report focuses on the major changes we performed on the lagtraj algorithm and provides an overview of the ARM Lagrangian Large-Scale Forcing Data (ARMLAGTRAJ) value-added products.

54 ENVIRONMENTAL SCIENCES↗

Data-Driven Method for Groundwater-Level Mapping and Monitoring-Well Network Optimization at Hanford

This report summarizes the initial results and outcomes of a physics-informed, data-driven groundwater level (GWL) mapping capability for the Hanford Site. GWL mapping at Hanford is typically conducted annually and requires a significant amount of computational and expert resources, and it does not allow assessment of the informational value of specific monitoring wells. The proposed method produces spatially and temporally resolved fields consistent with sparse, irregularly sampled, and nonuniformly distributed well measurements. Implemented successfully, this capability will allow rapid mapping of groundwater levels and provide an opportunity to optimize monitoring activities (both location and sampling frequency) based on data information value evaluation. The approach integrates a diffusion-based generative model – trained on MODFLOW simulation data from the Plateau-to-River (P2R) model – with score-based data assimilation (SDA), allowing observation-conditioned mapping without retraining for each monitoring-network layout.

54 ENVIRONMENTAL SCIENCES↗

A Feasibility Study on the Integration of Human Performance Data From Diverse Sources Based on the Complexity of a Proceduralized Task

Securing the safety of socio-technical systems including nuclear facilities is the upmost goal to ensure their sustainability because historical records demonstrate that the performance degradation of human operators (e.g., human errors) is one of the crucial contributors to the occurrence of unexpected events resulting in extensive casualties and financial losses. This implies that the collection of human performance data in diverse conditions with which they could be faced during the operation of nuclear facilities. As this collection requires significant resources, it is necessary to resolve how to accomplish it with limited resources. To address this challenge, as suggested in the SHEEP framework, it is indispensable to extract valuable insights after integrating various kinds of human performance data obtained from different sources. However, a practical method to soundly integrate them seems to be still incomplete. Accordingly, the applicability of TACOM (Task Complexity) measure is investigated as a tool to identify useful information based on the integration of human performance data observed from different simulation conditions. As a result, it is expected that the TACOM measure would play an important role in addressing the technical challenge in securing human performance data.

99 GENERAL AND MISCELLANEOUS↗

Phosphoproteomics Modifications in Women with Rheumatoid Arthritis─Application of Web-Based Software to Enhance Data Visualization

Individuals with rheumatoid arthritis (RA) are at increased risk of functional disability, cardiovascular disease, and obesity, all of which are influenced by dysregulated skeletal muscle. Here, this pilot study aims to identify phosphoproteomics changes in RA skeletal muscle and visualize modifications through development of a web-based app designed to promote user-friendly data interpretation and visualization. NanoLC–MS/MS analysis was performed on vastus lateralis biopsies from three women with RA and matched healthy controls. Differential analysis was performed using the Limma R package. Kinase substrate enrichment analysis (KSEA) predicted changes in kinase activity. RA muscle displayed 35 upregulated and 60 downregulated phosphosites, including the cytoskeletal proteins TTN (Ser33201, Ser33013, Ser20925), NEB (Ser2219, Thr254, Ser33013, Ser20925), FLNA (Ser1459), and LASP1 (Ser146). Compared to healthy controls, KSEA predicted decreased activity of several kinases in RA muscle, including PRKACA and CDKs. All such changes were visualized by use of our web-based app. Overall, phosphoproteome analysis reveals signaling alterations in RA skeletal muscle linked to cytoskeletal proteins, representing candidate disease biomarkers; these modifications can be explored through use of our web-based software.

phosphoproteomics↗