Search NASASearch

SEARCH · Search NASA

Results for “data augmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A Novel Segmentation Algorithm for the ARM User Facility All-Sky Imagers Using Machine Learning Applications

Cloud cover plays a pivotal role in modulating the Earth's energy budget through the reflection of incoming solar radiation and the trapping of outgoing longwave radiation. Ground-based all-sky imagers offer an objective assessment of cloud cover that can be used to estimate solar irradiance, classify cloud types, track cloud movement, and serve as a benchmark 10 for the evaluation of satellite and reanalysis data products. The Atmospheric Radiation Measurement (ARM) user facility has utilized all-sky imagers for more than 25 years to monitor cloud cover and augment its comprehensive suite of atmospheric measurements. Following the retirement of its Total Sky Imager (TSI), ARM recently deployed the TSI’s successor, the All Sky Imager (ASI-16 camera systems). To provide a smooth transition and continuity to the vast amount of knowledge gathered by the TSI over the years, while addressing typical deployment issues, we developed a novel pixel segmentation algorithm, 15 the ASI Sky Cover (ASISKYCOVER). ASISKYCOVER builds on the different strengths and properties of the TSI processing algorithm while integrating machine learning techniques, ensuring data validity and accuracy across diverse atmospheric conditions. It enhances cloud cover characterization with new features such as artifact detection and uncertainty quantification. ASISKYCOVER also includes cloud cover estimates for near-zenith (narrow field-of-view) and reduces susceptibility to false detections. This study introduces ASISKYCOVER, details its algorithm framework, and demonstrates its capabilities using a 20 year-long dataset from the ARM Southern Great Plains site. Comparisons with co-located TSI data and other ARM measurements, such as zenith-pointing radars and lidars, are presented, underscoring the ASISKYCOVER’s potential to improve cloud cover analyses and data evaluation efforts, as well as to be integrated into higher-level data products that synergize instrument suites to generate new and insightful information

Silber, Israel

Hydropower potential derived from streamflow extremes for Alaska, USA

Alaska is an expansive region known for its abundant natural resources, including thousands of miles of streams and rivers. These rivers represent potential opportunities for future hydropower development that could provide reliable energy supply for local communities. There is limited long-term high temporal resolution streamflow data available for the region, making data-driven estimates of potential hydropower and its variability across the state challenging. This study provides a novel data-driven approach for hydropower capacity estimation across Alaska. We use supervised machine learning to develop a relationship between the daily and peak flow duration curves in order to augment the size of our dataset from 44 sites to 67 sites. We perform a stochastic hydropower estimation across the 67 sites and identify approximately 1000 MW of total potential hydropower capacity distributed across these sites. Our study provides the first step towards more comprehensive hydropower estimation for this critical region, highlighting the need for future work integrating high-resolution spatial data, community needs, and economic constraints in estimates of potential hydropower development in Alaska.

Hydropower

Simplified Model and Approach to Transform Infrared Surface Temperature to Film Effectiveness in a Conjugate Heat Transfer Experiment

This paper describes a simplified engineering model based on a one-dimensional thermal resistance network. The model is used to develop a new method to relate film cooling effectiveness and heat transfer augmentation to local overall cooling effectiveness in a conjugate flat plate experiment. This paper presents experimental proof-of-concept data to demonstrate the potential for this model. In contrast to previous approaches, neither the wall heat flux nor the adiabatic wall temperature is required to estimate the local film cooling performance parameters. The model predicts surface temperatures that are within the experimental uncertainties over the range for which the model is trained, and to within five percent when the model is extrapolated to higher coolant channel Reynolds numbers. This paper is relevant to conjugate test rigs that can measure the hot surface temperature distribution with and without film cooling. This information may also be relevant to designers as a method to approximate surface temperatures or used as an approximate heat transfer model for optimization studies.

advanced gas turbines

Toward real-time optimization through model reduction and model discrepancy sensitivities

Optimization problems arise in a range of scenarios, from optimal control to model parameter estimation. In many applications, such as the development of digital twins, it is essential to solve these optimization problems within wall-clock-time limitations. However, this is often unattainable for complex systems, such as those modeled by nonlinear partial differential equations. One strategy for mitigating this issue is to construct a reduced-order model (ROM) that enables more rapid optimization. In particular, the use of nonintrusive ROMs—those that do not require access to the full-order model at evaluation time—is popular because they facilitate the computation of optimization solutions within the wall-clock time requirements. However, the optimization solution will be unreliable if the iterates move outside the ROM training data. This article proposes the use of hyper-differential sensitivity analysis with respect to model discrepancy (HDSA-MD) as a computationally efficient tool to augment ROM-constrained optimization and improve its reliability. The proposed approach consists of two phases: (i) an offline phase where several full-order model evaluations are computed to train the ROM, and (ii) an online phase where a ROM-constrained optimization problem is solved, a limited number of full-order model evaluations are computed, and HDSA-MD is used to enhance the optimization solution. Numerical results are demonstrated for two examples, atmospheric contaminant control and wildfire ignition location estimation, in which a ROM is trained offline using inaccurate atmospheric data. In conclusion, the HDSA-MD update yields a significant improvement in the ROM-constrained optimization solution using only one full-order model evaluation online with corrected atmospheric data.

PDE-constrained optimization

Generative large language models for predictive maintenance planning

Maintenance planning and the generation of necessary components for tasks can prove time-consuming and complex. Automating the creation of recurring or similar tasks by leveraging previous planning packages and data, while uncovering insights to automate planning package generation, presents an opportunity to conserve valuable time and resources. This work aims to harness the textual and probabilistic capabilities of large language models (LLMs) to automate the generation of planning packages. Utilizing diverse data sources ranging from raw data to handwritten text, both singular and collaborative LLMs are trained and tested. Results demonstrate their capability to generate essential planning package components, effectively replicating the statistical patterns in the data. This demonstrates the use of these tools inside a digital asset for automated planning. This work outlines a methodology for constructing datasets, a training suite, and evaluation methods for LLM-based textual and conversational planning tools utilized in an asset digital twin. Results indicate that the fine-tuned models generate estimated planning information within the statistical ranges observed in real maintenance data. The models achieve high accuracy (>90%) in document question-answering and instruction generation tasks. Furthermore, the conversational retrieval-augmented generation (RAG) assistant system achieves 100% document retrieval accuracy, while conversational information capture exceeds 98% across the majority of work-package assistant modules.

97 MATHEMATICS AND COMPUTING

Simplified Model and Approach to Transform Infrared Surface Temperature to Film Effectiveness in a Conjugate Heat Transfer Experiment

In the pursuit of more efficient gas turbines, film cooling is a critical technology. This article describes a simplified engineering model based on a one-dimensional thermal resistance network. The model is used to relate film-cooling effectiveness and heat transfer augmentation to local overall cooling effectiveness in a conjugate flat plate experiment. Here, this article presents experimental proof-of-concept data to demonstrate the potential for this model. In contrast to previous approaches, neither the wall heat flux nor the adiabatic wall temperature is required to estimate the local film-cooling performance parameters. The model predicts surface temperatures that are within the experimental uncertainties over the range for which the model is trained and to within five percent when the model is extrapolated to higher coolant channel Reynolds numbers. This article is relevant to conjugate test rigs that can measure the hot-surface temperature distribution with and without film cooling. This information may also be relevant to designers as a method to approximate surface temperatures or used as an approximate heat transfer model for optimization studies.

advanced turbines

Collection and Analysis of Telemetry for CyOTE Heuristics (CATCH)

The Collection and Analysis of Telemetry for CyOTE Heuristics (CATCH) provides a framework for augmenting an organization’s existing security controls with CyOTE developed analyses. CATCH collects, stores, analyzes, and creates STIX reports on anomalous data. CATCH connects the CyOTE analysis framework together with the MITRE ICS ATT&CK® patterns and highlights areas of improvement and further research. This tool is designed to enhance an organization’s security controls by providing a structured approach to collecting, storing, analyzing, and reporting anomalous data.

99 GENERAL AND MISCELLANEOUS

Digitizing and Enhancing Accessibility of the Fusion Safety Archives

This project focuses on the digitization and public accessibility to the Fusion Safety Archives at the Idaho National Laboratory. The first phase involves a thorough review of each document in the physical archives to determine its online availability. For documents that are available online, PDF copies and unique identifiers are collected for database integration. Documents not available online are delivered to Red Inc. for digitization. Additionally, defunct storage devices such as diskettes are sent to INL’s archival department for data retrieval where possible. The second phase of the project involves the creation of a comprehensive database to house the digital copies of the archives. The database will facilitate easy access and management of the digitized documents. Following the database creation, we plan to train a Retrieval-Augmented Generation (RAG) based AI on publicly available documents. The trained AI will be integrated into a front-facing application, allowing the public to easily access information from the Fusion Safety Archives. This project aims to preserve valuable historical data, improve accessibility, and promote transparency in fusion safety research.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY

X-ray spectroscopy of multi-temperature plasmas using the differential emission measure formalism

We present a theoretical construct that nominally underlies spectroscopic data analysis of multi-temperature plasmas, known as the differential emission measure (DEM). From a data analytic perspective, the DEM formalism is used to derive temperature distributions from line spectra that are formed in the presence of temperature gradients and by time integrations of evolving plasmas. From a modeling perspective, DEMs are convenient intermediaries between radiation hydrodynamics simulations and spectroscopic measurements acquired in the laboratory. The DEM concept and its associated methodologies were originally developed by spectroscopists working with astrophysical data. We borrow from these earlier investigations. In this manuscript, intended primarily as a tutorial, we discuss the basic concepts, but also augment various aspects of the theory by the way of extension and example, including a detailed treatment of various weighting and averaging schemes, intended to mitigate ambiguities that often arise when reporting temperature information. We focus on high-temperature plasmas that are not in local thermodynamic equilibrium and the x-ray spectra that they produce, although the core ideas presented here are applicable to spectroscopy in other energy bands. A few examples involving the derivation and manipulation of model DEMs in simple geometries are provided.

Liedahl, Duane A. [Lawrence Livermore National Lab

Assessing the Expansion of Ground-Motion Sensing Capability in Smart Cities via Internet Fiber-Optic Infrastructure

Monitoring ground motion in smart cities can improve the public safety by providing critical insights on natural and anthropogenic hazards, for example, earthquakes, landslides, explosions, infrastructure failures, and so forth. Although seismic activity is typically measured using dedicated point sensors (e.g., geophones and accelerometers), techniques such as distributed acoustic sensing have demonstrated the utility of using fiber-optic cable to detect seismic activity over comparable distances. In this article, we present the results of a study that quantifies the expansion in an area monitored for low-amplitude ground-motion events by augmenting existing point sensors with the internet fiber-optic cable infrastructure. Here we begin by describing our methodology, which utilizes geospatial data on point sensors and internet optical fiber deployed in metropolitan statistical areas (MSAs) in the United States. We extend these data to identify the area that can be monitored by (1) considering the observed seismic noise data in target locations, (2) applying the model from Wilson et al. (2021) to understand the potential coverage area gains using optical fiber sensing, and (3) optimizing the selection of fiber segments to maximize coverage and minimize deployment costs. We implement our methodology in ArcGIS to assess the additional area that can be monitored for low-amplitude ground-motion events (i.e., magnitude >0.5) by utilizing internet fiber-optic cables in the 100 most populous MSAs in the United States. We find that the addition of internet fiber-based sensors in MSAs would increase the area monitored on average by over an order of magnitude from 1% to 12%, if the subset of fiber cable segments that maximize coverage and minimize deployment costs is chosen even if only 20% of all fibers are used.

58 GEOSCIENCES

Analog-to-digital converter based on voltage-controlled superconducting devices

The increasing demand for cryogenic electronics in superconducting and quantum computing systems calls for ultra-energy-efficient data conversion architectures that remain functional at deep cryogenic temperatures. Here, in this work, we present the first design of a voltage-controlled superconducting flash analog-to-digital converter (ADC) based on a voltage-controlled quantum-enhanced Josephson junction field-effect transistor (JJFET). Exploiting its strong gate tunability and transistor-like behavior, the JJFET offers a scalable alternative to conventional current-controlled superconducting devices while aligning naturally with CMOS-style design methodologies. Building on our previously developed Verilog-A compact model calibrated to experimental data, we design and simulate a three-bit JJFET-based flash ADC targeted for integration within cryogenic control and readout circuitry in quantum computing. The core comparator block is realized through careful bias current selection and augmented with a three-terminal nanocryotron to precisely define reference voltages. Cascaded JJFET comparators ensure robust voltage gain, cascadability, and logic-level restoration across stages. Simulation results demonstrate accurate quantization behavior with ultra-low power dissipation, underscoring the feasibility of voltage-driven superconducting mixed-signal circuits. This work establishes a critical step toward unifying superconducting logic and data conversion, paving the way for scalable cryogenic architectures in quantum–classical co-processors, low-power artificial intelligence accelerators, and next-generation energy-constrained computing platforms.

Analog-to-digital converter

REFSafE: A RAG-Enabled Framework for Predictive Risk Analysis and Automated Safety Report Generation in Mission-Critical Environments

Operational safety in mission-critical environments requires AI systems that are accurate, interpretable, and resistant to hallucination. We present an agentic Retrieval-Augmented Generation (RAG) framework, REFSafe, for grounded hazard analysis and automated safety report generation. The system integrates Large Language Models (LLMs) with structured operational data, historical incident repositories, policy documents, and external authoritative sources. Through iterative agentic reasoning, the framework retrieves, verifies, and synthesizes evidence prior to generation, enforcing citation-backed outputs with explicit source attribution (documents, links, and prior events) to ensure traceability and trust. To mitigate hallucinations and unsupported claims, all risk assessments and forecasts are constrained to retrieved evidence, with confidence signals derived from retrieval relevance and source consistency. A transparent pipeline enables subject matter experts (SMEs) to validate predictions, and provide structured feedback, forming a continuous performance calibration loop. Preliminary deployment demonstrates improved reliability in hazard detection and safety/vulnerability report generation. This work advances trustworthy, evidence-grounded AI for predictive safety intelligence in mission-critical operations.

Das, Sanjay [ORNL] (ORCID:0009000542591915)

Use of Frit‐Disc Crucible Sets to Make Solution Growth More Quantitative and Versatile

The recent availability of step‐edge, frit‐disc crucible sets (generally sold as Canfield Crucible Sets or CCS) has led to multiple innovations associated with the group's use of solution growth. The use of CCS allows for the clean separation of liquid from solid phases during the growth process. This clean separation enables the reuse of the decanted liquid, either allowing for simple, economic, savings associated with recycling expensive precursor elements or allowing for the fractionation of a growth into multiple, small steps, revealing the progression of multiple solidifications. Clean separation of liquid from solid phases also allows for the determination of the liquidus line (or surface) and the creation, or correction, of composition–temperature phase diagrams. The reuse of clean decanted liquid has also allowed to prepare liquids ideally suited for the growth of large single crystals of specific phases by tuning the composition of the melt to the optimal composition for growth of the desired phase, often with reduced nucleation sites. Finally, it is discussed how solution growth and CCS use can be harnessed to provide a plethora of composition–temperature data points defining liquidus lines or surfaces with differing degrees of precision to either test or anchor artificial intelligence and/or machine‐learning‐based attempts to augment and extend the limited experimentally determined database.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Computational toolkit for predicting thickness of 2D materials using machine learning and autogenerated dataset by large language model

The thickness of 2D materials not only plays a crucial role in determining the performance of nanoelectronic and optoelectronic devices but also introduces complexities in predicting volume-dependent properties, such as energy storage capacity, due to the intrinsic vacuum within these materials. Although a plethora of experimental techniques, including but not limited to optical contrast, Raman spectroscopy, nonlinear optical spectroscopy, near-field optical imaging, and hyperspectral imaging, facilitate the measurement of 2D material thickness, comprehensive data for many materials remain elusive. Over the past decade, the exponential proliferation of 2D materials and their heterostructures has outstripped the capabilities of conventional experimental and computational approaches. In this evolving landscape, machine learning (ML) has emerged as an indispensable tool, offering a scalable approach to augment these traditional methodologies. Addressing the critical gap, we introduce THICK2D—Thickness Hierarchy Inference and Calculation Kit for 2D Materials. This Python-based computational framework harnesses an autogenerated thickness database, developed using large language models, and advanced ML algorithms to facilitate the rapid and scalable estimation of material thickness, relying solely on crystallographic data. To demonstrate the utility and robustness of THICK2D, we successfully used the toolkit to predict the thickness of more than 8000 2D-based materials, sourced from two extensive 2D materials databases. THICK2D is disseminated as an open-source utility, accessible on GitHub at https://github.com/gmp007/THICK2D, and archived on Zenodo at https://10.5281/zenodo.11216648.

Ekuma, Chinedu E. (ORCID:0000000258527556)

Towards a RAG-based summarization for the Electron Ion Collider

Abstract The complexity and sheer volume of information — encompassing documents, papers, data, and other resources — from large-scale experiments demand significant time and effort to navigate, making the task of accessing and utilizing these varied forms of information daunting, particularly for new collaborators and early-career scientists.To tackle this issue, a Retrieval Augmented Generation (RAG)-based Summarization AI for EIC (RAGS4EIC) is under development. This AI-Agent not only condenses information but also effectively references relevant responses, offering substantial advantages for collaborators. Our project involves a two-step approach: first, querying a comprehensive vector database containing all pertinent experiment information; second, utilizing a Large Language Model (LLM) to generate concise summaries enriched with citations based on user queries and retrieved data. We describe the evaluation methods that use RAG assessments (RAGAs) scoring mechanisms to assess the effectiveness of responses. Furthermore, we describe the concept of prompt template based instruction-tuning which provides flexibility and accuracy in summarization. Importantly, the implementation relies on LangChain [1], which serves as the foundation of our entire workflow. This integration ensures efficiency and scalability, facilitating smooth deployment and accessibility for various user groups within the Electron Ion Collider (EIC) community. This innovative AI-driven framework not only simplifies the understanding of vast datasets but also encourages collaborative participation, thereby empowering researchers. As a demonstration, a web application has been developed to explain each stage of the RAG Agent development in detail. The application can be accessed athttps://rags4eic-ai4eic.streamlit.app.[A tagged version of the source code can be found inhttps://github.com/ai4eic/EIC-RAG-Project/releases/tag/AI4EIC2023_PROCEEDING.]

Instruments & Instrumentation

Immersive Analytics in Critical Spatial Domains: From Materials to Energy Systems

Immersive analytics (IA) leverages virtual reality, augmented reality, and mixed reality to transform how users interact with complex datasets across domains such as science, industry, and education. These immersive technologies offer spatial and multimodal environments that foster intuitive exploration, but they also introduce challenges related to cognitive load, interface design, and system performance. Here, this article presents a comprehensive review of visualization techniques, interaction models, and multimodal inputs utilized in IA. Drawing on case studies in scientific visualization, industrial training, and educational communication, we examine both the potential and limitations of current systems. Finally, we propose future research directions, focusing on real‐time collaboration, adaptive user interfaces, and scalable data exploration strategies to advance the field.

99 - GENERAL AND MISCELLANEOUS

How the Galaxy–Halo Connection Depends on Large-scale Environment

We investigate the connection between galaxies, dark matter halos, and their large-scale environments at z = 0 with Illustris TNG300 hydrodynamic simulation data. We predict stellar masses from subhalo properties to test two types of machine learning (ML) models: explainable boosting machines (EBMs) with simple galaxy environment features and E(3)-invariant graph neural networks (GNNs). The best-performing EBM models leverage spherically averaged overdensity features on 3 Mpc scales. Interpretations via SHapley Additive exPlanations also suggest that in the context of the TNG300 galaxy–halo connection, simple spherical overdensity on ∼3 Mpc scales is more important than cosmic web distance features measured using the DisPerSE algorithm. Meanwhile, a GNN with connectivity defined by a fixed linking length, L, outperforms the EBM models by a significant margin. As we increase the linking length scale, GNNs learn important environmental contributions up to the largest scales we probe (L = 10 Mpc). We conclude that 3 Mpc distance scales are most critical for describing the TNG galaxy–halo connection using the spherical overdensity parameterization, but that information on larger scales, which is not captured by simple environmental parameters or cosmic web features, can further augment these models. Our study highlights the benefits of using interpretable ML algorithms to explain models of astrophysical phenomena, and the power of using GNNs to flexibly learn complex relationships directly from data while imposing constraints from physical symmetries.

79 ASTRONOMY AND ASTROPHYSICS

Optimizing Porous Transport Layer Porosity for Proton Exchange Membrane Water Electrolysis

An empirical model is presented that describes anode-side losses related to porous transport layer (PTL) morphology in proton exchange membrane water electrolysis (PEMWE). The model is based on an advanced voltage breakdown analysis that links various overpotentials to PTL morphology. Custom Ti PTLs, spanning uncommonly low porosities (22 - 31%), were fabricated and analyzed with X-ray CT to obtain pore and particle size distributions. Particle size distributions were consistent across samples with an average particle diameter of 12.0?..mu..m, whereas average pore diameters ranged from 6.0 to 7.0?..mu..m. The PTLs were tested in standard PEMWE cell assemblies with anode catalyst loadings of 0.1 mgIr cm-2 to obtain polarization curves, electrochemical impedance spectra, and augmented Tafel analysis. The PTL-dependent anode side losses were deconvoluted and assigned to excess utilization, concentration, ion transport resistance, and electrical contact resistance overpotentials. The data and model reveal an optimal 20 - 28% PTL porosity region where utilization and contact resistance overpotentials are minimized without triggering concentration and ion transport losses related to water deprivation. The optimal PTL porosity depends on the operating current density and is demonstrated at realistic PEMWE water flow rates to establish PTL design guidance for operation at scale.

08 HYDROGEN