Search NASA⌕ Search

SEARCH · Search NASA

Results for “data provenance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Hardware-in-the-Loop Testing of Wide-Area Damping Controller for Field Implementation in Large-scale Power Grid

In our previous work, an adaptive measurement-driven wide-area damping controller (WADC) for suppressing inter-area oscillations has been proposed and a hardware prototype was developed and validated through hardware-in-the-loop tests. As a continuation of the work, this paper introduces a WADC software prototype to handle the realistic challenges for field implementation in the control room of the power grid. The WADC software is developed and operated as an openPDC adapter with a graphical user interface (GUI) to monitor the WADC inputs and output, the communication delays and other variables. The software prototype has been fully tested through an enhanced hardware-in-the-loop (HIL) test setup. Its performance is verified under various realistic communication uncertainties, such as random time delays and data losses, with different communication protocols. The experiment results have proven the WADC software can deliver sufficient damping to suppress the targeted oscillation mode in handling various communication uncertainties for future field deployment.

Jia, Xinlan↗

Hutchinson Trace Estimation for high-dimensional and high-order Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) have proven effective in solving partial differential equations (PDEs), especially when some data are available by seamlessly blending data and physics. However, extending PINNs to high-dimensional and even high-order PDEs encounters significant challenges due to the computational cost associated with automatic differentiation in the residual loss function calculation. Herein, we address the limitations of PINNs in handling high-dimensional and high-order PDEs by introducing the Hutchinson Trace Estimation (HTE) method. Starting with the second-order high-dimensional PDEs, which are ubiquitous in scientific computing, HTE is applied to transform the calculation of the entire Hessian matrix into a Hessian vector product (HVP). This approach not only alleviates the computational bottleneck via Taylor-mode automatic differentiation but also significantly reduces memory consumption from the Hessian matrix to an HVP’s scalar output. We further showcase HTE’s convergence to the original PINN loss and its unbiased behavior under specific conditions. Comparisons with the Stochastic Dimension Gradient Descent (SDGD) highlight the distinct advantages of HTE, particularly in scenarios with significant variability and variance among dimensions. We further extend the application of HTE to higher-order and higher-dimensional PDEs, specifically addressing the biharmonic equation. By employing tensor-vector products (TVP), HTE efficiently computes the colossal tensor associated with the fourth-order high-dimensional biharmonic equation, saving memory and enabling rapid computation. The effectiveness of HTE is illustrated through experimental setups, demonstrating comparable convergence rates with SDGD under memory and speed constraints. Additionally, HTE proves valuable in accelerating the Gradient-Enhanced PINN (gPINN) version as well as the Biharmonic equation. Overall, HTE opens up a new capability in scientific machine learning for tackling high-order and high-dimensional PDEs.

Curse of dimensionality↗

ATcT — Active Thermochemical Tables Python Interface

SF-25-140 atct is a lightweight, Python client for the ATcT v1 API that enables programmatic access to high-accuracy thermochemical data and turnkey reaction-enthalpy analysis. The package implements full v1 endpoint coverage (species lookup by ATcT ID, name, formula, SMILES, InChI, CAS RN; covariance queries; health checks) with robust error handling, retries, and environment-based configuration for local/production endpoints. Beyond data retrieval, atct provides rigorously implemented reaction calculators that propagate uncertainties via either (i) a conventional independent-errors method (0 K or 298.15 K) or (ii) covariance-aware propagation using provided covariances at 298.15 K. Typed data classes ensure transparent, reproducible data structures and carry ATcT Thermochemical Network (TN) version identifiers for provenance. Dual import paths and comprehensive examples facilitate integration into research pipelines, enabling reproducible thermochemical calculations, automated validation, and downstream method development.

Bross, DavidHamilton [Argonne National Laboratory ↗

Optimizing Facility Operations by Applying Machine Learning to the Army Reserve Enterprise Building Control System (Final Report)

Thousands of U.S. Department of Defense (DoD) buildings have building automation systems (BASs) and/or advanced meters. Although these systems have a wealth of data, performance optimization requires time and expertise to review and act on that information. Machine learning (ML) can provide automated and actionable insights to controls operators. This demonstration implemented proven ML methods on the Army Reserve Enterprise Building Control System. ML refers to algorithms that “learn” from data and improve their performance on a given task over time. In the buildings domain these tasks range from predicting future energy consumption, to identifying operational issues before faults occur, to optimizing control decisions. To learn, ML requires input data, which – for buildings – typically consists of instrument data such as energy consumption data and subsystem controls information such as set-point temperatures, and context data consisting of information such as the physical location of the building, the area of the building, and the weather. ML models use the relationships learned from the input data to make predictions with new, previously unseen, data. The team was able to investigate and successfully implement the following ML use cases: labeling consumption data as anomalous or non-anomalous; baseline whole-building load prediction (unknown fault status); fault detection (validation not possible); and site prioritization for energy-related projects. Due to the constraints of the project, interventions were not able to be implemented during the demonstration; therefore, assessments of operational cost savings and maintenance avoided could not be performed. The project has been presented at two leading national building conferences and two additional publications to peer-reviewed journals are currently in preparation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Predicting High‐Resolution Spatial and Spectral Features in Mass Spectrometry Imaging with Machine Learning and Multimodal Data Fusion

Recent advancements in molecular Mass Spectrometry Imaging have sparked interest in integrating high spatial resolution methods with molecular mass-spectrometry-based chemical imaging. Fusion-based algorithms have proven effective in generating high spatial-resolution molecular mass spectra. However, a significant challenge stems from the differing physical mechanisms underlying image generation and data upsampling techniques, potentially leading to discrepancies in integrated information channels. Integrating physical constraints into data processing workflows is essential to tackle this issue. In this study, we propose an innovative approach that merges data from Fourier transform ion cyclotron resonance (FTICR), time-of-flight matrix-assisted laser desorption/ionization, and time-of-flight secondary ion mass spectrometry imaging techniques. By leveraging FT-ICR's unparalleled spectral resolution and ToF-SIMS's exceptional spatial resolution, we achieve submicron spatial resolution, enabling the observation of intact molecular species with remarkable spectral precision. Canonical correlation analysis is employed to incorporate physical constraints. Through sophisticated image processing and machine learning techniques, the results of this fusion hold significant promise for advancing our comprehension of complex systems and unveiling concealed molecular intricacies.

canonical correlation analysis↗

Irradiation Testing of Ultrasonic Transducers

Ultrasonic technologies offer the potential for high accuracy and resolution in-pile measurement of numerous parameters, including geometry changes, temperature, crack initiation and growth, gas pressure and composition, and microstructural changes. Many Department of Energy-Office of Nuclear Energy (DOE-NE) programs are exploring the use of ultrasonic technologies to provide enhanced sensors for in-pile instrumentation during irradiation testing. For example, the ability of single, small diameter ultrasonic thermometers (UTs) to provide a temperature profile in candidate metallic and oxide fuel would provide much needed data for validating new fuel performance models. Other efforts include an ultrasonic technique to detect morphology changes (such as crack initiation and growth) and acoustic techniques to evaluate fission gas composition and pressure. These efforts are limited by the lack of existing knowledge of ultrasonic transducer material survivability under irradiation conditions. To address this need, the Pennsylvania State University (PSU) was awarded an Advanced Test Reactor National Scientific User Facility (ATR NSUF) project to evaluate promising magnetostrictive and piezoelectric transducer performance in the Massachusetts Institute of Technology Research Reactor (MITR) up to a fast fluence of at least 1021 n/cm2 (E> 0.1 MeV). This test will be an instrumented lead test; and real-time transducer performance data will be collected along with temperature and neutron and gamma flux data. By characterizing magnetostrictive and piezoelectric transducer survivability during irradiation, test results will enable the development of novel radiation tolerant ultrasonic sensors for use in Material and Test Reactors (MTRs). The current work bridges the gap between proven out-of-pile ultrasonic techniques and in-pile deployment of ultrasonic sensors by acquiring the data necessary to demonstrate the performance of ultrasonic transducers

Daw, J.↗

Announcing the Biomedical Data Translator: Initial Public Release

ABSTRACT The growing availability of biomedical data offers vast potential to improve human health, but the complexity and lack of integration of these datasets often limit their utility. To address this, the Biomedical Data Translator Consortium has developed an open‐source knowledge graph–based system—Translator—designed to integrate, harmonize, and make inferences over diverse biomedical data sources. We announce here Translator's initial public release and provide an overview of its architecture, standards, user interface, and core features. Translator employs a scalable, federated, knowledge graph framework for the integration of clinical, genomic, pharmacological, and other biomedical knowledge sources, enabling query retrieval, inference, and hypothesis generation. Translator's user interface is designed to support the exploration of knowledge relationships and the generation of insights, without requiring deep technical expertise and gradually revealing more detailed evidence, provenance, and confidence information, as needed by a given user. To demonstrate Translator's application and impact, we highlight features of the user interface in the context of three real‐world use cases: suggesting potential therapeutics for patients with rare disease; explaining the mechanism of action of a pipeline drug; and screening and validating drug candidates in a model organism. We discuss strengths and limitations of reasoning within a largely federated system and the need for rich concept modeling and deep provenance tracking. Finally, we outline future directions for enhancing Translator's functionality and expanding its data sources. Translator represents a significant step forward in making complex biomedical knowledge more accessible and actionable, aiming to accelerate translational research and improve patient care.

Research & Experimental Medicine↗

Proton beam power limits for stationary water-cooled tungsten target with different cladding materials

The proton beam power limit for a solid-tungsten spallation target is largely determined by beam induced thermomechanical structural loads and decay heat power deposition, while its lifetime is limited by radiation damage and fatigue life of the target materials. In this paper, we studied the power limits of a stationary water-cooled solid tungsten target concept. Tantalum clad tungsten was considered as a reference case. Being a low activation material, zircaloy 2 cladding option was studied and its decay heat driven power limit was compared with the reference case. Zirconium alloys have proven operations records in spallation target and nuclear fission environments, supported by materials data obtained from post irradiation examinations. Recent study also demonstrated feasibility of diffusion bonding zirconium to tungsten using vanadium foil inter layer. Particle transport simulations code FLUKA was used to calculate energy deposition and decay heat power deposition in the target, based on the beam parameters technically feasible at the Second Target Station of the Spallation Neutron Source at Oak Ridge National Laboratory. The energy deposition data were used for flow, thermal, and structural analyses to determine the beam intensity limit on the target concept studied. The decay heat deposition data were used to calculate the transient temperature evolution in the tungsten volumes in a loss of coolant accident (LOCA) scenario to determine its beam power limit. For a 1.3 GeV proton beam, the power limit on a stationary target was 400 kW for a tantalum clad target model and 800 kW for a zircaloy 2 clad target model.

Lee, Yong Joong↗

Stability and Convergence of Solutions to Stochastic Inverse Problems Using Approximate Probability Densities

Data-consistent inversion is designed to solve a class of stochastic inverse problems where the solution is a pullback of a probability measure specified on the outputs of a quantities of interest (QoI) map. Here, this work presents stability and convergence results for the case where finite QoI data result in an approximation of the solution as a density. Given their popularity in the literature, separate results are proven for three different approaches to measuring discrepancies between probability measures: f-divergences, integral probability metrics, and L p metrics. In the context of integral probability metrics, we also introduce a pullback probability metric that is well-suited for data-consistent inversion. This fills a theoretical gap in the convergence and stability results for data-consistent inversion that have mostly focused on convergence of solutions associated with approximate maps. Numerical results are included to illustrate key theoretical results with intuitive and reproducible test problems that include a demonstration of convergence in the measure-theoretic "almost" sense.

97 MATHEMATICS AND COMPUTING↗

The VIBES Are Shifting: Assessing Emergent Capabilities in Multi-Modal Models

Researchers assessing open-source domains such as the internet, and particularly those studying information conflict, often have no single prescribed workflow. In the course of their research, they may need to perform a diverse array of tasks far beyond simply identifying an ever-changing set of objects. These analytical tasks can include ascertaining the provenance of images, understand an image in the context of accompanying text-based data, or identifying indicators of digital image manipulation. They must further be able to do this at the scale of tens of thousands of images or more. Traditional machine vision models have typically lacked the flexibility and breadth of performance sufficient for these needs. The research team from Pacific Northwest National Laboratory assessed the performance of a single baseline CLIP ViT-L model against a series of analytical tasks relevant for the study of online information conflict.

97 MATHEMATICS AND COMPUTING↗

yProv4ML: Effortless provenance tracking for machine learning systems

The rapid growth in interest in deep learning and foundation models (FMs) in particular, has attracted the attention of a diverse range of researchers thanks to their generalization ability. However, the advent of these techniques has also brought to light the lack of transparency and rigor in the way development is pursued. In particular, the inability to determine the number of epochs and other hyperparameters in advance presents challenges in identifying the best model. To address this challenge, machine learning frameworks such as MLFlow can automate the collection of this type of information. However, these tools capture data using proprietary formats and pose little attention to lineage. This paper proposes yProv4ML, a framework that captures provenance information generated during machine learning processes in PROV-JSON format, with minimal code modification.

Machine learning↗

Physics-Informed Neural Network (PINN) Prediction of Mixed Mass-Heat-Crystallization Limited Methane Hydrate Formation and Dissociation in Micro-Confinement

The creation and use of Physics-Informed Neural Networks (PINNs) for simulating the dynamics of methane hydrate formation and dissociation will be presented. The PINN framework's main benefit is its capacity to impose physical consistency with only a partial comprehension of the governing equations. This makes the algorithm especially useful for systems with little experimental evidence or a lack of theoretical knowledge. A strong basis for forecasting methane hydrate behavior over the verified operating ranges of 30.0-80.9 bar pressure and 1.0-4.0 K sub-cooling conditions is provided by the combination of conductive heat transfer equations and mixed mass-transfer–crystallization kinetics. PINNs were more accurate at predicting the mixed mass-heat-crystallization limited kinetics than conventional Artificial Neural Networks (ANNs), demonstrating remarkable predictive accuracy for methane hydrate production over the ANN model. The efficiency of incorporating physical limitations from first principles into machine learning frameworks for methane hydrate crystallizations is reinforced by these findings. For hydrate-related applications in energy generation, carbon sequestration, and climate modelling, our study establishes PINNs as a computational tool that is both scalable and efficient. The proven capacity to close the gap between conventional physics-based simulations and solely data-driven models creates new opportunities for expedited hydrate research and practical applications.

Hartman, Ryan L [NYU Tandon School of Engineering]↗

Generalization of Deep-Learning Models for Classification of Local Distance Earthquakes and Explosions across Various Geologic Settings

Although accurately classifying signals from earthquakes and explosions at local distance (<250 km) remains an important task for seismic network operations, the growing volume of available seismic data presents a challenge for analysts using traditional source discrimination techniques. In recent years, deep-learning models have proven effective at discriminating between low-magnitude earthquakes and explosions measured at local distances, but it is not clear how well these models are capable of generalizing across different geological settings. To address the issue of generalization between regions, we train deep-learning models (convolutional neural networks [CNNs]) on time–frequency representations (scalograms) of three-component earthquake and explosion signals from eight different regions in the continental United States. We explore scenarios where models are trained on data from all regions, individual regions, or all but one region. We find that although CNN models trained on individual regions do not necessarily generalize well across different settings, models trained on multiple regions that include diverse path coverage generalize to new regions, with station-level accuracy of up to 90% or more for data sets from unseen regions. In general, CNN-based discrimination models significantly outperform models based on uncorrected P/S ratio (measured in the 10–18 Hz frequency band), even when CNN models are tested on data from entirely unseen regions.

58 GEOSCIENCES↗

Polyconvex neural network models of thermoelasticity

Machine-learning function representations such as neural networks have proven to be excellent constructs for constitutive modeling due to their flexibility to represent highly nonlinear data and their ability to incorporate constitutive constraints, which also allows them to generalize well to unseen data. Here, in this work, we extend a polyconvex hyperelastic neural network framework to (isotropic) thermo-hyperelasticity by specifying the thermodynamic and material theoretic requirements for an expansion of the Helmholtz free energy expressed in terms of deformation invariants and temperature. Different formulations which a priori ensure polyconvexity with respect to deformation and concavity with respect to temperature are proposed and discussed. The physics-augmented neural networks are furthermore calibrated with a recently proposed sparsification algorithm that not only aims to fit the training data but also penalizes the number of active parameters, which prevents overfitting in the low data regime and promotes generalization. The performance of the proposed framework is demonstrated on synthetic data, which illustrate the expected thermomechanical phenomena, and existing temperature-dependent uniaxial tension and tension-torsion experimental datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Transit Rider/Travel Behavior Inventory Survey - Minneapolis-St. Paul Metro - 1990

The 1990 transit on-board survey aimed to update the 1988 survey, which was conducted as part of the Preliminary Engineering Study for the Hennepin County Light Rail Transit System. The 1988 study received Regional Transit Board funding, and survey results could be applied to a mode split model for projecting ridership on the proposed Light Rail Transit System. However, this survey was designed with the 1990 Travel Behavior Inventory in mind. The intention had been to update the 1988 survey in 1990 to be compatible with the 1990 Travel Behavior Inventory data. The results of the 1990 survey were used primarily to create a table of observed transit trips between each of the 1,200 traffic analysis zones in the region. This trip table was used to calibrate a new mode split model, which estimated future year travel by mode. The 1990 update survey focused on new routes and routes that had changed significantly since 1988. To preserve compatibility with the 1988 survey, the same survey questionnaire card was used, together with the same survey procedures for data collection. The procedures randomly sampled bus patrons during the transit trip, asking key questions about the patron and the transit trip. The survey card was intended for patrons to fill out quickly so it could be completed during the transit trip. The questions focused on conditions that have proven over time to significantly influence ridership. In all, a total of 20,126 valid survey records were processed. Adding surrogate trips, the survey data file is composed of a total of 27,159 trip records . About 10% of the records were filled out by persons who had answered more than one questionnaire.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Accelerating Control Systems with GitOps: A Path to Automation and Reliability

GitOps is a foundational approach for modernizing infrastructure by leveraging Git as the single source of truth for declarative configurations. The poster explores how GitOps transforms traditional control system infrastructure, services and applications by enabling fully automated, auditable, and version-controlled infrastructure management. Cloud-native and containerized environments are shifting the ecosystem not only in the IT industry but also within the computational science field, as is the case of CERN and Diamond Light Source among other Accelerator/Science facilities which are slowly shifting towards modern software and infrastructure paradigms. The ACORN project, which aims to modernize Fermilab’s control system infrastructure and software is implementing proven best-practices and cutting-edge technology standards including GitOps, containerization, infrastructure as code and modern data pipelines for control system data acquisition and the inclusion of AI/ML in our accelerator complex.

Gonzalez, M. [Fermilab]↗

Identifying Sample Provenance From SEM/EDS Automated Particle Analysis via Few-Shot Learning Coupled With Similarity Graph Clustering

Automated particle analysis (APA) provides a vast amount of compositional data via energy-dispersive X-ray spectroscopy along with size and shape data via scanning electron microscopy for individual particles in a sample. In many instances, APA data are leveraged to support identification of the source of a sample based on the detection of particles of a specific composition. Often, the particles that provide context make up a minuscule portion of the sample. Additionally, the interpretation of complex samples can be difficult due to the diversity of compositions both in the mixture and within a particle. In this work, we demonstrate a method to compute and cluster similarity graphs that describe inter-particle relationships within a sample using a multi-modal few-shot learning neural network. Here, as a proof-of-concept, we show that samples known to have been exposed to gunshot residue can be distinguished from samples occasionally mistaken for gunshot residue. Our workflow builds upon standard APA techniques and data processing methods to unveil additional information in a readily interpretable and quantitatively comparable format.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Assessment of Accelerated Stress Testing Data for Silicon Photovoltaics Using Tensor Decomposition Methods

The photovoltaic (PV) industry is simultaneously targeting long warranties and new materials/designs for high-energy-yield modules, requiring an advanced methodology to forecast long-term durability of products with un-proven materials combinations. Extended, sequential, and combined stress testing methods are gaining popularity for assessing durability of PV modules/materials beyond the early-stage mortalities. Importantly, multiple degradation mechanisms can proceed simultaneously, and their separate contributions to the overall power loss should ideally be quantified. This work examines the use of data-driven tools towards developing a strategy for faster learning cycles in accelerated stress testing.

accelerated stress testing↗