Search NASA⌕ Search

SEARCH · Search NASA

Results for “DATA BASES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Accelerating data acquisition with FPGA-based edge machine learning: a case study with LCLS-II

New scientific experiments and instruments generate vast amounts of data that need to be transferred for storage or further processing, often overwhelming traditional systems. Edge machine learning (EdgeML) addresses this challenge by integrating machine learning (ML) algorithms with edge computing, enabling real-time data processing directly at the point of data generation. EdgeML is particularly beneficial for environments where immediate decisions are required, or where bandwidth and storage are limited. In this paper, we demonstrate a high-speed configurable ML model in a fully customizable EdgeML system using a field programmable gate array (FPGA). Our demonstration focuses on an angular array of electron spectrometers, referred to as the ‘CookieBox,’ developed for the Linac Coherent Light Source II project. The EdgeML system captures 51.2 Gbps from a 6.4 GS s −1 analog to digital converter and is designed to integrate data pre-processing and ML inside an FPGA. Our implementation achieves an inference latency of 0.2 µs for the ML model, and a total latency of 0.4 µs for the complete EdgeML system, which includes pre-processing, data transmission, digitization, and ML inference. The modular design of the system allows it to be adapted for other instrumentation applications requiring low-latency data processing.

97 MATHEMATICS AND COMPUTING↗

Deliverable 6.7-Final Technical Report: Development Summary and Evaluation of the Solar Uncertainty Integrator (SUNI) Software

The Data Quality and Uncertainty Integration Project was a three-year effort to address stakeholder needs for assessing solar radiation resource data quality based on existing tools for estimating radiometer measurement uncertainties and assessing post-measurement data quality. The annual research objectives for the project addressed a logical progression of effort needed to achieve the ultimate project goal of developing the Solar Uncertainty Integrator (SUNI) software. This final technical report summarizes the development process for achieving these key research objectives and addresses the outreach and code development efforts in the final year of the project to develop a new solar irradiance data uncertainty integration software package.

14 SOLAR ENERGY↗

On the Abuse and Detection of Polyglot Files

A polyglot is a file that is valid in two or more formats. Polyglot files pose a problem for file-upload and generative AI web interfaces that rely on format identification to determine how to securely handle incoming files. In this work we found that existing file-format and embedded-file detection tools, even those developed specifically for polyglot files, fail to reliably detect polyglot files used in the wild. To address this issue, we studied the use of polyglot files by malicious actors in the wild, finding 30 polyglot samples and 15 attack chains that leveraged polyglot files. Using knowledge from our survey of polyglot usage in the wild---the first of its kind---we created a novel data set based on adversary techniques. We then trained a machine learning detection solution, PolyConv, using this data set. PolyConv achieves a precision-recall area-under-curve score of 0.999 with an F1 score of 99.20% for polyglot detection and 99.47% for file-format identification, significantly outperforming all other tools tested. We developed a content disarmament and reconstruction tool, ImSan, that successfully sanitized 100% of the tested image-based polyglots, which were the most common type found via the survey. Our work provides concrete tools and suggestions to enable defenders to better defend themselves against polyglot files, as well as directions for future work to create more robust file specifications and methods of disarmament.

Oesch, T [ORNL] (ORCID:0000000269091022)↗

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is continuous generation of an extremely large amount of equipment reliability (ER) data. These data elements come in textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) forms. They provide system engineers with valuable insights and information by discovering anomalous behaviors or degradation trends, identifying possible causes behind such behaviors and trends, and predicting their direct consequences. This paper directly targets the knowledge generation from ER data by putting “data into context.” We employ model-based system engineering (MBSE) of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by first identifying which of the developed MBSE elements they are referring to. This task is harder for textual data since the information contained in issue or maintenance reports needs to be “understood” by a computational tool. We called this process “knowledge extraction” since our methods extract knowledge from textual data. Last, once numeric and textual ER data elements have been processed and “understood,” we discover possible cause-effect relations among them. This is performed by observing whether a logical connection through the MBSE models exists, and if there is a temporal relationship among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 - MATHEMATICS AND COMPUTING↗

Design optimization of MAPS-based detectors using a data-driven fast simulation approach

A parametric simulation tool for pixel sensors is presented. A realistic pixel response is simulated purely based on measurement input, without requiring detailed knowledge of the underlying manufacturing process. As such, it provides an efficient alternative to the use of Technology Computer-Aided Design simulations, which typically depend on proprietary process information. Due to its parametric approach, the package is fast and thus particularly useful for larger detector systems and high hit rate environments. This work presents measurements, simulation and its validation for the MALTA2 sensor. It is a small collection electrode monolithic active pixel sensor produced in the Tower 180 nm complementary metal-oxide-semiconductor imaging process. Modifications to the sensor’s periphery, mainly in the hit merger, are studied in order to optimize the performance for tracking and calorimetry. This optimization is of special interest as part of the MALTA3 sensor redesign in the 65 nm Tower Partners Semiconductor Co. process.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗

User Guide for Sample Reduction at GP-SANS

This manual is intended as a quick guide for data reduction of GP-SANS data. It includes all necessary steps to do the data reduction based on absolute calibration using the open beam method and how to transfer the reduced data to the personal computer system. If any errors are coming up so that the reduction script is not functioning as intended, please contact the instrument scientist.

97 MATHEMATICS AND COMPUTING↗

Constraints on fifth forces and ultralight dark matter from OSIRIS-REx target asteroid Bennu

It is important to test the possible existence of fifth forces, as ultralight bosons that would mediate these are predicted to exist in several well-motivated extensions of the Standard Model. Recent work indicated asteroids as promising probes, but applications to real data are lacking so far. Here we use the OSIRIS-REx mission and ground-based tracking data for the asteroid Bennu to derive constraints on fifth forces. Our limits are strongest for mediator masses m ~ (10 -18 -10 -17 ) eV, where we currently achieve the tightest bounds. These can be translated to a wide class of models leading to Yukawa-type fifth forces, and we demonstrate how they apply to U(1) B dark photons and baryon-coupled scalars. Our results demonstrate the potential of asteroid tracking in probing well-motivated extensions of the Standard Model and ultralight bosons near the fuzzy dark matter range.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Towards Resilient Near Real-Time Analysis Workflows in Fusion Energy Science

Nuclear fusion holds the promise of an endless source of energy. Several research experiments across the world and joint modeling and simulation efforts between the nuclear physics and high performance computing communities are actively preparing the operation of the International Thermonuclear Experimental Reactor (ITER). Both experimental reactors and their simulated counterparts generate data that must be analyzed quickly and in a resilient way to support decision making for the configuration of subsequent runs or prevent a catastrophic failure. However, the cost if the traditional techniques used to improve the resilience of analysis workflows, i.e., replicating datasets and computational tasks, becomes prohibitive with explosion of the volume of data produced by modern instruments and simulations. Therefore, we advocate in this paper for an alternate approach based on data reduction and data streaming. The rationale is that by allowing for a reasonable, controlled, and guaranteed loss of accuracy it becomes possible to transfer smaller amounts of data, shorten the execution time of analysis workflows, and lower the cost of replication to increase resilience. We develop our research and development roadmap towards resilient near real-time analysis workflows in fusion energy science and present early results showing that data streaming and data reduction is a promising way to speed up the execution and improve the resilience of analysis workflows.

Suter, Fred↗

Does class matter? Understanding differential pandemic recovery via a building typology

This study investigates the recovery of building-level footfall from the COVID-19 pandemic using privacy-preserving mobile devices-based footfall data within 60 downtown areas in the USA and Canada. Using clustering, we identify five distinct building typologies based on their characteristics, including rent, quality and recovery rates. The results reveal significant variation of recovery rates by building features. We find negative relationships with footfall recovery for both the percentage of office and remote work tenants and building quality. In contrast, buildings with traditional work tenants and retail functions achieve higher recovery rates. We also test the ‘flight to quality’ hypothesis via on our typology results. High-quality office buildings (Class A+) continue to have high rents but experience low physical footfall recovery, which suggests that this class is not as resilient as portrayed. The findings thus suggest the importance of considering both economic and footfall resilience in evaluating the performance of office buildings.

Covid-19↗

Anomalous Spin Precession Frequency Analysis in the Muon $g-2$ Experiment at Fermilab

The Muon $g-2$ experiment at Fermilab aims to measure the muon anomalous magnetic moment with an unprecedented precision of 140 parts per billion (ppb). Data collection concluded in June 2023, and analysis of the largest dataset (2021-2023) is underway. Previous publications based on data from 2018-2020 established the experimental foundation. This document provides an overview of the measurement of the muon anomalous spin precession frequency ($\omega_a$) and the associated systematic corrections. The precision of these results directly tests the Standard Model's completeness, making the experiment a cornerstone in the field of particle physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is that the amount of equipment reliability (ER) data being continuously generated are extremely large. These data elements come in different forms: textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) and they provide system engineers with valuable insights and information regarding the discovery of anomalous behaviors or degradation trends, the identification of the possible causes behind such behaviors and trends, and the prediction of their direct consequences. This paper directly targets the generation of knowledge from ER data by putting “data into context”. Here, we employ model-based system engineering (MBSE) models of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by identifying first which elements of the developed MBSE elements they are referring to. This task is much harder for textual data since the information contained in issue or maintenance reports needs to “be understood” by a computational tool. Here we called this process “knowledge extraction” where our methods to extract knowledge from textual data. Lastly, once numeric and textual ER data elements have been processed and “understood”, we discover possible cause-effect relations among them. This is performed by observing if a logical connection through the MBSE models exists, and if there is a temporal relation among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 MATHEMATICS AND COMPUTING↗

INTEGRATE – Inverse Network Transformations for Efficient Generation of Robust Airfoil and Turbine Enhancements

The INTEGRATE (Inverse Network Transformations for Efficient Generation of Robust Airfoil and Turbine Enhancements) project developed a new inverse-design capability for the aerodynamic design of wind turbine rotors using invertible neural networks. Training data was obtained from improved turbulence and transition models for RANS and hybrid RANS/LES solvers with machine-learned physics-based data-augmented corrections and then using the resulting neural-network(s) augmented RANS model to run thousands of 2-D and 3-D CFD simulations.

17 WIND ENERGY↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory↗

The MUSIC Critical Benchmark and Nuclear Data

The Measurement of Uranium Subcritical and Critical (MUSIC) experiment was a series of measurements of critical and subcritical configurations of bare highly enriched uranium. The goal was to compare measurement methods, analysis techniques, and simulation methods across regimes of criticality and to provide high-quality validation of 235 U nuclear data. A benchmark evaluation of the two critical configurations of the MUSIC experiments will soon be published in the release of the International Criticality Safety Benchmark Evaluation Project Handbook. The recent execution of the experiment aids in proper quantification of model simplifications and all uncertainties associated with the experiment. Historical benchmark evaluations are heavily relied on for uranium nuclear data validation despite the fact that the same level of documentation and comparable uncertainty analysis may not be present. The MUSIC evaluation is less likely to include “unknown unknowns” that could impede accurately modeling the system. Presented are both highly detailed and very simplified models, which represent the experimental configurations accurately, aiding the users of the benchmark for nuclear data or transport code validation. The sensitivities of k eff to nuclear data and nuclear data–related uncertainties are very similar between this experiment and previous bare uranium sphere experiments. In addition, the nuclear data uncertainties to any nuclides other than 235 U are small. For all these reasons, the recently evaluated MUSIC benchmark critical configurations could prove very useful for 235 U nuclear data validation. Currently, major libraries have good agreement with the experimental results, within 200 pcm for all nuclear data libraries, and within one standard deviation of the experimental result for most. Suggested nuclear data adjustments based on MUSIC and Lady Godiva are also presented, with posterior improvements to both the agreement in k eff and the uncertainty associated with the nuclear data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Uncertainty in inventories for life cycle assessment: State‐of‐the‐art, challenges, and new technologies

Uncertainty is a critical factor that can hinder the quality and potential applications of life cycle assessment (LCA) results. A prominent source of uncertainty stems from the life cycle inventory (LCI) data. Various methodologies exist to estimate the uncertainty associated with LCI data, primarily based on the widely used structured pedigree matrix approach or the computationally intensive Monte Carlo simulation. This perspective review explores how new technologies (e.g., computational algorithms and data collection methods) from data science and related fields can contribute to identifying, quantifying, and reducing uncertainty in LCI modeling. A brief overview of the sources of uncertainty in LCI modeling and how they are addressed in current LCA practice is provided. Additionally, several new technologies are identified, and the potential benefits of their implementation in reducing uncertainties in LCI modeling are discussed. This perspective review concludes by identifying potential areas that require further development for these technologies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗