Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Smart Pixel Sensors for the HL-LHC

Large-scale particle physics experiments produce tens of terabytes of data every second. Innovative methods to manage the data rate at the HL-LHC, which expects to operate at 10x the luminosity of what the LHC was initially designed for, are needed. AI-on the chip provides a way to intelligently filter out low momentum clusters in the pixel detector. This will open up an opportunity to use the pixel detector for the first time in the CMS Level-1 trigger, and lead to increased sensitivity to new physics measurements and searches. We have taped out our first chip, which incorporates a $p_T$ filtering algorithm on an ASIC chip. Our initial $p_T$ filtering algorithm considers clusters that are tracked by CMS. We will report on ongoing studies seeking to enhance the performance of our filter by utilizing unsupervised learning on untracked clusters, thus increasing background rejection.

43 PARTICLE ACCELERATORS↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Streaming Readout and Data-Stream Processing With ERSAP

With the exponential growth in the volume and complexity of data generated at high-energy physics and nuclear physics research facilities, there is an imperative demand for innovative strategies to process this data in real or near-real-time. Given the surge in the requirement for high-performance computing, it becomes pivotal to reassess the adaptability of current data processing architectures in integrating new technologies and managing streaming data. This paper introduces the ERSAP framework, a modern solution that synergizes flow-based programming with the reactive actor model, paving the way for distributed, reactive, and high performance in data stream processing applications. Additionally, we unveil a novel algorithm focused on time-based clustering and event identification in data streams. The efficacy of this approach is further exemplified through the data-stream processing outcomes obtained from the recent beam tests of the EIC prototype calorimeter at DESY.

Vardan, Gyurjyan↗

Data for Soil Fertility Management for Sustainable Miscanthus × giganteus Production: Increased Tiller Weight from Nitrogen Management Explains Yield Gains in Aged Miscanthus

Aging-related yield decline in Miscanthus × giganteus (miscanthus) remains a major constraint to sustainable biomass production. This study evaluated how nitrogen (N) management and soil fertility influence yield-component traits and productivity in aging miscanthus. Trials were conducted at two sites established in 2008 at the University of Illinois Energy Farm, Urbana, IL. (i) The Sun Grant trial received 0, 60, and 120 kg N ha−1 annually until 2015. Starting 2021, half of each plot received 60 or 120 kg N ha−1, resulting in six legacy-contemporary treatments: 0N–0N, 0N–120N, 60N–0N, 60N–60N, 120N–0N, 120N–120N. (ii) The Energy Farm trial remained unfertilized until 2014, when one half of each plot received 56 kg N ha−1, forming two treatments: 0N–0N, 0N–56N. Sun Grant trial results showed N fertilization increased tiller density (tillers m−2) and tiller weight (g tiller−1) in juvenile to early-mature miscanthus (2011–2015). After N withdrawal, both traits declined (20 % and 40 %), though legacy effects persisted in tiller weight in the aging stands (2020–2023). Contemporary N had little effect on tiller density but increased tiller weight by 34 %–77 %, resulting in 23 %–106 % higher machine-harvested biomass yield in 0–120N, 60-60N, and 120-120N plots. At the Energy Farm trial, 0N–56N plots yielded 59 %–108 % more biomass than 0N–0N. Soil total N increased (Sun Grant: 47 % by 2020; Energy Farm: 58 % by 2023), while Mehlich-3 P (42 %–44 %) and K (21 %–46 %) declined. These findings identify tiller weight as a key determinant of biomass yield in aging miscanthus and highlight the need for P and K management for long-term productivity.

miscanthus↗

Modular Subsurface Sensors and Integrated Software for Advanced Subsurface Characterization and Monitoring using Unoccupied Vehicles

The advent and subsequent proliferation of autonomous airborne, waterborne, and groundbased vehicles (i.e., “drones”) promises to broadly transform the geosciences and associated industries, including fossil energy exploration and development, mineral resource exploration and development, water-resource management, and environmental remediation. For geophysical characterization and monitoring, the prospect of programming highly repeatable and low-cost drone missions for subsurface imaging will allow for deployments in hazardous and previously inaccessible areas. Coupled with autonomous workflows for data processing, management, and visualization, drone-based geophysical characterization and monitoring will enable unprecedented, real-time insight into diverse subsurface properties and processes of scientific and engineering importance. Toward this end, the objectives of this Lab Directed Research and Development (LDRD) project were to develop new (1) instrumentation for dronebased electromagnetic induction (EMI) geophysical imaging, including separated transmitter and receivers and associated electronics, (2) software for real-time data telemetry, processing, management, and visualization. Although EMI has been previously deployed using unoccupied aerial systems (UASs), these applications failed to capitalize on the game-changing capabilities of drone platforms. Whereas drone-based data acquisition allows for collection of rich, three-dimensional (3D) multi-offset/multi-angle configurations between transmitters and receivers, past efforts have relied on conventional instrumentation that was designed for ground-based data collection with the transmitter and a single receiver housed in the same unit; nor did these previous applications demonstrate real-time delivery of results to support rapid management decisions in the field. In this 1-year project, we (1) designed and constructed new lightweight independent transmitter and receiver antenna platforms that communicate with a laptop computer; (2) developed software to control data acquisition, manage/transfer data, and visualize data as its collected; and (3) demonstrated the operation of the new hardware and software systems in a ground-based field test. Our work entails major technological advances for EMI and established a foundation on which to build a new drone-based, real-time geophysical EMI imaging capability to support diverse challenges facing the nation.

47 OTHER INSTRUMENTATION↗

Pushing the limits of NAND technology scaling with ferroelectrics

Artificial intelligence (AI) continues to drive transformative advancements across various industries. The data-intensive nature of AI training (and inferencing) has resulted in the generation of unprecedented volumes of data with machine-generated content surpassing human-generated data by more than 100-fold in 2025. Efficiently managing this data influx necessitates advanced digital storage technologies. However, traditional NAND flash memory, which is critical for supporting data flows in AI systems—alongside high-bandwidth memory, for AI training—faces fundamental scaling limitations as it approaches the 1000-layer milestone, encompassing more than 40 trillion transistors. This article delves into the potential of hafnia-based ferroelectric materials as a breakthrough solution to these challenges. Recent advancements indicate that the intrinsic limitations of ferroelectric field-effect transistors (FEFETs) can be mitigated through material and device-level engineering. These advancements enable FEFETs to meet the stringent density, reliability, and scalability requirements of future three-dimensional NAND technology. The role of ferroelectrics in addressing NAND scaling challenges and expanding storage capabilities presents a promising avenue for meeting the storage demands of the AI-driven era.

3D NAND↗

Population structure limits the use of genomic data for predicting phenotypes and managing genetic resources in forest trees

There is overwhelming evidence that forest trees are locally adapted to climate. Thus, genecological models based on population phenotypes have been used to measure local adaptation, infer genetic maladaptation to climate, and guide assisted migration. However, instead of phenotypes, there is increasing interest in using genomic data for gene resource management. We used whole-genome resequencing and common-garden experiments to understand the genetic architecture of adaptive traits in black cottonwood. We studied the potential of using genome-wide association studies (GWAS) and genomic prediction to detect causal loci, identify climate-adapted phenotypes, and inform gene resource management. We analyzed population structure by partitioning phenotypic and genomic (single-nucleotide polymorphism) variation among 840 genotypes collected from 91 stands along 16 rivers. Most phenotypic variation (60 to 81%) occurred among populations and was strongly associated with climate. Population phenotypes were predicted well using genomic data (e.g., predictive abilityr> 0.9) but almost as well using climate or geography (r> 0.8). In contrast, genomic prediction within populations was poor (r< 0.2). We identified many GWAS associations among populations, but most appeared to be spurious based on pooled within-population analyses. Hierarchical partitioning of linkage disequilibrium and haplotype sharing suggested that within-population genomic prediction and GWAS were poor because allele frequencies of causal loci and linked markers differed among populations. Given the urgent need to conserve natural populations and ecosystems, our results suggest that climate variables alone can be used to predict population phenotypes, delineate seed zones and deployment zones, and guide assisted migration.

Science & Technology - Other Topics↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

Do we have globally representative data to understand soil processes?

Understanding and modeling soils and soil organic matter (SOM) are central to a variety of human needs, from food production to ecosystem management. Soil data have been collected for over a century, but the global spatial and process representativeness of soil data remains unclear. We assessed the representativeness of currently available soil data that could be used to understand a variety of SOM processes. We used 16 open-source soil databases and data from over 281,000 unique locations globally, categorizing the databases into three main data types necessary to understand SOM processes: soil carbon stocks and fluxes, mechanistic drivers of these stocks and fluxes, and soil carbon gain or loss potential. We found that stock and driver data have extensive global coverage. However, data on soil carbon gain or loss potential, particularly data describing change in soils over time such as time series data, are severely limited in their global coverage. We conclude that while significant strides have been made in measuring soil carbon stocks and fluxes, and their drivers, we are limited in global data related to changes in soils over time. Our recommendations for soil data generators are to ensure precise metadata reporting and prioritizing sampling in underrepresented areas like tropical, arctic, mountainous, wetland and arid regions. We also encourage designing revisit schemes that explicitly support change detection and reporting multi-modal datasets that can aid in model development. Targeted measurement of low coverage soil data types and regions is necessary for a range of applications including current and future biogeochemical predictions, and their management and policy implications.

carbon fluxes↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Framework of compressive sensing and data compression for 4D-STEM

Four-dimensional Scanning Transmission Electron Microscopy (4D-STEM) is a powerful technique for high-resolution and high-precision materials characterization at multiple length scales, including the characterization of beam-sensitive materials. However, the field of view of 4D-STEM is relatively small, which in absence of live processing is limited by the data size required for storage. Furthermore, the rectilinear scan approach currently employed in 4D-STEM places a resolution- and signal-dependent dose limit for the study of beam sensitive materials. Improving 4D-STEM data and dose efficiency, by keeping the data size manageable while limiting the amount of electron dose, is thus critical for broader applications. Here we introduce a general method for reconstructing 4D-STEM data with subsampling in both real and reciprocal spaces at high fidelity. The approach is first tested on the subsampled datasets created from a full 4D-STEM dataset, and then demonstrated experimentally using random scan in real-space. The same reconstruction algorithm can also be used for compression of 4D-STEM datasets, leading to a large reduction (100 times or more) in data size, while retaining the fine features of 4D-STEM imaging, for crystalline samples.

4D-STEM↗

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE↗

Real Time Phasor Analytics (RTPA) and RTPA-SCR System Strength Online Tool

This presentation showcases the Real-Time Phasor Analytics (RTPA) framework for monitoring inertia and assessing system strength in power grids. RTPA is an open-source tool designed to standardize access to data from Power Management Units (PMUs) and Phasor Data Concentrators (PDCs). It facilitates real-time connectivity to multiple PDCs in accordance with the IEEE C37.118-2 standard and supports asynchronous data stream integration. Additionally, RTPA can simulate a PDC server streaming C37.118-2 data and provides Python bindings for seamless interaction with the framework, eliminating the need for direct Rust programming.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Investigation of fast and efficient lossless compression algorithms for macromolecular crystallography experiments

Structural biology experiments benefit significantly from state-of-the-art synchrotron data collection. One can acquire macromolecular crystallography (MX) diffraction data on large-area photon-counting pixel-array detectors at framing rates exceeding 1000 frames per second, using 200 Gbps network connectivity, or higher when available. In extreme cases this represents a raw data throughput of about 25 GB s −1 , which is nearly impossible to deliver at reasonable cost without compression. Our field has used lossless compression for decades to make such data collection manageable. Many MX beamlines are now fitted with DECTRIS Eiger detectors, all of which are delivered with optimized compression algorithms by default, and they perform well with current framing rates and typical diffraction data. However, better lossless compression algorithms have been developed and are now available to the research community. Here one of the latest and most promising lossless compression algorithms is investigated on a variety of diffraction data like those routinely acquired at state-of-the-art MX beamlines.

36 MATERIALS SCIENCE↗