Search NASASearch

SEARCH · Search NASA

Results for “data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Hyperspectral segmentation of plants in fabricated ecosystems

Hyperspectral imaging provides a powerful tool for analyzing above-ground plant characteristics in fabricated ecosystems, offering rich spectral information across diverse wavelengths. This study presents an efficient workflow for hyperspectral data segmentation and subsequent data analytics, minimizing the need for user annotation through the use of ensembles of sparse mixed scale convolution neural networks. The segmentation process leverages the diversity of ensembles to achieve high accuracy with minimal labeled data, reducing labor-intensive annotation efforts. To further enhance robustness, we incorporate image alignment techniques to address spatial variability in the dataset. Downstream analysis focuses on using the segmented data for processing spectral data, enabling monitoring of plant health. This approach provides a scalable solution for spectral segmentation, and facilitates actionable insights into plant conditions in complex, controlled environments. Our results demonstrate the utility of combining advanced machine learning techniques with hyperspectral analytics for high-throughput plant monitoring.

Zwart, Petrus H.

A Performance Model of In-Situ Techniques

The computational capacity of High-Performance Computing (HPC) systems increases continuously with the rapid development of central processing units (CPUs) and graphic processing units (GPUs), while the in-/output (IO) subsystem develops relatively slowly and storage capacity is also limited. Data-intensive applications, which are designed to leverage the high computational capacity of HPC resources, typically generate a considerable amount of data for post-processing visualizations and data analytics. The limited IO speed and storage space could lead to constraints in the actual performance of these applications and, therefore, scientific discovery. In-situ techniques, where data is visualized/analysed while still in memory rather than through disk, can contribute to alleviating these problems as they can reduce or even fully avoid data writing/reading through the IO subsystem to/from storage. However, the overall efficiency of insitu techniques crucially depends on the characteristics of both the in-situ tasks and the applications, and the resource distribution among them. Therefore, choosing the right in-situ approach (synchronous, asynchronous, or hybrid) and resource allocation is essential to minimize overhead and maximize the benefits of concurrent execution. In this paper, we present a performance model of in-situ techniques to find the most beneficial in-situ approach and the preferred resource configuration. We verify the high accuracy of our approach with over 6800 measurements and provide use cases with different applications.

Ju, Yi [Max Planck Computing and Data Facility, Ga

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics

Development of National New Construction Weighting Factors for the Commercial Building Prototype Analyses (2008-2022)

The U.S. Department of Energy (DOE) tasked Pacific Northwest National Laboratory (PNNL) with updating commercial building construction weights for the purpose of estimating national and state-by-state energy savings impacts of changes made to various commercial energy codes and standards. A similar activity was last completed by PNNL in 2020 using disaggregate construction volume data acquired from the Dodge Data & Analytics database (formerly McGraw Hill) for the years 2003-2018 (Lei et al, 2020). As time passes, changes in economic and social demand reshape construction volume trends. For the current update, PNNL reviewed the same data source with the latest construction data for the years 2008-2022. For commercial building analyses, PNNL typically uses a suite of 16 prototype buildings simulated in the 19 ASHRAE climate zones with 16 of them present in the United States. The 2008-2022 commercial building weighting factors were derived using the same approach employed to develop the 2003-2018 set (Lei et al, 2020). Applying the construction volume data from the database to the prototypes and climate zones resulted in the following new construction area-based weighting factors. Table ES.1 shows the weighting factors including all building categories found in the database, and Table ES.2 shows the weighting factors normalized to include only buildings represented by the 16 prototypes. Section 3.0 also includes national- and state-level weighting factors by area and building count.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

A digital twin platform for building performance monitoring and optimization: Performance simulation and case studies

Advancements in sensor technology, data analytics, affordable compute, and communication infrastructure have paved the way for Digital Twin technology in optimizing building operations and controls. This study presents the development of an open and interoperable web-based Digital Twin platform for integrating diverse data streams and facilitating effective user interactions. The platform utilizes modern technologies for the web framework and time-series data management, ensuring scalability and responsiveness. The backend supports seamless integration of diverse data sources and emulators, incorporating data from building sensors and meters, external weather Application Programming Interfaces, and advanced EnergyPlus simulation models of the building and its energy systems including the Distributed Energy Resources that are formulated in Functional Mockup Units. A simulation case study was conducted with FlexLab, a test facility on Lawrence Berkeley National Laboratory campus. The case study includes normal operations, Distributed Energy Resource integration, and power outage scenarios, to illustrate the Digital Twin’s ability to provide critical insights into energy performance and thermal resilience. The results demonstrated the platform’s potential as a decision-support tool for optimizing building energy performance and enhancing resilience against extreme weather events. Future work will focus on deploying the Digital Twin platform to a real building for field validation, extending its capabilities to cover more scenarios such as bidirectional Electric Vehicle interactions, and enhancing user engagement.

EnergyPlus

The Vertebrate Breed Ontology: Toward Effective Breed Data Standardization

Abstract Background Limited universally-adopted data standards in veterinary medicine hinder data interoperability and therefore integration and comparison; this ultimately impedes the application of existing information-based tools to support advancement in diagnostics, treatments, and precision medicine. Hypothesis/Objectives A single, coherent, logic-based standard for documenting breed names in health, production, and research-related records will improve data use capabilities in veterinary and comparative medicine. Animals No live animals were used. Methods The Vertebrate Breed Ontology (VBO) was created from breed names and related information compiled from the Food and Agriculture Organization of the United Nations, breed registries, communities, and experts, using manual and computational approaches. Each breed is represented by a VBO term that includes breed information and provenance as metadata. VBO terms are classified using description logic to allow computational applications and Artificial Intelligence–readiness. Results VBO is an open, community-driven ontology representing over 19 500 livestock and companion animal breed concepts covering 49 species. Breeds are classified based on community and expert conventions (e.g., cattle breed) and supported by relations to the breed's genus and species indicated by National Center for Biotechnology Information (NCBI) Taxonomy terms. Relationships between VBO terms (e.g., relating breeds to their foundation stock) provide additional context to support advanced data analytics. VBO term metadata includes synonyms, breed identifiers/codes, and attributed cross-references to other databases. Conclusion and Clinical Importance The adoption of VBO as a standard for breed names in databases and veterinary electronic health records enhances veterinary data interoperability and computability, supporting precision medicine.

Veterinary Sciences

AI Driven Optimization of Public Transit

This project explores the application of AI-driven methods to optimize public transit operations for the Chattanooga Area Regional Transportation Authority (CARTA). By leveraging data analytics, machine learning, and predictive modeling, the initiative seeks to enhance system efficiency, improve rider experience, and support sustainability goals. This research, supported by the National Science Foundation and the U.S. Department of Energy, integrates real-time transit data with advanced computational tools to inform decision-making, optimize routes, and balance operational demands. The work exemplifies a forward-looking model for mid-sized cities aiming to modernize mobility systems through intelligent technology integration.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Enhancing ZFP: A Statistical Approach to Understanding and Reducing Error Bias in a Lossy Floating-Point Compression Algorithm

The amount of data generated and gathered in scientific simulations and data collection applications is continuously growing, putting mounting pressure on storage and bandwidth concerns. A means of reducing such issues is data compression; but, lossless data compression is typically ineffective when applied to floating-point data. Thus, users tend to apply a lossy data compressor, which allows for small deviations from the original data. It is essential to understand how the error from lossy compression impacts the accuracy of the data analytics. Thus, we must analyze not only the compression properties but the error as well. In this paper, we provide a statistical analysis of the error caused by ZFP compression, a state-of-the-art, lossy compression algorithm explicitly designed for floating-point data. We show that the error is indeed biased and propose simple modifications to the algorithm to neutralize the bias and further reduce the resulting error.

97 MATHEMATICS AND COMPUTING

Improving Additive Manufactured Component Performance through Multi-Scale Microstructure Simulation and Process Optimization

The purpose of this project was to utilize computational tools to understand the relationships between processing, microstructure, and properties for additively manufactured (AM) aluminum alloys for automotive applications, and to provide an engineering solution for helping to optimize process conditions. The project leverages ORNL developments in computational modeling, including AM process modeling, phase-field based microstructure evolution predictions, and data analytics techniques for mapping process conditions to material outcomes. The project utilized an Al-Cu-Mn-Zr alloy as a model material for studying formation of defects and microstructural features in response to variations in process conditions. Based on both pre-existing experimental data and simulation results, statistical process maps were constructed to identify regions of process space with minimal defect formation and advantageous microstructures and properties. The software tools used for this purpose were successful disseminated to GM, who were able to successful compile the relevant HPC codes within their own computing ecosystem and perform initial calculations to reproduce ORNL results.

36 MATERIALS SCIENCE

Catalyzing deep decarbonization with federated battery diagnosis and prognosis for better data management in energy storage systems

Industrial data analytics methods play a central role in improving energy storage performance and efficiency, impacting the future of electrified transportation and renewable electricity generation. However, significant challenges hinder the large-scale deployment of batteries. Conventional methods rely on centralized collection and processing of fleet-level data, leading to database size issues and privacy concerns due to potential data breaches. To enable scalable deployment of battery management systems, this article proposes a federated battery diagnosis and prognosis model, which distributes the processing of battery standard current-voltage-time-usage data in a privacy-preserving manner. Instead of transferring the raw data, this approach communicates only the locally processed parameters, thus reducing communication load and preserving data confidentiality. The federated model offers a paradigm shift in battery health management through privacy-preserving distributed methods for battery data processing and lifetime prediction, ensuring the reliable and sustainable deployment of lithium-ion batteries in a rapidly evolving world.

asset health management

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES

Revealing systematic changes in the transcriptome during the transition from exponential growth to stationary phase

ABSTRACT The composition of bacterial transcriptomes is determined by the transcriptional regulatory network (TRN). The TRN regulates the transition from one physiological state to another. Here, we use independent component analysis to monitor the composition of the transcriptome during the transition from the exponential growth phase to the stationary phase. With Escherichia coli K-12 MG1655 as a model strain, we trigger the transition using carbon, nitrogen, and sulfur starvation. We find that (i) the transition to the stationary phase accompanies common transcriptome changes, including increased stringent responses and reduced production of cellular building blocks and energy regardless of the limiting element; (ii) condition-specific changes are strongly associated with transcriptional regulators ( e.g. , Crp, NtrC, CysB, Cbl) responsible for metabolizing the limiting element; and (iii) the shortage of each limiting element differentially affects the production of amino acids and extracellular polymers. This study demonstrates how the combination of genome-scale datasets and new data analytics reveals the fundamental characteristics of a key transition in the life cycle of bacteria. IMPORTANCE Nutrient limitations are critical environmental perturbations in bacterial physiology. Despite its importance, a detailed understanding of how bacterial transcriptomes are adjusted has been limited. By utilizing independent component analysis (ICA) to decompose transcriptome data, this study reveals key regulatory events that enable bacteria to adapt to nutrient limitations. The findings not only highlight common responses, such as the stringent response, but also condition-specific regulatory shifts associated with carbon, nitrogen, and sulfur starvation. The insights gained from this work advance our knowledge of bacterial physiology, gene regulation, and metabolic adaptation.

Lim, Hyun Gyu (ORCID:0000000204692388)

X-ray spectroscopy of multi-temperature plasmas using the differential emission measure formalism

We present a theoretical construct that nominally underlies spectroscopic data analysis of multi-temperature plasmas, known as the differential emission measure (DEM). From a data analytic perspective, the DEM formalism is used to derive temperature distributions from line spectra that are formed in the presence of temperature gradients and by time integrations of evolving plasmas. From a modeling perspective, DEMs are convenient intermediaries between radiation hydrodynamics simulations and spectroscopic measurements acquired in the laboratory. The DEM concept and its associated methodologies were originally developed by spectroscopists working with astrophysical data. We borrow from these earlier investigations. In this manuscript, intended primarily as a tutorial, we discuss the basic concepts, but also augment various aspects of the theory by the way of extension and example, including a detailed treatment of various weighting and averaging schemes, intended to mitigate ambiguities that often arise when reporting temperature information. We focus on high-temperature plasmas that are not in local thermodynamic equilibrium and the x-ray spectra that they produce, although the core ideas presented here are applicable to spectroscopy in other energy bands. A few examples involving the derivation and manipulation of model DEMs in simple geometries are provided.

Liedahl, Duane A. [Lawrence Livermore National Lab

Tutorial on In Situ and Operando (Scanning) Transmission Electron Microscopy for Analysis of Nanoscale Structure–Property Relationships

In situ and operando (scanning) transmission electron microscopy [(S)TEM] is a powerful characterization technique that uses imaging, diffraction, and spectroscopy to gain nano-to-atomic scale insights into the structure–property relationships in materials. This technique is both customizable and complex because many factors impact the ability to collect structural, compositional, and bonding information from a sample during environmental exposure or under application of an external stimulus. In the past two decades, in situ and operando (S)TEM methods have diversified and grown to encompass additional capabilities, higher degrees of precision, dynamic tracking abilities, enhanced reproducibility, and improved analytical tools. Much of this growth has been shared through the community and within commercialized products that enable rapid adoption and training in this approach. This tutorial aims to serve as a guide for students, collaborators, and nonspecialists to learn the important factors that impact the success of in situ and operando (S)TEM experiments and assess the value of the results obtained. As this is not a step-by-step guide, readers are encouraged to seek out the many comprehensive resources available for gaining a deeper understanding of in situ and operando (S)TEM methods, property measurements, data acquisition, reproducibility, and data analytics.

(S)TEM

Comparability of Liquid Chromatography Tandem Mass Spectrometry Analysis of Dissolved Organic Matter across Laboratories

Non-targeted liquid chromatography tandem highresolution mass spectrometry (LC−MS/MS) is increasingly applied for the structure-resolved chemical analysis of dissolved organic matter (DOM). With new developments in MS instrumentation and analysis software, the approach has gained substantial momentum over the past decade. However, achieving high-quality analytical data that is reproducible and comparable across laboratories can be a bottleneck in non-targeted metabolomics and organic matter chemical analysis, especially for data reuse in repository-scale analyses. Understanding the capabilities as well as challenges of comparing LC−MS/MS data from different laboratories is necessary for inferring global trends from public data sets. To illuminate instrumentation factors that drive differences and variability, we used a standardized data analysis pipeline, including classical (CMN) and featurebased molecular networking (FBMN), to analyze data from a ring trial by 24 laboratories on identical sample sets of algal and DOM extracts that were mixed in predefined concentrations and spiked with standards. Our results showed that data sets from similar mass spectrometer types with unified instrument parameters were qualitatively comparable, resolving the same general trends and shared mass spectral features. Interlaboratory comparability was best for high-intensity features, while low-intensity features showed greater detection variability. Our analysis also highlights challenges when comparing data from instruments with different acquisition rates or operating with less standardized methods. Lastly, we provide recommendations for data integration, public data sharing, standardization, and best practices for standardized LC−MS/MS data acquisition, which will be critical for long-term time series and intercomparability of DOM chemical analyses.

DOM

Analytical capabilities for iodine detection: Review of possibilities for different applications

This Review summarizes a range of analytical techniques that can be used to detect, quantify, and/or distinguish between isotopes of iodine (e.g., long-lived 129 I, short-lived 131 I, stable 127 I). One reason this is of interest is that understanding potential radioiodine release from nuclear processes is crucial to prevent environmental contamination and to protect human health as it can incorporate into the thyroid leading to cancer. It is also of interest for evaluating iodine retention performances of next-generation iodine off-gas capture materials and long-term waste forms for immobilizing radioiodine for disposal in geologic repositories. Depending upon the form of iodine (e.g., molecules, elemental, and ionic) and the matter state (i.e., solid, liquid, and gaseous), the available options can vary. In addition, several other key parameters vary between the methods discussed herein, including the destructive vs nondestructive nature of the measurement process (including in situ vs ex situ measurement options), the analytical data collection times, and the amount of sample required for analysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Emulation and detection of physical faults and cyber-attacks on building energy systems through real-time hardware-in-the-loop experiments

The increasing use of remote or mobile access, integrated wearable technologies, data exchange, and cloud-based data analytics in modern smart buildings is steering the building industry towards open communication technologies. The increased connectivity and accessibility could lead to more cyber-attacks in smart buildings. On the other hand, physical faults (e.g., HVAC -heating, ventilation, and air-conditioning faults) may have similar adverse impacts as those from the cyber-attacks on building energy systems, such as occupant discomfort, energy wastage, and equipment downtime. However, current physical behavior-based anomaly detection methods fail to differentiate between cyber-attacks and physical faults in building energy systems. Moreover, the challenge in collecting real-world threat data with ground truth has led researchers to rely on numerical models with user-defined assumptions, which may not accurately reflect real-world conditions due to the lack of in-situ experimental datasets. To address these challenges and gaps, this paper presents a flexible hardware-in-the-loop (HIL) testbed for generating cyber-attack and physical fault datasets and demonstrating threat detection algorithms in a real building automation system (BAS) environment. This testbed combines hardware (i.e., real BAS with local HVAC controllers and a physical network) with software (i.e., high-fidelity models to represent behaviors of building envelope and HVAC energy systems), enabling emulations of realistic threats. Five HIL experiments, including one baseline without any threats, two with physical faults, and two with cyber-attacks, were conducted to generate datasets containing detailed network traffic and system states. A joint classification framework, incorporating a network analyzer and a physical HVAC fault detector, was proposed to automatically detect cyber-physical abnormalities on BAS at both the network and the physical HVAC levels. The network analyzer comprises a conditional random fields (CRF) based command validator and a statistics-based detection strategy. The fault detector employs a weather and schedule-based pattern matching and feature-based principal component analysis (WPM-FPCA) method. Evaluation of the classification using four metrics from the multi-class confusion matrix revealed an average accuracy of 90.2%, recall of 89.7%, precision of 88.5% and F1-score of 89.2%. Finally, these results demonstrate that the proposed joint classification framework can effectively differentiate between specific types of cyber-attacks (e.g., device reinitialization attack, network Denial-of-Service attack) and physical faults (e.g., air handling unit operational fault, cooling coil valve stuck) in real time for improved building energy management.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI