Search NASASearch

SEARCH · Search NASA

Results for “data lifecycle”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING

Recommendations for Best Practices for Data Preservation and Open Science in HEP

These recommendations are the result of reflections by scientists and experts who are, or have been, involved in the preservation of high-energy physics data. The work has been done under the umbrella of the Data Lifecycle panel of the International Committee of Future Accelerators (ICFA), drawing on the expertise of a wide range of stakeholders. A key indicator of success in the data preservation efforts is the long-term usability of the data. Experience shows that achieving this requires providing a rich set of information in various forms, which can only be effectively collected and preserved during the period of active data use. The recommendations are intended to be actionable by the indicated actors and specific to the particle physics domain. They cover a wide range of actions, many of which are interdependent. These dependencies are indicated within the recommendations and can be used as a road map to guide implementation efforts. These recommendations are best accessed and viewed through the web application, see https://icfa-data-best-practices.app.cern.ch/

Campana, Simone [CERN]

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

A critical review and meta-analysis of energy demand, carbon footprint, and other environmental impacts from carbon fiber manufacturing

The demand for carbon fibers and carbon fiber-reinforced polymers (CFRPs) is rapidly growing due to their outstanding mechanical properties and potential to enhance sustainability, particularly for lightweighting applications. However, carbon fibers are typically produced from fossil-based feedstocks, involve energy-intensive processes, and have limited options for sustainable end-of-life management or circularity. Despite these challenges, the energy demand and lifecycle environmental implications of their production remain poorly understood. Here, we conduct a critical literature review and meta-analysis of carbon fiber manufacturing, revealing significant variations in reported energy demand, carbon footprint, and lifecycle inventory data. Our analysis makes two novel contributions. First, we identify key underlying factors driving these variations. Second, we highlight that carbon fiber, far from being a homogeneous product, has grades varying substantially in mechanical properties, end-use markets, energy intensity of manufacturing processes, and therefore environmental impacts—an aspect often underrepresented in life cycle assessments. We assert that current data are insufficient for reliably evaluating environmental impacts, posing a risk of misleading decision-making. Addressing this gap requires new lifecycle inventory datasets clearly incorporating carbon fiber heterogeneity and key influencing factors identified in this study. Additionally, we propose actionable recommendations, including a checklist, to advance sustainability in the carbon fiber sector.

CED

ACTIVE

The Automated Control Testbed for Integration, Verification, and Emulation (ACTIVE) framework is a software platform designed to support the optimized operation and management of a wide range of building types. It enables the development, testing, and validation of diverse control strategies, including AI-based, rule-based, and model-based approaches. The platform facilitates a seamless transition from simulation-based evaluation of control strategies to real-world field validation and deployment. ACTIVE supports the full building management lifecycle, encompassing data acquisition and management, system monitoring, optimized control, adaptive learning services, device dispatch and coordination, as well as advanced analytics and visualization. Together, these capabilities provide an integrated environment for improving building performance, operational efficiency, reducing energy cost, and reliability.

Smith, Robert [Oak Ridge National Laboratory (ORNL

WholeTraveler Anonymized Data Phase 1

Phase 1 of the WholeTraveler Study data collection consisted of an online-only survey. This survey captured data on three categories of observable variation in the population relevant to transportation decisions. First, the survey collected traditional demographic data such as age, gender, income, and education level. Second, it collected data across personality, psychological, and preference categories. This included: 1. The "Big Five" inventory personality traits: openness to new experience, conscientiousness, extroversion, agreeableness, and neuroticism; 2. Risk and time preferences; and 3. Environmental preferences. Third, the survey collected data on historical behavior patterns including: 1. Adoption of (as well as interest in) new technologies or innovations (e.g., smartphones, PEVs, solar panels, adaptive cruise control [ACC]); 2. Car ownership history and current car ownership status; 3. Recent mode use across different time scales (e.g., previous week, previous month, previous year); and 4. Timing of major life events such as starting a family as well as overall lifecycle trajectory patterns. Data from Phase 1 and Phase 2 are linked by a unique respondent identifier. Anonymized versions of the Phase 1 and Phase 2 data are both available on Livewire.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

A Novel Concentrating Solar Weathering Apparatus for Experimental Validation of Multi-Modal Degradation Models

High-performance coatings for Concentrating Solar Power (CSP) receivers are subjected to remarkable environmental stressors during normal operations. Applied to the receiver tubes, these coatings serve to maximize the solar absorptivity of the receiver, transferring as much heat as possible from the solar collectors into the heat-transfer fluid (HTF). The lifecycle of these coatings is not well-defined, and the harsh operational conditions make them difficult to test. NREL has designed, built, and tested an apparatus to expose these samples to design levels of environmental stress and well beyond, into accelerated and destructive conditions. The chamber is actively cooled, monitored, and has the capability to supply humidification for cycling tests, allowing us to test multi-modal degradation and failure conditions at high temperature, high flux, and high humidity conditions. These conditions can catalyze high-temperature oxidation, mechanical degradation, and other modes of absorptivity loss seen in selective solar receiver coatings. The experimental data can feed lifecycle models for expensive and necessarily resilient materials, offering insights to aid maintenance schedules, technoeconomic analysis, and material industry performance benchmarks. This presentation will demonstrate the apparatus design and performance, as well as initial results for aging on a selective receiver coating.

14 SOLAR ENERGY

Predictive Phenomics Initiative Project Dataset Catalog Collection

The Predictive Phenomics Science & Technology Initiative (PPI) at Pacific Northwest National Laboratory are tackling the grand challenge of understanding and predicting phenotype by identifying the molecular basis of function and enable function-driven design and control of biological systems. Research projects within this initiative are divided into three Thrust Areas (TAs): TA1) Enhancing Multi-Scale Phenomics Measurements, TA2) Identifying Molecular Patterns of Biological Function, and TA3) Computational Methods - Phenotypic Signatures. In efforts to enable discovery, reproducibility, and reuse of PPI-funded digital research data generated or used through the course of the proposed research-funded lifecycles, all corresponding digital data assets conducted under the Laboratory Directed Research and Development Program at PNNL are linked to this PPI dataset catalog collection.

59 BASIC BIOLOGICAL SCIENCES

Predictive Phenomics Initiative Project Dataset Catalog Collection

The Predictive Phenomics Science & Technology Initiative (PPI) at Pacific Northwest National Laboratory are tackling the grand challenge of understanding and predicting phenotype by identifying the molecular basis of function and enable function-driven design and control of biological systems. Research projects within this initiative are divided into three Thrust Areas (TAs): TA1) Enhancing Multi-Scale Phenomics Measurements, TA2) Identifying Molecular Patterns of Biological Function, and TA3) Computational Methods - Phenotypic Signatures. In efforts to enable discovery, reproducibility, and reuse of PPI-funded digital research data generated or used through the course of the proposed research-funded lifecycles, all corresponding digital data assets conducted under the Laboratory Directed Research and Development Program at PNNL are linked to this PPI dataset catalog collection.

59 BASIC BIOLOGICAL SCIENCES

IPC-Fusion (Infrastructure Perception and Control (IPC): Multisensor Data Fusion Software) [SWR-25-153]

As part of the National Laboratory of the Rockies' (NLR’s) Infrastructure Perception and Control Laboratory, the IPC-Fusion toolkit provides a probabilistic, scalable, multi-sensor fusion framework that integrates (late-stage fusion) heterogeneous object detection data from traffic sensors to enable robust, real-time tracking of roadway occupants. The algorithmic design of the toolkit is motivated by the need for creating a digital twin of traffic at the edge in a scalable and affordable manner. The software operates by combining object-level measurements (such as position and velocity) from a suite of sensors (such as radar, lidar, camera) using Kalman filtering and probabilistic data association techniques to overcome individual sensor limitations and achieve superior tracking performance in complex traffic zones. The framework addresses key challenges including heterogeneous measurement uncertainties, asynchronous data streams, varying spatiotemporal data resolutions, robust data association, and adaptive object lifecycle management. Validated on real-world traffic intersection data including vehicles and pedestrians, IPC-Fusion demonstrates enhanced tracking reliability across scenarios involving occlusions, sensor failures, and varying traffic densities, supporting the broader IPC initiative's goal of transforming transportation infrastructure through advanced perception capabilities for intelligent transportation systems, traffic safety applications, and autonomous vehicle support.

Sandhu, Rimple [National Laboratory of the Rockies

DEMOS (Demographic Microsimulator Tool for Longitudinal Synthetic Population) [SWR-25-135] related to NLR SWR-26-076

The Demographic Microsimulator (DEMOS) is an agent-based simulation framework used to model the evolution of population demographic characteristics and lifecycle events, such as education attainment, marital status, and other key transitions. DEMOS modules are designed to capture the interdependencies between short-term and long-term lifecycle events, which are often influential in downstream transportation and land-use modeling. A key feature of DEMOS is its ability to track changes in an agent’s demographic status from year t to year t + 1. This structure allows the model to evolve populations over any user-defined time horizon. As a result, DEMOS is well suited for analyzing medium- and long-term transportation-related decisions, including household vehicle transactions (e.g., purchasing, selling, or replacing vehicles) and work location choices. Core features of DEMOS include the modeling of more than ten lifecycle events, behaviorally realistic patterns informed by long-running panel data, explicit representation of interdependencies among lifecycle processes, and a flexible, modular simulation architecture. A technical memorandum describing DEMOS is available here. The memorandum provides an overview of the framework’s functionality, model structure, input and output data, and its applications in transportation planning and broader policy analysis contexts. Interested readers are also encouraged to consult the paper listed below for additional details on the DEMOS methodology. Sun, Bingrong, Shivam Sharda, Venu M. Garikapati, Mohamed Amine Bouzaghrane, Juan Caicedo, Srinath Ravulaparthy, Isabel Viegas de Lima, Ling Jin, C. Anna Spurlock, and Paul Waddell. "Demographic Microsimulator for Integrated Urban Systems: Adapting Panel Survey of Income Dynamics to Capture the Continuum of Life." Transportation Research Record (2025): 03611981251333339.

Sun, Bingrong [National Laboratory of the Rockies

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS

Powering Circularity Through Data Reporting and Collection

Sustainability and Circular Economy have many metrics for evaluation. Calculating mass intensity, energy return on investment, financial payback, and recycling rate for proposed technology changes and lifecycle management can support decision making. Robust data with modeling tools can perform these calculations, informing good decision making. Analyses show that reliability is more critical than recyclability. Improved data gathering and tool accessibility will support our industry to make more circular choices for PV lifecycle management.

14 SOLAR ENERGY

AI-Ready Control System for the Fermilab Accelerator Complex

Reliable, high-intensity operation of the Fermilab Accelerator Complex is critical to the success of the Long-Baseline Neutrino Facility and Deep Underground Neutrino Experiment. We describe the requirements and infrastructure necessary to support routine use of artificial intelligence and machine learning (AI/ML) in the accelerator control system. Three capabilities are identified: a machine learning operations (MLOps) framework standardizing the lifecycle of AI/ML automation from data management through deployment and monitoring; a data quality framework defining and enforcing standards required to build trustworthy AI/ML applications; and workflow integration with large language models to assist physicists, engineers, and operators with information retrieval, code development, and routine analysis. Use cases spanning beam diagnostics, beam control, and support system automation illustrate the technical requirements across the complex.

43 PARTICLE ACCELERATORS

Integrated System Planning: Emerging Software Requirements in the Power Industry

Power system planning software remains fragmented across organizational boundaries, with specialized tools for capacity expansion, production cost modeling, power flow, and dynamic analysis operating on incompatible data models and assumptions. This article argues that the fragmentation is not merely a technical problem but a predictable consequence of Conway's law: software architectures mirror the departmental structures within which they are developed. Regulatory milestones like Federal Energy Regulatory Commission (FERC) Order 888 formalized these divisions, but the roots trace back to the distinct engineering disciplines-mechanical, chemical, and electrical-that staffed generation and transmission planning departments in vertically integrated utilities. As the industry moves toward integrated system planning (ISP) that coordinates generation, transmission, and distribution investment decisions, the software ecosystem must evolve accordingly. We identify five categories of software requirements to enable this transition: coherent data inputs decoupled from individual applications, unified and extensible data schemas, modular component representations that support multiple abstraction levels, lifecycle management of planning datasets, and well-defined application programming interface (API) contracts that separate data exchange from algorithmic control. We examine how these requirements interact with three common workflow patterns-serial gate clearing, sequential multiapplication, and convergence oriented-and discuss the interface design principles each demands. We then outline a vision for platform-based planning architectures where specialized analytical services compose through standardized interfaces and where artificial intelligence (AI)/machine learning (ML) tools augment decision support within a disciplined software infrastructure. The practices proposed here offer a path from today's siloed tool collections toward collaborative planning ecosystems capable of handling the complexity of modern power system transformation.

24 POWER TRANSMISSION AND DISTRIBUTION