Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Systems Engineers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING↗

Project Phase 1 Report: Reducing Data Center Peak Cooling Demand and Energy Costs With Cold Underground Thermal Energy Storage (Cold UTES)

Cold Underground Thermal Energy Storage (Cold UTES) is an ultra-long duration grid energy storage technology. With Cold-UTES, low-cost grid power is converted to cold thermal energy and stored in the native subsurface rock at the point of use. Cold UTES is one approach within the general category of engineered geothermal systems. Cold UTES for peak-hour cooling of data centers (DCs) was studied for deployment in Maricopa County, Arizona and Loudoun County, Virgina using thermal storage capacities from 4 GWh-th to over 1,000 GWh-th (>1 Terawatt-hour). The two sites have different power and transmission systems, available grid energy resources, daily and seasonal load profiles, weather conditions, and grid regulatory requirements. The study results indicate high value for both locations and because of this, likely indicates value across most of the US and the world. The basis of the study was a 1,000 MW-e hourly electric use of DC computing and auxiliary loads, which was modeled as 1 GW-th of thermal load to a dry cooled heat rejection system - i.e. a cooling system that does not consume water. The electric power required for cooling the DC varies as the air temperature changes. In cold weather only the dry-coolers are used, with an electrical load for cooling load as low as 10 MW-e. In hot summer hours, chillers and dry- coolers are required, which raises the electrical load for cooling load to as much as 300 MW-e. The continuous and peak cooling electrical loads result in a grid interconnection requirement of no less than 1,300 MW-e. From both a grid and thermal design modeling perspective the 1.3 GW-e could either be a single facility or result from the total load at multiple sites.

15 GEOTHERMAL ENERGY↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗

Visual Systems Mapping to Define and Compare Woody Biomass LCAs for Sustainable Systems

The challenge addressed in this research centres on the need to choose between several biomass sources and energy production processes, while supporting rural economies and resilience of forest systems. A key barrier to effective decision-making for strategies using biomass is the lack of standardized and transparent life cycle assessment (LCA) baselines. These baselines are critical for assessing the impacts of biomass strategies but often vary due to regional factors and chosen simplifying assumptions of the LCAs. However, omitting key variables can mean the LCA omits key feedback and balancing loops relevant to fully assessing impacts of the change or test scenario. To address these complexities, this project employs a systems engineering approach: visual systems mapping. This technique is used to define the boundaries and dynamic behaviours of LCA baselines, enhancing transparency. By examining five literature sources and their documented baseline scenarios, the systems mapping case-studies demonstrates an approach to documenting and archiving these baselines. Recommendations are that visual systems mapping should be used to document key assumptions, such as baselines, of LCAs. Further, where possible open data repositories should hold key information about LCA baselines and reproducible workflows (e.g., using open-source tools) should be used to improve transparency and comparability in LCAs. Given the consensus within the broader scientific community on the importance of replicable data practices, this research reinforces the need for standardized frameworks and systems engineering tools in LCAs. This research demonstrates a pathway to more transparent, standardized, and comparable LCAs, that may bolster decisions for biomass systems.

Davis, Maggie [ORNL] (ORCID:0000000181319328)↗

Model-based economic analysis under uncertainty for PFAS treatment by granular activated carbon and ion exchange technologies

Recent drinking water regulations have imposed the need for per- and polyfluoroalkyl substances (PFAS) remediation. In response, treatment facilities may be required to retrofit existing treatment schemes to treat PFAS below maximum contaminant levels (MCLs). Adsorption technologies such as granular activated carbon (GAC) and ion exchange (IX) have been demonstrated to be effective; however, there are limited techno-economic metrics available which provide guidance on technology selection and design for diverse PFAS-containing source water conditions. Process systems engineering (PSE) tools which can traditionally perform these analyses are hindered by the data availability, model validity, and understanding of treatment phenomena for emerging contaminants. This work employs published data regressions, statistical models, process models, techno-economic analyses, and other process systems tools in a model-based uncertainty framework to consider the limitations of emerging contaminant research. Through this analysis framework, economic results are provided as probabilistic distributions based on the uncertainty of the models and diverse conditions that treatment facilities experience.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving the User Interface of the DeepLynx Data Warehouse

DeepLynx is an open-source ontology-based data warehouse created by INL to support the creation and life cycle of digital engineering projects, with a particular emphasis on digital twins [1]. Digital twins are systems that represent physical assets and process in a real-time digital environment [1]. Most well-known commercial data warehouses use Graphical User Interfaces (GUIs) for users to interact with their systems [3]. Limited publications have addressed the design of these interfaces and understanding of their target users. The current users and development team acknowledge the need to improve the current UI, not just for aesthetics but to improve functionality and workflow of DeepLynx. Traditional data warehouse users are developers, data scientists and business analysts [2]. DeepLynx users have a vast range of experience using data warehouses, and diverse roles, including engineers, scientists and management positions. Because there is a broader audience of target users for DeepLynx than a typical data warehouse, it is essential that DeepLynx has a useable and intuitive user interface. To achieve this the team performed human-computer interaction methods, including a Heuristic Evaluation of current UI using Neilsen’s Usability Heuristic, create personas based on current users by designing a user survey, data analysis and develop of personas. Followed by a redesign of the UI following using Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design in industry standard software Figma. Lastly a Heuristic Evaluation of new UI design, using Neilsen’s Usability Heuristic and User testing of redesign UI and have a group of users complete a Thinking Aloud Test of the new UI. Preliminary results of the Heuristic Evaluation of current UI arise issue with Consistency and Standards, Visibility of System Status, Match System and Real World and Recognition Rather than Recall. These issues were addressed in the proposed redesign by applying Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design. Next steps include formalized list of lessons learned and design implications for future publications.

97 MATHEMATICS AND COMPUTING↗

Data Structure Alchemy

In an increasingly more data-driven world, the project set out to uncover the first principles of data-structure design, chart the immense design space they form, and build automation that can synthesize an optimal structure, or even a whole storage engine, for any given workload, hardware platform, and cost target. Data structures are at the center of every computational system and are directly responsible for its performance. Two core technical thrusts were defined: 1) Mapping design spaces for key data-centric abstractions (filters, hash functions, storage-engine layouts, neural-network topologies, blockchain protocols, image layouts, etc.). 2) Developing search & synthesis algorithms, initially analytical cost models, later neural-guided bi-level optimisers that navigate sextillions of candidate designs in seconds and materialise the best one as ready‐to-run code. This report distills the key insights, accomplishments, and impact.

97 MATHEMATICS AND COMPUTING↗

Learning dynamical systems from data: An introduction to physics-guided deep learning

Modeling complex physical dynamics is a fundamental task in science and engineering. Traditional physics-based models are first-principled, explainable, and sample-efficient. However, they often rely on strong modeling assumptions and expensive numerical integration, requiring significant computational resources and domain expertise. While deep learning (DL) provides efficient alternatives for modeling complex dynamics, they require a large amount of labeled training data. Furthermore, its predictions may disobey the governing physical laws and are difficult to interpret. Physics-guided DL aims to integrate first-principled physical knowledge into data-driven methods. It has the best of both worlds and is well equipped to better solve scientific problems. Recently, this field has gained great progress and has drawn considerable interest across discipline Here, we introduce the framework of physics-guided DL with a special emphasis on learning dynamical systems. We describe the learning pipeline and categorize state-of-the-art methods under this framework. We also offer our perspectives on the open challenges and emerging opportunities.

97 MATHEMATICS AND COMPUTING↗

ACDC (Automated Campbell Diagram Code) [SWR-26-042]

This application provides a web-based graphical user interface to generating Campbell Diagrams and visualizing mode shapes for OpenFAST turbine models. Determining the aeroelastic stability and dynamic characteristics of wind turbines is a critical step in turbine design and analysis. Historically, extracting natural frequencies and mode shapes from OpenFAST—the industry-standard whole-turbine simulation code—has been a fragmented and tedious process. It required manual model configuration, command-line linearization execution, and complex post-processing via proprietary scripts to handle rotating-frame dynamics. To address these workflow bottlenecks, we present the Automated Campbell Diagram Code (ACDC), an open-source graphical software tool developed by the National Laboratory of the Rockies (NLR) under the DOE-funded Distributed Wind Aeroelastic Modeling (dWAM) project. ACDC streamlines the end-to-end linearization and stability analysis workflow into a single, intuitive cross-platform application. The software guides users through OpenFAST model configuration, definition of operating points, and the automated execution of steady-state trim and linearization simulations. Under the hood, ACDC automates the complex mathematical post-processing steps required for rotating systems, including Multi-Blade Coordinate (MBC) transformations, eigenanalysis, and advanced modal tracking utilizing the Modal Assurance Criterion (MAC) and spectral clustering. Finally, ACDC processes these results to automatically generate Campbell diagrams and features a robust 3D visualization engine to animate full-system mode shapes. By eliminating the reliance on external post-processing environments and manual data manipulation, ACDC significantly accelerates dynamic analysis and lowers the barrier to entry for wind energy researchers and engineers.

Summerville, Brent [National Laboratory of the Roc↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

Report on Initial Sodium Testing on the Thermal Hydraulic Experimental Test Article (THETA) (Fiscal Year 2024 Final Report)

The Thermal Hydraulic Experimental Test Article (THETA) is a facility that is used to develop sodium components and instrumentation as well as to acquire experimental data for validation of reactor thermal hydraulic and safety analysis codes. The facility simulates nominal thermal hydraulic conditions as well as protected/unprotected loss of flow accidents in a sodium-cooled fast reactor (SFR). High fidelity distributed temperature profiles of the developed flow field may be acquired with Rayleigh backscatter based optical fiber temperature sensors. The facility was designed in partnership with systems code experts to tailor the experiment to ensure the most relevant and highest quality data for code validation. THETA is comprised of a traditional primary coolant and secondary coolant system. The primary system is submerged in the pool of sodium and consists of a pump, electrically heated core, intermediate heat exchanger, and connected piping and thermal barriers (redan). The secondary system, located outside of the sodium pool, consists of a pump, sodium to air heat exchanger, and connected piping and valves. In fiscal year 2023, thermal stratification tests were completed with the primary system online, while the secondary system was being constructed [1]. These tests had shown that the core barrel and intermediate heat exchanger (IHX) outlet required increased thermal insulation. The THETA primary system was removed from METL, cleaned, thermal insulators installed, and then inserted into METL Test Vessel 4. At the time of this writing the THETA primary and secondary system are operational. During this fiscal year 100+ hours of testing was completed to characterize thermal hydraulic phenomena associated with steady state and transient conditions in a pool type liquid metal cooled reactor. A majority of the testing campaign was completed to satisfy the experimental data acquisition requirements for the GAIN Voucher with Oklo, CRADA 2021-21121. THETA is still operational at the time of this publication and future testing is planned for fiscal year 2025. Work is underway to publish existing and future data to an online database to facilitate collaboration with SFR engineers looking to validate their systems code or computational fluid dynamics models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Radiation Effects on Network on Chips (NoC) Laboratory Directed Research and Development (LDRD) project

This project was motivated by State-of-the-Art (SOTA) technology that incorporates Network on Chips (NOC) for efficient data communication across the various computer kernels. For example, on the AMD Versal Field Programmable Gate Arrays (FPGA), an NoC has been incorporated for fast data communication from the programmable logic and other computer kernels (processing system, adaptable intelligence engines, etc.). The radiation effects on the legacy technology of this FPGA, such as the programmable logic, are well understood, and established methods exist to measure cross-sections when new families/generations are released; however, newly incorporated technologies, such as the NoC, are not fully understood and could introduce new failure points into the mission space.

36 MATERIALS SCIENCE↗

Quantum-Inspired Bayesian Sampling for Uncertainty Quantification and Machine Learning (Final Technical Report)

With increasing simulation and measurement data, machine learning and artificial intelligence have been widely used in computational decision-making of complex engineering systems. The resulting tools, such as uncertainty quantification solvers, reinforcement learning, and physics-informed machine learning, have achieved great success in critical DOE tasks such as material discovery and design, energy system modeling and control, and numerical weather and climate prediction. A core topic in scientific machine learning and artificial intelligence is Bayesian inference: given an observed data set, people want to estimate the posterior distribution of a (possibly large) number of hidden parameters. Due to the flexibility and weak assumptions, Bayesian sampling has been the mainstream Bayesian inference solvers despite the rapid progress of approximate Bayesian inference. Classical Bayesian sampling methods such as Markov-chain Monte Carlo suffer from a low-acceptance rate due to the random walk nature, therefore state-of-the-art techniques use Hamiltonian Monte Carlo and its variants to efficiently draw posterior samples in a high dimension. The key idea of Hamiltonian Monte Carlo and its variants is to simulate the Hamiltonian dynamics of a classical particle with a fixed mass, and their performance significantly degrades when the posterior distribution is highly spiky or has multiple modes. Leveraging the idea of quantum physics, this project has investigated new theory, algorithms and applications of Bayesian inference (especially Bayesian sampling). The main results include: (1) novel quantum-inspired Bayesian sampling methods that can lead to better accuracy for challenging multi-modal or spiky distributions, (2) more scalable machine learning framework leveraging tensor-compressed Bayesian inference, and (3) Bayesian and sampling approaches for verifying the robustness of continuous and binary neural networks.

97 MATHEMATICS AND COMPUTING↗

WRS Capabilities Booklet [Slides]

WRS is the digital backbone of the Weapons Program—delivering trusted data assets, cyber-assured software and systems, and AI-enabling software—that transform insights into decisive action. We empower physicists, engineers, researchers, and scientists to think faster, act strategically, and stay ahead in an ever-evolving threat landscape. Our efforts ensure critical nuclear weapons data remains secure, accessible, and usable—supporting mission-critical work, informed decision making, and scientific advancement at LANL and across the Nuclear Security Enterprise (NSE).

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Basic Research Needs for Inverse Methods for Complex Systems under Uncertainty

Inverse problems, which aim to infer unknown properties of a system using experimental and observational data, are central to addressing many of the U.S. Department of Energy’s (DOE) most critical scientific and engineering challenges. Accurate, computationally efficient, and data-efficient solutions to inverse problems are essential for advancing DOE mission-critical science drivers, including analyzing data from large-scale experimental facilities, optimizing fusion reactor performance, accelerating materials discovery, enhancing geophysical imaging, improving wildfire predictions, and enabling autonomous systems and digital twins. However, these problems are becoming increasingly complex, often involving nonlinear, highdimensional, and interconnected systems and models that span multiple physics and scales, while relying on data with varying quantity, quality, and information content. Compounding these challenges is the uncertainty inherent in DOE-relevant systems, where errors in inputs, noise in data, incompleteness of data, and discrepancies between models and reality constrain the accuracy and precision of solutions. At the same time, the convergence of recent scientific computing trends—scientific machine learning, artificial intelligence, and computing advances such as exascale computing—is creating unprecedented opportunities for tackling these challenges. The cross-cutting nature of inverse problems, combined with their growing complexity and rapidly evolving data and algorithmic demands, strongly motivates the formulation of a prioritized research agenda to maximize their capabilities and impact. In response to this need, DOE’s Advanced Scientific Computing Research (ASCR) program in the Office of Science convened the Workshop on Basic Research Needs for Inverse Problems for Complex Systems Under Uncertainty in June 2025. This workshop brought together experts across disciplines to identify grand challenges and major opportunities in the field. Through collaborative discussions, the workshop defined transformative research directions aimed at addressing the mathematical, statistical, and computational challenges posed by inverse problems under uncertainty. As a result of these efforts, four priority research directions (PRDs) were identified to guide future research and development in this area. These PRDs, summarized below, represent a roadmap for advancing the foundational science and mathematics of inverse problems, enabling robust, scalable, and uncertainty-aware solutions that are critical for DOE applications.

97 MATHEMATICS AND COMPUTING↗

Reliable statistics-based detection and investigation of anomalies in a SMART valve system

Reliable anomaly detection and diagnosis are critical for the safe operation of complex engineered systems. This study presents a unified framework that integrates statistical, model-based, and data-driven techniques for anomaly detection and investigation, demonstrated on SMART valve systems in hybrid energy applications. Four detection methods—mean deviation, seasonal extreme studentized deviate, ARIMA forecasting, and matrix profiling—were implemented and compared. Matrix profiling was particularly effective in revealing subtle deviations and hidden relationships among variables. Anomaly investigation was performed by analyzing variable-level and grouped signal profiles, with system topology incorporated to distinguish primary faults from propagated effects. Grouping signals by type enhanced interpretability, enabling accurate localization of anomalies across multi-dimensional datasets. Experimental results confirmed the framework's capability to consistently detect and isolate anomalies while providing actionable insights into system interdependencies. The proposed methodology offers a robust, interpretable, and scalable solution for condition monitoring, with potential applications in safety-critical domains such as nuclear energy, aerospace, and process industries.

ARIMA models↗

2010-2012 California Household Travel Survey

The 2010-2012 California Household Travel Survey (CHTS) was administered by the California Department of Transportation, which collected demographic and travel behavior characteristics for residents across the entire state. At the time, it was the largest such regional or statewide survey ever conducted in the United States. Detailed travel behavior information was obtained from more than 42,500 households via multiple data-collection methods, including computer-assisted telephone interviewing, online and mail surveys, wearable (7,574 participants) and in-vehicle (2,910 vehicles) global positioning system devices, and on-board diagnostic sensors that gathered data directly from a vehicle's engine. Details of personal travel behavior were gathered within the region of residence, inter-regionally within the state, and in adjoining states and Mexico. The survey sampling plan was designed to ensure an accurate representation of the entire population of the state. The CHTS included additional features, such as vehicle-acquisition decisions, parking choices, work schedules and flexibility, use of toll lanes/priced facilities, and walk and bicycle trips, to support advanced model development.

1Hz data↗