Search NASA⌕ Search

SEARCH · Search NASA

Results for “data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Ring Pull Strain Analysis Version 1.1

This report details an analysis package, Ring Pull Strain Analysis (RPSA), that can be used to present and quantify digital image correlation (DIC) data as it relates to a gaugeless ring pull test. Gaugeless ring pull is a testing technique for mechanical testing of small annular samples, usually cut from a thin-walled tube. DIC data is often necessary for this kind of test because bending moments present on the ring cause a non-uniform strain distribution and localized measurements are necessary. In addition, the annular geometry of a ring lends itself to a polar representation, which is not present with typical DIC analysis methods. RPSA was made to calculate and plot the polar representation of strain from standard pre-processed DIC data of a gaugeless ring pull test. Further analysis can be done on ring pull including a quasi-uniaxial tensile analysis and coating analysis, which are also performed by RPSA. In addition, due to the universality of DIC plotting and ring pull test analysis, RPSA can accommodate a wide variety of tests, though it is tailored for ring pull testing. This report details how RPSA works, including the theory, assumptions, and logic behind the calculations and the structure of the program.

36 MATERIALS SCIENCE↗

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES↗

Model-agnostic likelihood for the reinterpretation of the 𝐵 + → 𝐾 + ⁢$𝑣\bar{𝑣}$ measurement at Belle II

We recently measured the branching fraction of the 𝐵 + → 𝐾 + ⁢$𝑣\bar{𝑣}$ decay using 362 fb −1 of on-resonance 𝑒 + ⁢𝑒 − collision data under the assumption of Standard Model kinematics, providing the first evidence for this decay. To facilitate future reinterpretations and maximize the scientific impact of this measurement, we publicly release the full analysis likelihood along with all necessary material required for reinterpretation under arbitrary theoretical models sensitive to this measurement. In this work, we demonstrate how the measurement can be reinterpreted within the framework of the weak effective theory. Using a kinematic reweighting technique in combination with the published likelihood, we derive marginal posterior distributions for the Wilson coefficients, construct credible intervals, and assess the goodness of fit to the Belle II data. For the weak effective theory Wilson coefficients, the posterior mode of the magnitudes |𝐶 VL +𝐶 VR |, |𝐶 SL +𝐶 SR |, and |𝐶 TL | corresponds to the point (11.3, 0.0, 8.2). The respective 95% credible intervals are [1.9, 16.2], [0.0, 15.4], and [0.0, 11.2].

bottom quark↗

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning for Automated Weld Quality Monitoring and Control

Resistance Spot Welding (RSW) is a critical process in the automotive industry, valued for its cost-effectiveness, short cycle time, and robustness. However, achieving consistent high-quality joints remains challenging due to the complex interplay of various factors, like materials, processes, and manufacturing uncertainties, etc. Under the collaborative project between Oak Ridge National Laboratory (ORNL) and General Motors (GM), we have developed a robust and expansible machine learning (ML) framework aimed at enhancing quality control in RSW. By harnessing the power of machine learning, we have developed the ability to ensure every aspect of the welding process, from the initial process design stage to the final weld joint quality. The framework operates by analyzing a variety of data streams, including in-line process signals, process parameters, materials, and postprocessed weld joint data. Through this analysis, the models have been trained to detect deviations from optimal quality standards, leveraging their ability to identify signature data patterns and anomalies within in-line signals and construct complex correlations between these signals and weld quality parameters. Meanwhile, the machine learning framework is designed to adapt to a variety of materials, including high strength steels and aluminum alloys, etc. Its flexible architecture facilitates the incorporation of diverse data sources and features, enabling precise modeling and prediction across a broad range of material properties and weld quality variables. The expansible ML frameworks represent a promising transformation in weld quality monitoring and control, empowering industry to achieve high levels of efficiency, consistency, and reliability in manufacturing.

99 GENERAL AND MISCELLANEOUS↗

Machine Learning for Automated Weld Quality Monitoring and Control

Resistance Spot Welding (RSW) is a critical process in the automotive industry, valued for its cost-effectiveness, short cycle time, and robustness. However, achieving consistent high-quality joints remains challenging due to the complex interplay of various factors, like materials, processes, and manufacturing uncertainties, etc. Under the collaborative project between Oak Ridge National Laboratory (ORNL) and General Motors (GM), we have developed a robust and expansible machine learning (ML) framework aimed at enhancing quality control in RSW. By harnessing the power of machine learning, we have developed the ability to ensure every aspect of the welding process, from the initial process design stage to the final weld joint quality. The framework operates by analyzing a variety of data streams, including in-line process signals, process parameters, materials, and postprocessed weld joint data. Through this analysis, the models have been trained to detect deviations from optimal quality standards, leveraging their ability to identify signature data patterns and anomalies within in-line signals and construct complex correlations between these signals and weld quality parameters. Meanwhile, the machine learning framework is designed to adapt to a variety of materials, including high strength steels and aluminum alloys, etc. Its flexible architecture facilitates the incorporation of diverse data sources and features, enabling precise modeling and prediction across a broad range of material properties and weld quality variables. The expansible ML frameworks represent a promising transformation in weld quality monitoring and control, empowering industry to achieve high levels of efficiency, consistency, and reliability in manufacturing.

42 ENGINEERING↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Review of searches for vector-like quarks, vector-like leptons, and heavy neutral leptons in proton–proton collisions at $\sqrt{s} = 13$ TeV at the CMS experiment

The LHC has provided an unprecedented amount of proton–proton collision data, bringing forth exciting opportunities to address fundamental open questions in particle physics. These questions can potentially be answered by performing searches for very rare processes predicted by models that attempt to extend the standard model of particle physics. The data collected by the CMS experiment in 2015–2018 at a center-of-mass energy of 13 TeV can be used to test the standard model with high precision and potentially uncover evidence for new particles or interactions. An interesting possibility is the existence of new fermions with masses ranging from the MeV to the TeV scale. Such new particles appear in many possible extensions of the standard model and are well motivated theoretically. New fermions may explain the appearance of three generations of leptons and quarks, the mass hierarchy across these generations, and the nonzero neutrino masses. In this report, the results of searches targeting vectorlike quarks, vector-like leptons, and heavy neutral leptons at the CMS experiment are summarized. The complementarity of current searches for each type of new fermion is discussed, and combinations of several searches for vector-like quarks are presented. The discovery potential for some of these searches at the High-Luminosity LHC is also discussed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Quantifying and Optimizing the Energy Benefits of Mass Timber Construction

The International Mass Timber Alliance (IMTA) is a global organization of industry leaders, engineers, scientists, and associations dedicated to advancing mass timber construction. Its mission is to generate and disseminate scientific data supporting the development of standardized construction and energy efficient practices that promote the adoption of mass timber worldwide. IMTA collaborated with Oak Ridge National Laboratory (ORNL) to leverage ORNL’s expertise in building envelope modeling and testing to evaluate how mass timber construction can reduce peak heating and cooling demand, lower overall energy use, and improve resilience during power outages. A previous study of 80 mass timber buildings in Finland found measured energy use up to 50% lower than predicted by simulation. This project aimed to validate and extend those findings for U.S. buildings through analytical modeling, laboratory testing, and full-scale building evaluations. The research focused on the thermal performance of low-embodied-energy wall assemblies, such as cross-laminated timber (CLT) panels and log walls, with particular attention to the effects of thermal inertia on indoor comfort and energy performance. While mass timber’s structural and fire-resistance properties are well documented, its whole-building thermal behavior has received limited attention. Field data, simulation results, and resilience testing from this study will inform future modeling practices, design guidelines, and construction practices by quantifying the unique thermal and demand-flexibility benefits of mass timber construction.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The lipidomics reporting checklist a framework for transparency of lipidomic experiments and repurposing resource data

The rapid increase in lipidomic studies has led to a collaborative effort within the community to establish standards and criteria for producing, documenting, and disseminating data. Creating a dynamic checklist that condenses key information about lipidomic experiments into common terminology will enhance the field's consistency, comparability, and repeatability. Here, we describe the structure and rationale of the established Lipidomics Minimal Reporting Checklist to increase transparency in lipidomics research.

59 BASIC BIOLOGICAL SCIENCES↗

Catalyzing deep decarbonization with federated battery diagnosis and prognosis for better data management in energy storage systems

Industrial data analytics methods play a central role in improving energy storage performance and efficiency, impacting the future of electrified transportation and renewable electricity generation. However, significant challenges hinder the large-scale deployment of batteries. Conventional methods rely on centralized collection and processing of fleet-level data, leading to database size issues and privacy concerns due to potential data breaches. To enable scalable deployment of battery management systems, this article proposes a federated battery diagnosis and prognosis model, which distributes the processing of battery standard current-voltage-time-usage data in a privacy-preserving manner. Instead of transferring the raw data, this approach communicates only the locally processed parameters, thus reducing communication load and preserving data confidentiality. The federated model offers a paradigm shift in battery health management through privacy-preserving distributed methods for battery data processing and lifetime prediction, ensuring the reliable and sustainable deployment of lithium-ion batteries in a rapidly evolving world.

asset health management↗

Enriching the physics program of the CMS experiment via data scouting and data parking

Specialized data-taking and data-processing techniques were introduced by the CMS experiment in Run 1 of the CERN LHC to enhance the sensitivity of searches for new physics and the precision of standard model measurements. These techniques, termed data scouting and data parking, extend the data-taking capabilities of CMS beyond the original design specifications. The novel data-scouting strategy trades complete event information for higher event rates, while keeping the data bandwidth within limits. Data parking involves storing a large amount of raw detector data collected by algorithms with low trigger thresholds to be processed when sufficient computational power is available to handle such data. The research program of the CMS Collaboration is greatly expanded with these techniques. The implementation, performance, and physics results obtained with data scouting and data parking in CMS over the last decade are discussed in this Report, along with new developments aimed at further improving low-mass physics sensitivity over the next years of data taking.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

United States Nuclear Power Reactor Used Nuclear Fuel Database and Applications

The Unified Database (UDB) within STANDARDS serves as the foundational data infrastructure for managing the United States' spent nuclear fuel inventory of 315,111 discharged assemblies totaling 91,036 metric tons of heavy metal. The database organizes this complex inventory through over 200 interconnected tables structured into eight primary attribute categories, supporting integrated analyses across storage, transportation, and disposal domains. Data enters the UDB through the GC-859 Nuclear Fuel Data Survey, which transitioned to web-based collection in 2023, improving data quality through real-time validation. The UDB enables automated generation of input files for nuclear safety analyses, reducing preparation time from weeks to hours while maintaining traceability. Applications include national inventory reporting, Certificate of Compliance assessments, and facility optimization. The three-tier distribution model balances accessibility with security requirements for federal agencies, national laboratories, and research organizations. The UDB provides essential data infrastructure as spent fuel management transitions from site-specific to integrated national campaigns.

Stefanovic, Peter↗

Deep-learning-based canopy height model generation from sub-meter resolution panchromatic satellite imagery

Canopy height models (CHMs) with sufficient resolution to distinguish individual trees are useful for a variety of applications. However, standard techniques to acquire such data, such as airborne lidar surveying, are often prohibitively expensive. Deep learning techniques for generating CHMs from high-resolution imagery are an attractive option to reduce costs. To date, success with these methods has been demonstrated using multichannel aerial photography and specialized satellite data products derived from multiple sensors, neither of which is commonly available at temporal resolutions finer than one year. Here we demonstrate a method to generate sub-meter resolution CHMs in three forests in California using a more abundant data source: sub-meter resolution, panchromatic satellite imagery from a single sensor. We show that phenology and species composition play important roles in model transferability; when trained using imagery from a single conifer forest in autumn, the model performs well on autumn imagery from a second conifer forest several hundred kilometers distant with no re-training. With modest additions to the training dataset, the same model generates minimally biased estimates of canopy height in both conifer and deciduous forests during multiple seasons. Because the model operates on satellite data with global coverage and a relatively short return interval, we propose its suitability to extrapolate tree-level canopy height data to remote regions and conduct high-temporal resolution monitoring of forest structure. We furthermore demonstrate the workflow’s applicability to fire modeling by conducting simulations in forests populated by trees measured using both this approach and airborne lidar surveying. We find minimal differences in fire behavior relative to a baseline case in which only statistical distributions of tree height and crown area are known. This result underscores the value of forest structural information derived from our workflow for improving the fidelity of wildland fire simulations, among other ecological applications.

54 ENVIRONMENTAL SCIENCES↗

Search for charged-lepton flavor violation in the production and decay of top quarks using trilepton final states in proton-proton collisions at $\sqrt{s}$ =13 TeV

A search is performed for charged-lepton flavor violating processes in top quark (𝑡) production and decay. The data were collected by the CMS experiment from proton-proton collisions at a center-of-mass energy of 13 TeV and correspond to an integrated luminosity of 138 fb −1 . The selected events are required to contain one opposite-sign electron-muon pair, a third charged lepton (electron or muon), and at least one jet of which no more than one is associated with a bottom quark. Boosted decision trees are used to distinguish signal from background, exploiting differences in the kinematics of the final states particles. The data are consistent with the standard model expectation. Upper limits at 95% confidence level are placed in the context of effective field theory on the Wilson coefficients, which range between 0.024–0.424 TeV −2 depending on the flavor of the associated light quark and the Lorentz structure of the interaction. These limits are converted to upper limits on branching fractions involving up (charm) quarks, 𝑡 → 𝑒⁢𝜇⁢𝑢 (𝑡 → 𝑒⁢𝜇⁢𝑐), of 0.032⁢(0.498) × 10 −6 , 0.022⁢(0.369) × 10 −6 , and 0.012⁢(0.216) × 10 −6 for tensorlike, vectorlike, and scalarlike interactions, respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗