Search NASA⌕ Search

SEARCH · Search NASA

Results for “data usage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

From roads to roofs: How urban and rural mobility influence building energy consumption

In this article, understanding the relationship between travel behavior and building energy use at an urban scale is crucial for developing effective energy management strategies. Mobility patterns significantly impact building occupancy, which in turn affects energy consumption. However, existing methods often focus on individual buildings, whereas geographical influences on energy usage are not adequately examined. This study addresses this gap by using transportation origin-destination (OD) data to estimate building occupancy and energy. The proposed method assigns OD trips from census block groups to the building level, incorporating building, travel survey, and census data to derive building occupancy profiles. This method was applied to urban and rural areas with 4062 buildings in 70 census block groups. We found that the OD-informed occupancy profile exhibits smoother energy consumption patterns compared with that of Department of Energy reference occupancy profiles. Our analysis reveals distinct building energy consumption patterns among groups with long and short commutes, emphasizing the effect of commute times and work schedules on residential energy usage. This framework is useful for practitioners in transportation agencies and utility companies, enabling the estimation of building energy based on mobility patterns. Overall, this study shows the potential of integrating transportation and building energy data to inform cross-sector energy management strategies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The critical importance of software for HEP

Particle physics has an ambitious and broad global experimental programme for the coming decades. Large investments in building new facilities are already underway or under consideration. Scaling the present processing power and data storage needs by the foreseen increase in data rates in the next decade for HL-LHC is not sustainable within the current budgets. As a result, a more efficient usage of computing resources is required in order to realise the physics potential of future experiments. Software and computing are an integral part of experimental design, trigger and data acquisition, simulation, reconstruction, and analysis, as well as related theoretical predictions. A significant investment in computing and software is therefore critical. Advances in software and computing, including artificial intelligence (AI) and machine learning (ML), will be key for solving these challenges. Making better use of new processing hardware such as graphical processing units (GPUs) or ARM chips is a growing trend. This forms part of a computing solution that makes efficient use of facilities and contributes to the reduction of the environmental footprint of HEP computing. The HEP community already provided a roadmap for software and computing for the last EPPSU, and this paper updates that, with a focus on the most resource critical parts of our data processing chain.

97 MATHEMATICS AND COMPUTING↗

Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization

Domain-specific visualizations sometimes focus on narrow, albeit important, tasks for one group of users. This focus limits the utility of a visualization to other groups working with the same data. While tasks elicited from other groups can present a design pitfall if not disambiguated, they also present a design opportunity—namely, the development of visualizations that support multiple groups. This development choice presents a trade-off of broadening the scope but limiting support for the more narrow tasks of any one group, which in some cases can enhance the overall utility of the visualization. We investigate this scenario through a design study where we develop Guidepost, a notebook-embedded visualization of data that helps scientists assess compute wait times, machine learning researchers understand prediction accuracy, and system maintainers analyze usage trends. We adapt the use of personas for visualization design from existing literature in the HCI and design domains, applying them to categorize tasks based on their uniqueness across stakeholder personas. Under this model, tasks shared between all groups should be supported by interactive visualizations and tasks unique to each group can be deferred to scripting with notebook-embedded visualization design. We evaluate our visualization through real-world case studies and a task-focused evaluation with nine participants. We observe that together, Guidepost's visual encodings, interactions, and export capabilities support the tasks of our differing personas.

97 MATHEMATICS AND COMPUTING↗

ComStock Measure Documentation: High-Efficiency Rooftop Unit

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - high-efficiency rooftop unit (RTU) - and briefly introduces key results. The full public data set can be accessed on the Comstock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Development of a Distribution Optimal Power Flow Federate for Open-Source OEDI-SI Platform

Increasing numbers of distributed generators in the electric power distribution networks require developing a control strategy to optimize solutions in real time. Linearized optimal distribution flow development has seen growth and acceptance in the distribution systems literature for efficiently modeling the \glspl{opf} for distribution systems. This paper examines the implementation and integration procedure for linearized optimal distribution flow federate to \gls{oedisi} platform. Specifically, we discuss i) the usage of the \gls{oedisi} platform, ii) obtaining a tractable solution using developed \gls{opf} federate, and iii) validation of solutions and bench-marking the \gls{oedisi} platform with developed \gls{opf} federate using OpenDSS. In brief, we demonstrate how a general linearized optimal distribution flow federate can be developed and integrated with a co-simulation environment to mimic real-world examples. The efficacy of the proposed method is demonstrated using the IEEE 123-bus test system under different scenarios to obtain a tractable solution and compare its results.

Sadnan, Rabayet↗

On the Abuse and Detection of Polyglot Files

A polyglot is a file that is valid in two or more formats. Polyglot files pose a problem for file-upload and generative AI web interfaces that rely on format identification to determine how to securely handle incoming files. In this work we found that existing file-format and embedded-file detection tools, even those developed specifically for polyglot files, fail to reliably detect polyglot files used in the wild. To address this issue, we studied the use of polyglot files by malicious actors in the wild, finding 30 polyglot samples and 15 attack chains that leveraged polyglot files. Using knowledge from our survey of polyglot usage in the wild---the first of its kind---we created a novel data set based on adversary techniques. We then trained a machine learning detection solution, PolyConv, using this data set. PolyConv achieves a precision-recall area-under-curve score of 0.999 with an F1 score of 99.20% for polyglot detection and 99.47% for file-format identification, significantly outperforming all other tools tested. We developed a content disarmament and reconstruction tool, ImSan, that successfully sanitized 100% of the tested image-based polyglots, which were the most common type found via the survey. Our work provides concrete tools and suggestions to enable defenders to better defend themselves against polyglot files, as well as directions for future work to create more robust file specifications and methods of disarmament.

Oesch, T [ORNL] (ORCID:0000000269091022)↗

Archi: Agentic Operations at the CMS Experiment

We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them. An instance of Archi has been deployed for the Computing Operations team of the CMS experiment at CERN's LHC since February 2026 as a support agent for technical operators, offering retrieval and analysis capabilities by combining documentation, historical data, and live monitoring systems. We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels. The system proves effective at operational tasks, resolving real-world queries posed by CMS operators. We also observe that locally-hosted, open-weight models perform competitively, enabling fully private management of sensitive data.

Lugato, Pietro [MIT; CERN]↗

Atmospheric Radiation Measurement (ARM) airborne field campaign data products between 2013 and 2018

Airborne measurements are pivotal for providing detailed, spatiotemporally resolved information about atmospheric parameters and aerosol and cloud properties, thereby enhancing our understanding of dynamic atmospheric processes. For 30 years, the US Department of Energy (DOE) Office of Science supported an instrumented Gulfstream 1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) Data Center and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated dataset was recently developed covering the final 6 years of G-1 operations (2013 to 2018, https://doi.org/10.5439/1999133; Mei and Gaustad, 2024). The integrated dataset includes data collected from 236 flights (766.4 h), which covered the Arctic, the US Southern Great Plains (SGP), the US West Coast, the eastern North Atlantic (ENA), the Amazon Basin in Brazil, and the Sierras de Córdoba range in Argentina. These comprehensive data streams provide much-needed insight into spatiotemporal variability in the thermodynamic quantities and aerosol and cloud properties for addressing essential science questions in Earth system process studies. This paper describes the DOE ARM merged G-1 datasets, including information on the acquisition, data collection challenges and future potentials, and quality control processes. It further illustrates the usage of this merged dataset to evaluate the Energy Exascale Earth System Model (E3SM) with the Earth System Model Aerosol–Cloud Diagnostics (ESMAC Diags) package.

54 ENVIRONMENTAL SCIENCES↗

Diaspora: Resilience-Enabling Services for Real-Time Distributed Workflows

The need for real-time processing to enable automated decision making and experimental steering has driven a shift from high-performance computing workflows on a centralized system to a distributed approach that integrates remote data sources, edge devices, and diverse compute facilities. Under this paradigm, data can be processed close to the source where it is generated, thus reducing latency and bandwidth usage. System resilience is thus a key challenge, requiring distributed workflows to survive component failures and to meet stringent quality-of-service requirements, which results in the need to mitigate anomalies such as congestion and low availability of resources. To address these challenges, we propose Diaspora, a unified resilience framework that is inspired by event-driven communication patterns used in public clouds. Specifically, we propose an event fabric that extends across sites, facilities, and computations to provide timely, reliable, and accurate information about data, application, and resource status. On top of the event fabric, we build resilience-enabling services that combine QoS-aware data streaming, resilient data views, resilient compute and data resources, and anomaly detection and prediction, all of which collectively enhance workflow resilience for these scientific cases.

Rao, Nageswara↗

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Sierra/SD - Its2Sierra - User's Manual - 5.20

The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within SierraMechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

97 MATHEMATICS AND COMPUTING↗

Sierra/SD – Its2Sierra – User's Manual – 5.22

The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within Sierra Mechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

97 MATHEMATICS AND COMPUTING↗

Achieving Unprecedented CO 2 Utilization InCO 2 Concrete™: System Design, Product Development and Process Demonstration

Anthropogenic sources of carbon dioxide are generated from a number of sources, but the key among these are ordinary Portland cement (OPC) production and combustion of fossil fuels. Cement production is the largest global CO 2 source from the mineral decomposition of carbonates. This is due to the clinkering process whereby limestone (mainly consisting of CaCO 3 ) is decomposed into CaO and CO 2 , and combined with silica rich clays at high temperatures to form clinkers (i.e. the four key minerals that comprise cement). The high temperature range of 1400 – 1550°C required for this process accounts for up to 60% of the generated CO 2 from cement production. Combination of the limestone decomposition and thermal requirements of the clinkering process causes cement production to contribute 8-9% of annual global CO 2 emissions. Combustion of fossil fuels (coal, oil and gas) was shown to contribute a much larger portion of global CO 2 emissions. As of 2018, combustion of fossil fuels accounted for 65% of global CO 2 , where 41% was derived from stationary sources for electricity and heat generation and the other 24% was related to transport. To reduce these contributions, key steps forward in CO 2 utilization technologies are required. Therefore, a CO 2 mineralization technology (CO 2 mineralization concrete) to reduce the OPC content in concrete, while utilizing flue gas emissions from fossil fuel combustion has been developed to address both areas simultaneously. This Reversa™ technology utilizes low-carbon cementation agents produced by in situ CO 2 mineralization (“mineral carbonation reactions”) to offer a promising alternative to OPC. CO 2 mineralization relies upon the reaction of dissolved CO 2 with inorganic alkaline reactants to precipitate mineral carbonates (e.g., CaCO 3 ), which bind proximate particles and achieve cementation. Herein, a concrete green body, which is composed of a mixture of binder, water, and mineral aggregates, is exposed to CO 2 borne in industrial flue gas streams. This manner of CO 2 mineralization allows the production of construction components that feature equivalent engineering attributes as their OPC-based counterparts while featuring a much smaller embodied carbon intensity (eCI). The purpose of this project is to demonstrate the feasibility of the Reversa process evolving from a TRL-3 technology at the bench-scale up to TRL-6 technology at the pilot-scale. The reliability of the Reversa technology was tested to prove the effective production of three standard industrial concrete products selected during the course of the project. The results detailed herein will demonstrate the evolution of this technology to the industrial scale. The culmination of this work resulted in 9 production runs completed at the National Carbon Capture Center (NCCC), Wilsonville, AL, using natural gas (NG) flue gas as the CO 2 source. Over the course of the production runs at NCCC, the CO 2 utilization as a function of time, 24-h CO 2 uptake, electricity usage, and 28-d net area compressive strength recorded for each run. Collection of this data will be used to determine the success of the demonstration goals: (1) achieving in excess of 0.2gCO 2 /g reactant , (2) achieving greater than 50% reduction in global warming potential compared to standard produced units, and (3) ensuring compliance of carbonated concrete with industry standard specifications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sierra/SD – Its2Sierra – User’s Manual (V.5.24)

The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within Sierra Mechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

97 MATHEMATICS AND COMPUTING↗

Sierra/SD – Its2Sierra – User’s Manual – (V.5.26)

The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within Sierra Mechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Sierra/SD – Its2Sierra – User’s Manual – 5.28

Incrementally updated theory documents associated with new software release. The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within Sierra Mechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

97 MATHEMATICS AND COMPUTING↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗