Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

APS upgrade: Commissioning the world’s first light source based on swap-out injection

The Advanced Photon Source (APS) has recently completed a major upgrade, replacing its 25-year-old storage ring with a cutting-edge hybrid seven-bend achromat lattice enhanced by six additional reverse bends. The new design achieves a natural emittance of 42 pm-rad, enabling the production of X-rays up to 500 times brighter than those generated by the original APS. A key innovation of the upgrade is the implementation of a swap-out injection scheme, which replaces entire depleted bunches instead of performing traditional top-up injection. This approach enables on-axis injection to accommodate for the reduced dynamic aperture resulting from strong focusing. This paper outlines the commissioning process, shares initial operating experience with swap-out injection, and presents performance data for new systems such as the bunch-lengthening cavity.

Sajaev, Vadim [Argonne, PHY]↗

District-Scale Analysis of Electricity Load and Strategies to Improve Energy Reliability Using Prototype District Models

Projected increases in electricity demand in the U.S. highlight the urgent need for effective load management to ensure grid reliability. As the building sector accounts for approximately 75% of electricity usage, enhancing energy efficiency and flexibility in this sector is crucial. Adopting district-level approaches offers significant advantages over traditional individual building analyses by enabling shared infrastructure and economies of scale. To navigate the data and computational challenges associated with modeling energy at the district level, prototype district models have been proposed as holistic, system-level solutions that capture complex interactions within typical configurations. This study presents these models as a reference tool for analyzing district-scale energy systems across various climate zones in the U.S. Developed with input from stakeholders, these models integrate varied building characteristics, inter-building connections, and energy system interactions. A case study utilizing the Urban Edge prototype district model, implemented on the URBANopt™ platform, evaluates multiple demand scenarios and the impact of distributed energy resources such as fuel-fired backup generators, photovoltaic systems, and batteries. Findings suggest that while new electric systems can significantly reduce annual energy use, they may also elevate peak electricity loads, with a notable 43% increase in heating-dominant climate zone 5B. The optimal backup power solutions vary based on location, influenced by factors such as utility rates and incentives. For example, PV and batteries perform well in high-cost regions like New York City, while diesel backup generators are more suitable for backup needs in climate zone 3A, such as Atlanta. Thus, this research highlights the importance of prototype district models for future district-scale energy planning.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An Efficient Checkpointing System for Large Machine Learning Model Training

As machine learning models increase in size and complexity rapidly, the cost of checkpointing in ML training became a bottleneck in storage and performance (time). For example, the latest GPT-4 model has massive parameters at the scale of 1.76 trillion. It is highly time and storage consuming to frequently writes the model to checkpoints with more than 1 trillion floating point values to storage. This work aims to understand and attempt to mitigate this problem. First, we characterize the checkpointing interface in a collection of representative large machine learning/language models with respect to storage consumption and performance overhead. Second, we propose the two optimizations: i) A periodic cleaning strategy that periodically cleans up outdated checkpoints to reduce the storage burden; ii) A data staging optimization that coordinates checkpoints between local and shared file systems for performance improvement.

machine learning, artificial intelligence↗

2024 NMDC Ambassador Training Materials [Slides]

The NMDC is a sustainable data discovery platform that promotes open science and shared-ownership across a broad and diverse community of researchers, funders, publishers, societies, and other collaborators. The NMDC aims to enable multi-omic microbiome research to accelerate scientific discovery. The NMDC is a Department of Energy funded program that is a collaboration between 3 National Laboratories: Lawrence Berkeley National Laboratory (LBNL), Los Alamos National Laboratory (LANL), and Pacific Northwest National Laboratory (PNNL).

54 ENVIRONMENTAL SCIENCES↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY25-Q4

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY26-Q1

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

VA Community Determinants of Health Data Curation Documentation FY26-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1 km grid) to another (e.g., U.S. Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, U.S. Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1 km grids. Some economic data may only be available at the ZIP code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., U.S. Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Shared Use Travel Behavior for Improving Rural Mobility: Insights from Greene County, Pennsylvania

Rural communities are considered disadvantaged communities as they suffer from a lack of transport options. Thus, rural regionsprovide less accessibility for commuters to reach their destination as opposed to urban regions. However, the issues of transport disadvantageand shared use mobility in rural areas within the United States (US) have not been well investigated. Furthermore, transport disadvantagediffers between communities and regions across the globe; thus, there is a need to study the behavioral choices of rural commuters within theUS context. This study contributes by analyzing the behavioral choices of rural communities within the US through a case study site ofWaynesburg, Pennsylvania, for adopting a shared use shuttle service. K-means clusters showed that trips from the survey data were a goodrepresentation of real trips from Ecolane. Furthermore, random parameter-based binary logit models were calibrated using data collected fromstudents, faculty, and residents in Waynesburg, Greene County, to study the behavioral choices of commuters. The findings for the faculty andstudents group revealed that prior experience with shared services increases the likelihood of using a shared shuttle. An important personalcharacteristic of inconvenience showed a higher propensity toward using existing modes as opposed to a shared shuttle. Such commutersvalue personal vehicles as more convenient as they have childcare responsibilities and varying schedules for work that require them to moveback and forth across locations, thus making a shared shuttle less attractive for them. The socioeconomic factors of age and gender show ahigher propensity for using shared shuttles. Furthermore, the findings from this study could be helpful for agencies in improving rural mobility andconsidering such shared mobility services for rural communities

42 ENGINEERING↗

An Integral Activity-Based Protein Profiling Method for Higher Throughput Determination of Protein Target Sensitivity to Small Molecules

Activity-based protein profiling (ABPP) is a chemoproteomic technique that uses small molecule probes to label active enzymes selectively and covalently in complex proteomes. Competitive ABPP, which involves treatment of the active proteome with an analyte of interest, is especially powerful for profiling how small molecules impact specific protein activities. Advances in higher throughput workflows have made it possible to generate extensive competitive ABPP data across diverse biological samples, making this approach highly appealing for characterizing shared and unique proteins affected by perturbations such as drug or chemical exposures. To use the competitive ABPP approach effectively to understand potential adverse effects of chemicals of concern (CoC), a wide range of concentrations may be needed, particularly for chemicals that lack potency or toxicity data. In this work, we present an integral competitive ABPP method that enables target sensitivity determination for different organophosphate (OP) pesticides as model toxicants. Using previously developed OP-ABPs, we optimized conditions for tandem mass tag (TMT) multiplexing of ABPP samples and compared conventional competitive ABPP involving samples at discrete paraoxon concentrations to pooled samples across that same concentration range. We then expanded our approach to compare protein target sensitivities toward two additional OP pesticides, chlorpyrifos oxon and malaoxon. The results showed that differences in integral intensities for the pooled competition sample can be used to evaluate the relative sensitivity of specific proteins without increasing the overall number of samples. For 8 CoC concentrations of interest, this strategy reduced the number of TMT plexes and the corresponding number of LC–MS/MS analyses 3-fold. In conclusion, we envision the integral ABPP (IABPP) method will provide a means to screen diverse chemicals more rapidly to identify both high and low sensitivity protein targets.

activity-based probes↗

Cloverleaf Data Artifacts for ArtIMis LDRD

This report summarizes the use of the open-source CloverLeaf/CloverLeaf3D mini-apps to generate synthetic data sets to train foundation models for the ArtIMis LDRD DI. These data artifacts are intended to be used by LANL collaborators and shared externally with our university and institutional partners. Note that CloverLeaf/CloverLeaf3D is not a LANL simulation code.

97 MATHEMATICS AND COMPUTING↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

ReachNow EV Driving Data From Seattle, WA, Portland, OR, and New York, NY

ReachNow provided Idaho National Laboratory (INL) with a dataset describing approximately 49,000 trips taken by customers and employees in approximately 100 BMW i3 EVs operating in ReachNow's free-floating car-sharing fleets in Seattle, WA, Portland, OR, and New York, NY between May 2016 and February 2017. Data fields include vehicle rental period start and end timestamps, the location where vehicles were parked at the start and end of rental periods, and distance driven during rental periods. A field categorizing the user during each rental period is also included. This field makes it possible to identify when vehicles were rented by customers and when vehicles were driven by fleet management team employees to reposition, charge, or service the vehicles.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Ba 4 RuMn 2 O 10 : A Noncentrosymmetric Polar Crystal Structure with Disordered Trimers

Phase-pure polycrystalline Ba 4 RuMn 2 O 10 was prepared and determined to adopt the noncentrosymmetric polar crystal structure (space group Cmc2 1 ) based on results of second harmonic generation, convergent beam electron diffraction, and Rietveld refinements using powder neutron diffraction data. The crystal structure features zigzag chains of corner-shared trimers, which contain three distorted face-sharing octahedra. The three metal sites in the trimers are occupied by disordered Ru/Mn with three different ratios: Ru1:Mn1 = 0.202(8):0.798(8), Ru2:Mn2 = 0.27(1):0.73(1), and Ru3:Mn3 = 0.40(1):0.60(1), successfully lowering the symmetry and inducing the polar crystal structure from the centrosymmetric parent compounds Ba 4 T 3 O 10 (T = Mn, Ru; space group Cmca). The valence state of Ru/Mn is confirmed to be +4 according to X-ray absorption near-edge spectroscopy. Ba 4 RuMn 2 O 10 is a narrow bandgap (~0.6 eV) semiconductor exhibiting spin-glass behavior with strong magnetic frustration and antiferromagnetic interactions.

36 MATERIALS SCIENCE↗