Search NASA⌕ Search

SEARCH · Search NASA

Results for “public availability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗

A publicly available PyTorch-$\mathrm{ABAQUS}$ $\mathrm{UMAT}$ deep-learning framework for level-set plasticity

Here this paper introduces a publicly available PyTorch-ABAQUS deep-learning framework of a family of plasticity models where the yield surface is implicitly represented by a scalar-valued function. In particular, our focus is to introduce a practical framework that can be deployed for engineering analysis that employs a user-defined material subroutine (UMAT/VUMAT) for ABAQUS, which is written in FORTRAN. To accomplish this task while leveraging the back-propagation learning algorithm to speed up the neural-network training, we introduce an interface code where the weights and biases of the trained neural networks obtained via the PyTorch library can be automatically converted into a generic FORTRAN code that can be a part of the UMAT/VUMAT algorithm. To enable third-party validation, we purposely make all the data sets, source code used to train the neural-network-based constitutive models, and the trained models available in a public repository. Furthermore, the practicality of the workflow is then further tested on a dataset for anisotropic yield function to showcase the extensibility of the proposed framework. A number of representative numerical experiments are used to examine the accuracy, robustness and reproducibility of the results generated by the neural network models.

36 MATERIALS SCIENCE↗

Impact Study of Thunderstorms on the US Power Grid Using Publicly Available Datasets

This work analyzes the impact of thunderstorms on the US power grid based on publicly available data. Since thunderstorms can bring lightning, heavy precipitation, and wind storms, analyzing their impact on the power system provides a combined correlation of lightning strikes, floods, and wind storms on power outages. This paper leverages publicly available thunderstorm datasets from the National Weather Service (NWS) and power outage datasets from Oak Ridge National Laboratory’s Environment for Analysis of Geo-Located Energy Information (EAGLE-I) to study the correlation between thunderstorms and power outages. This work is analyzing the patterns of thunderstorms from 2013-2022, which shows that the thunderstorms are not slowing down and will seem to continue their impact on human life in the future. This work also analyzes the monthly and yearly pattern of the impact of thunderstorms on power systems at the national, state, and county level.

Bhusal, Narayan↗

Publicly Available Molten Salt Data/Benchmarks

This presentation summarizes data for three molten salt reactors that has been made publicly available by Copenhagen Atomics. This information will be shared with the participants of the OECD/NEA International Reactor Physics Experiment Evaluation (IRPhE) Technical Review Group.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Deploying a Publicly Available and Living National Oil and Gas Well Geodatabase: CO2-Locate: A Dynamic Database & Tool

Emission mitigation and safe geologic carbon storage require an understanding of local wellbore infrastructure, yet well data are siloed across many entities. Addressing this challenge, the National Energy Technology Laboratory published CO2-Locate, an integrated and dynamic national wellbore geodatabase. CO2-Locate offers up to date well data spanning more than 40 federal, state, and tribal entities, as well as spatially summarized insights designed to support commercial, regulatory, and research communities as they strive to curb climate change through a national energy transition.

Pfander, Isabelle↗

Effect of image resolution on automated classification of chest X-rays

Deep learning (DL) models have received much attention lately for their ability to achieve expert-level performance on the accurate automated analysis of chest X-rays (CXRs). Recently available public CXR datasets include high resolution images, but state-of-the-art models are trained on reduced size images due to limitations on graphics processing unit memory and training time. As computing hardware continues to advance, it has become feasible to train deep convolutional neural networks on high-resolution images without sacrificing detail by downscaling. This study examines the effect of increased resolution on CXR classification performance. We used the publicly available MIMIC-CXR-JPG dataset, comprising 377,110 high resolution CXR images for this study. We applied image downscaling from native resolution to 2048 × 2048 pixels, 1024 × 1024 pixels, 512 × 512 pixels, and 256 × 256 pixels and then we used the DenseNet121 and EfficientNet-B4 DL models to evaluate clinical task performance using these four downscaled image resolutions. We find that while some clinical findings are more reliably labeled using high resolutions, many other findings are actually labeled better using downscaled inputs. We qualitatively verify that tasks requiring a large receptive field are better suited to downscaled low resolution input images, by inspecting effective receptive fields and class activation maps of trained models. Lastly, we show that stacking an ensemble across resolutions outperforms each individual learner at all input resolutions while providing interpretable scale weights, indicating that diverse information is extracted across resolutions.

47 OTHER INSTRUMENTATION↗

AI for Earthquake Physics

The core LANL program sponsored by Office of Science, Basic Energy Science, Chemical Sciences, Geosciences, and Biosciences (DOE-BES-CSGB) and led by PI Johnson aims to research earthquake faults to advance fault physics and earthquake hazards. All work completed is required to be made publicly available through publications and open-source codes supporting the published results. All routines are/will-be written in open source python and applied to publicly available data sets. These routines will format data from input into models, develop and test modeling frameworks for the problems addressed, and produce figures applicable to peer-reviewed manuscripts. All work is reviewed for Los Alamos Unlimited Release before submitting to a journal. This summary encompasses recently completed work and work to be complete for the duration of the program.

Johnson, Christopher↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

Vistransformers Explained

The Vistransformers Explained library is a collection of python notebooks that demonstrate the internal mechanics and uses of visual-transformer (ViT) machine learning models. The code implements, with mild modifications, ViT models that have been made publicly available through publication and GitHub code. The value added by this code is in-depth explanations of the mathematics behind the sub-modules of the ViT models, including original figures. Additionally, the library contains the code necessary to implement and train the ViT models. The library does not include example training data for the models; instead, it would rely on users generating their own datasets. The code is based on the PyTorch python library. It does not include any files other than python scripts, modules, or notebooks.

Callis, Skylar↗

Ground Water Age Predictor

Machine Learning script to predict groundwater ages based on auxiliary features in a publicly available dataset, based on publicly available software libraries. Code applied to data from the Groundwater Ambient Monitoring and Assessment (GAMA) program in California, including well location and construction information, chemical constituents and isotopic tracers, and land use metrics.

Chakraborty, Indrasis↗

Hyperparameter Studies for Vision Transformers Trained on High-Fidelity Simulations

This library is a collection of python modules that define, train, and analyze vision-transformer (ViT) machine learning models. The code implements, with mild modifications, ViT models that have been made publicly available through publication and GitHub code. The training data for these models is hydrodynamic simulation output in the form of numpy arrays. This library contains code to train these ViT models on the hydrodynamic simulation output with a variety of hyperparameters, and to compare the results of such models. Furthermore, the library contains definitions of simple convolutional neural network (CNN) machine learning architectures which can be trained on the same hydrodynamic simulation output. These are included as a reference point to compare the ViT models to. Additionally, the library includes trained ViT and CNN models and example input data for demonstration purposes. The code is based on the PyTorch python library.

Callis, Skylar↗

Monte Carlo Thought Search: Large Language Model Querying for Complex Scientific Reasoning in Catalyst Design

Discovering novel catalysts requires complex reasoning involving multiple chemical properties and resultant trade-offs, leading to a combinatorial growth in the search space. While large language models (LLM) have demonstrated novel capabilities for chemistry through complex instruction following capabilities and high quality reasoning, a goal-driven combinatorial search using LLMs has not been explored in detail. In this work, we present a Monte Carlo Tree Search-based approach that improves beyond state-of-the-art chain-of-thought prompting variants to augment scientific reasoning. We introduce two new reasoning datasets: 1) a curation of computational chemistry simulations, and 2) diverse questions written by catalysis researchers for reasoning about novel chemical conversion processes. We improve over the best baseline by 25.8\% and find that our approach can augment scientist's reasoning and discovery process with novel insights.\footnote{All resources will be publicly available upon publication.

Sprueill, Henry W.↗

Power System Waveform Datasets for Machine Learning

The desire for increased visibility across the electricity grid will necessarily increase the deployment of sensing and measurement devices and associated data management needs to unprecedented levels. For the existing sensing and measurement infrastructure, there remains a great amount of “value” yet to be extracted through advanced data management and analytics. Availability of more data will not, by itself, lead to changes in grid visibility, security, and resiliency. To create the predictive and prescriptive environment required to enable new markets and transactions for customer revenue and a reliable grid, the data must be collected, organized, evaluated, and analyzed using sophisticated algorithms to provide actionable information allowing operators and customers to reliably manage an increasingly complex grid. Progress in artificial intelligence (AI) has been largely driven by large, publicly available datasets that can be used to train AI algorithms such as MNIST, a database of handwritten images of digits, and ImageNet, an image database of everyday objects. These types of publicly available databases of real-world training datasets have been largely credited for advancement of image processing, computer vision, and deep learning algorithms that these use cases deploy. However, in the power systems industry to date, there are few databases with proper event labeling, and data access to a publicly available collection of power system event waveforms that will allow users to interact with grid signature data. Publicly available datasets of power system event waveforms, such as the DOE/EPRI dataset, often lack critical metadata or contain limited examples of each event type, and data formats vary widely across these datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Development of the United States GReenhouse Gas and Air Pollutants Emissions System (GRA 2 PES)

In the U.S., emissions of greenhouse gases and air pollutants are often developed independently. Here, we describe the GReenhouse gas And Air Pollutants Emissions System (GRA 2 PES), which provides gridded emissions of fossil-fuel carbon dioxide (ffCO 2 ) and 93 air quality (AQ) species for 17 combustion and non-combustion sectors at 4 km × 4 km spatial resolution across the contiguous US. We find that the AQ emissions most spatially correlated with ffCO 2 are nitrogen oxides (NO x , ρ = 0.67), followed by sulfur dioxide (SO 2 , ρ = 0.51), carbon monoxide (CO, ρ = 0.44), and fine particulate matter (PM 2.5 , ρ = 0.38). We evaluate GRA 2 PES ffCO 2 emissions with an ensemble of publicly available regional and global inventories at national (Normalized Mean Bias (NMB) = +1.4%), state (NMB = +1.5%, R 2 = 0.98), and urban (NMB = +11.5%, R 2 = 0.97) scales. Nationally, the differences of publicly available inventories from the ensemble average range from −10.0% to +5.7%, and consistency diverges at state and urban scales. We simulate GRA 2 PES ffCO 2 in a particle dispersion model and compare to measurements of radiocarbon ( 14 C)-derived ffCO 2 collected in Los Angeles (August 2021), with results suggesting that GRA 2 PES ffCO 2 may be low by 19% for this city, but well within model-observation differences for other publicly available inventories (−43% to +94%). GRA 2 PES AQ/ffCO 2 ratios converted to concentration space generally agree with field observations (NMB = +4%, log R 2 = 0.90). Lastly, we present a method by which to utilize GRA 2 PES to derive AQ emission fluxes from ffCO 2 emissions.

Lyu, Congmeng [National Oceanic and Atmospheric Ad↗

The cosmic waltz of Coma Berenices and Latyshev 2 (Group X)

Context. Open clusters (OCs) are fundamental benchmarks where theories of star formation and stellar evolution can be tested and validated. Coma Berenices (Coma Ber) and Latyshev 2 (Group X) are the second and third OCs closest to the Sun, making them excellent targets to search for low-mass stars and ultra-cool dwarfs. In addition, this pair will experience a flyby in 10–16 Myr, making it a benchmark to test pair interactions of OCs. Aims. We aim to analyse the membership, luminosity, mass, phase-space (i.e. positions and velocities), and energy distributions for Coma Ber and Latyshev 2 and test the hypothesis of the mixing of their populations at the encounter time. Methods. We developed a new phase-space membership methodology and applied it to Gaia data. With the recovered members, we inferred the phase-space, luminosity, and mass distributions using publicly available Bayesian inference codes. Then, with a publicly available orbit integration code and members’ positions and velocities, we integrated their orbits 20 Myr into the future. Results. In Coma Ber, we identified 302 candidate members distributed in the core and tidal tails. The tails are dynamically cold and asymmetrically populated. The stellar system called Group X is made of two structures: the disrupted OC Latyshev 2 (186 candidate members) and a loose stellar association called Mecayotl 1 (146 candidate members), and both of them will fly by Coma Ber in 11.3 ± 0.5 Myr and 14.0 ± 0.6 Myr, respectively, and each other in 8.1 ± 1.3 Myr. Conclusions. We study the dynamical properties of the core and tails of Coma Ber and also confirm the existence of the OC Latyshev 2 and its neighbour stellar association Mecayotl 1. Although these three systems will experience encounters, we find no evidence supporting the mixing of their populations.

79 ASTRONOMY AND ASTROPHYSICS↗