Search NASA⌕ Search

SEARCH · Search NASA

Results for “science data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Spatial profiling of the interplay between cell type- and vision-dependent transcriptomic programs in the visual cortex

How early sensory experience during “critical periods” of postnatal life affects the organization of the mammalian neocortex at the resolution of neuronal cell types is poorly understood. We previously reported that the functional and molecular profiles of layer 2/3 (L2/3) cell types in the primary visual cortex (V1) are vision-dependent [S. Chenget al.,Cell185, 311–327.e24 (2022)]. Here, we characterize the spatial organization of L2/3 cell types with and without visual experience. Spatial transcriptomic profiling based on 500 genes recapitulates the zonation of L2/3 cell types along the pial–ventricular axis in V1. By applying multitasking theory, we suggest that the spatial zonation of L2/3 cell types is linked to the continuous nature of their gene expression profiles, which can be represented as a 2D manifold bounded by three archetypal cell types. By comparing normally reared and dark reared L2/3 cells, we show that visual deprivation-induced transcriptomic changes comprise two independent gene programs. The first, induced specifically in the visual cortex, includes immediate-early genes and genes associated with metabolic processes. It manifests as a change in cell state that is orthogonal to cell-type-specific gene expression programs. By contrast, the second program impacts L2/3 cell-type identity, regulating a subset of cell-type-specific genes and shifting the distribution of cells within the L2/3 cell-type manifold. Through an integrated analysis of spatial transcriptomics with single-nucleus RNA-seq data, we describe how vision patterns cortical L2/3 cell types during the critical period.

Science & Technology - Other Topics↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Comparative Study of Large Language Model Architectures on Frontier

Large language models (LLMs) have garnered significant attention in both the AI community and beyond. Among these, the Generative Pre-trained Transformer (GPT) has emerged as the dominant architecture, spawning numerous variants. However, these variants have undergone pre-training under diverse conditions, including variations in input data, data preprocessing, and training methodologies, resulting in a lack of controlled comparative studies. Here we meticulously examine two prominent open-sourced GPT architectures, GPT-NeoX and LLaMA, leveraging the computational power of Frontier, the world’s first Exascale supercomputer. Employing the same materials science text corpus and a comprehensive end-to-end pipeline, we conduct a comparative analysis of their training and downstream performance. Our efforts culminate in achieving state-of-the-art performance on a challenging materials science benchmark. Furthermore, we investigate the computation and energy efficiency, and propose a computationally efficient method for architecture design. To our knowledge, these pre-trained models represent the largest available for materials science. Our findings provide practical guidance for building LLMs on HPC platforms.

Yin, Junqi↗

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science

This presentation explores the transformative potential of autonomous electron microscopy and artificial intelligence (AI) in accelerating materials science discovery, particularly for energy applications and materials operating in extreme environments. We discuss pioneering self-driving laboratories at NREL designed to intelligently probe material synthesis and degradation across multiple scales, aiming to rapidly bridge the gap between atomic-level understanding and the development of high-performance, reliable materials. Utilizing advanced machine learning techniques, such as few-shot learning and multimodal analysis integrating imaging and spectroscopy, we demonstrate methods to extract actionable descriptors for material behavior, quantify complex microstructural evolution, and statistically link synthesis parameters to defect populations. This AI-driven approach promises to accelerate the creation of predictive materials tailored for specific missions, enabling faster development cycles and enhanced material assurance.

36 MATERIALS SCIENCE↗

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science

This presentation explores the transformative potential of autonomous electron microscopy and artificial intelligence (AI) in accelerating materials science discovery, particularly for energy applications and materials operating in extreme environments. We discuss pioneering self-driving laboratories at NREL designed to intelligently probe material synthesis and degradation across multiple scales, aiming to rapidly bridge the gap between atomic-level understanding and the development of high-performance, reliable materials. Utilizing advanced machine learning techniques, such as few-shot learning and multimodal analysis integrating imaging and spectroscopy, we demonstrate methods to extract actionable descriptors for material behavior, quantify complex microstructural evolution, and statistically link synthesis parameters to defect populations. This AI-driven approach promises to accelerate the creation of predictive materials tailored for specific missions, enabling faster development cycles and enhanced material assurance.

97 MATHEMATICS AND COMPUTING↗

Identifying Sample Provenance From SEM/EDS Automated Particle Analysis via Few-Shot Learning Coupled With Similarity Graph Clustering

Automated particle analysis (APA) provides a vast amount of compositional data via energy-dispersive X-ray spectroscopy along with size and shape data via scanning electron microscopy for individual particles in a sample. In many instances, APA data are leveraged to support identification of the source of a sample based on the detection of particles of a specific composition. Often, the particles that provide context make up a minuscule portion of the sample. Additionally, the interpretation of complex samples can be difficult due to the diversity of compositions both in the mixture and within a particle. In this work, we demonstrate a method to compute and cluster similarity graphs that describe inter-particle relationships within a sample using a multi-modal few-shot learning neural network. Here, as a proof-of-concept, we show that samples known to have been exposed to gunshot residue can be distinguished from samples occasionally mistaken for gunshot residue. Our workflow builds upon standard APA techniques and data processing methods to unveil additional information in a readily interpretable and quantitatively comparable format.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SOMA: Observability, monitoring, and in situ analytics for exascale applications

With the rise of exascale systems and large, data-centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service-based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application-specific diagnostic data and performance data addresses these needs. Furthermore, we present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi-GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring.

97 MATHEMATICS AND COMPUTING↗

Population-level Dark Energy Constraints from Strong Gravitational Lensing using Simulation-Based Inference

In this work, we present a scalable approach for inferring the dark energy equation-of-state parameter ($w$) from a population of strong gravitational lens images using Simulation-Based Inference (SBI). Strong gravitational lensing offers crucial insights into cosmology, but traditional Monte Carlo methods for cosmological inference are computationally prohibitive and inadequate for processing the thousands of lenses anticipated from future cosmic surveys. New tools for inference, such as SBI using Neural Ratio Estimation (NRE), address this challenge effectively. By training a machine learning model on simulated data of strong lenses, we can learn the likelihood-to-evidence ratio for robust inference. Our scalable approach enables more constrained population-level inference of $w$ compared to individual lens analysis, constraining $w$ to within $1\sigma$.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Accelerating the identification of novel secondary metabolites in bioenergy plant root exudates using MicroED

Small molecule metabolites drive inter- and intraspecies communication and dependencies in diverse biological systems, yet a large proportion of these important chemical compounds remain uncharacterized in plants and microbes. Approximately 90% of the metabolites in root exudate profiles are unknown compounds, despite the importance of root exudate composition in plant-microbe interactions. We need advanced analytical capabilities that will support rapid discovery and structural elucidation of metabolites from biological samples that may be limited in quantity and high in complexity. To fill this gap, this project aimed to develop an integrated workflow involving metabolite extraction, separation, and crystallization from plant root exudates followed by characterization using nuclear magnetic resonance (NMR) spectroscopy, mass spectrometry, and microcrystal electron diffraction (MicroED). Using crude root exudates from sorghum, this project successfully developed higher throughput exudate fractionation strategies to obtain pure compounds for crystallization and identified crystals in multiple fractions that diffracted. Additional efforts to increase the throughput of high-quality crystal generation for MicroED, such as crystallization screening and crystallization chaperone exploration, will be needed to further advance root exudate metabolite identification. The overall optimized sample preparation process can then be integrated with the existing data collection and data analysis pipelines for MicroED at PNNL to facilitate more rapid natural product discovery.

59 BASIC BIOLOGICAL SCIENCES↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

Robust Dark Energy Constraints with the Dark Energy Spectroscopic Survey (Final Technical Report)

This project developed and applied advanced theoretical, computational, and data-analysis methodologies to extract robust and precise cosmological constraints from the Dark Energy Spectroscopic Instrument (DESI). The work focused on maximizing the scientific return of DESI through optimized survey strategy, novel higher-order clustering statistics, improved modeling of small-scale structure, and rigorous mitigation of observational systematics. Over the award period, the project made substantial contributions to DESI science planning, produced new methods for bispectrum and three-point correlation function analyses, advanced constraints on primordial non-Gaussianity, and delivered widely used software tools. The project also played a major role in training graduate students and a postdoctoral researcher who contributed directly to DESI key projects. The results have significantly enhanced the cosmological reach of DESI and provide a strong foundation for future surveys such as DESI-II and Stage-V experiments.

79 ASTRONOMY AND ASTROPHYSICS↗

Tansey_MICRE_D98_microphysics_V5

This data file contains retrievals of cloud-effective-radius and cloud-droplet-number-concentration for low clouds based on the Dong et al. 1998 retrieval technique for the MICRE campaign. The retrieval is based on observed downwelling broadband SW fluxes and microwave-radiometer-retrieved-liquid-water-path averaged on a 5 minute time scale (mean when low cloud is present). Details on the algorithm and analysis of results are given in Tansey et al. 2024 (submitted DOI: 10.22541/essoar.173482310.03809503/v1). The data is limited to the period 20160410 to 20160612 AND 20170103 to 20170214 during which time good quality microwave-radiometer measurements are available.

54 ENVIRONMENTAL SCIENCES↗

Tansey_MICRE_D98_microphysics_V5

This data file contains retrievals of cloud-effective radius and cloud-droplet-number concentration for low clouds based on the Dong et al. 1998 retrieval technique for the MICRE campaign. The retrieval is based on observed downwelling broadband shortwave (SW) fluxes and microwave-radiometer-retrieved-liquid-water-path averaged on a 5-minute time scale (mean when low cloud is present). Details on the algorithm and analysis of results are given in Tansey et al. 2024 (submitted DOI: 10.22541/essoar.173482310.03809503/v1). The data are limited to the period of 20160410 to 20160612 and 20170103 to 20170214 during which time good quality microwave radiometer measurements are available.

54 ENVIRONMENTAL SCIENCES↗

Comparing gas composition from fast pyrolysis of live foliage measured in bench-scale and fire-scale experiments

Background: Fire models have used pyrolysis data from oxidising and non-oxidising environments for flaming combustion. In wildland fires pyrolysis, flaming and smouldering combustion typically occur in an oxidising environment (the atmosphere). Aims: Using compositional data analysis methods, determine if the composition of pyrolysis gases measured in non-oxidising and ambient (oxidising) atmospheric conditions were similar. Methods: Permanent gases and tars were measured in a fuel-rich (non-oxidising) environment in a flat flame burner (FFB). Permanent and light hydrocarbon gases were measured for the same fuels heated by a fire flame in ambient atmospheric conditions (oxidising environment). Log-ratio balances of the measured gases common to both environments (CO, CO 2 , CH 4 , H 2 , C 6 H 6 O (phenol), and other gases) were examined by principal components analysis (PCA), canonical discriminant analysis (CDA) and permutational multivariate analysis of variance (PERMANOVA). Key results: Mean composition changed between the non-oxidising and ambient atmosphere samples. PCA showed that flat flame burner (FFB) samples were tightly clustered and distinct from the ambient atmosphere samples. CDA found that the difference between environments was defined by the CO-CO 2 log-ratio balance. PERMANOVA and pairwise comparisons found FFB samples differed from the ambient atmosphere samples which did not differ from each other. Conclusion: Relative composition of these pyrolysis gases differed between the oxidising and non-oxidising environments. This comparison was one of the first comparisons made between bench-scale and field scale pyrolysis measurements using compositional data analysis. Implications: These results indicate the need for more fundamental research on the early time-dependent pyrolysis of vegetation in the presence of oxygen.

54 ENVIRONMENTAL SCIENCES↗

Autonomous platform for solution processing of electronic polymers

The manipulation of electronic polymers’ solid-state properties through processing is crucial in electronics and energy research. Yet, efficiently processing electronic polymer solutions into thin films with specific properties remains a formidable challenge. We introduce Polybot, an artificial intelligence (AI) driven automated material laboratory designed to autonomously explore processing pathways for achieving high-conductivity, low-defect electronic polymers films. Leveraging importance-guided Bayesian optimization, Polybot efficiently navigates a complex 7-dimensional processing space. In particular, the automated workflow and algorithms effectively explore the search space, mitigate biases, employ statistical methods to ensure data repeatability, and concurrently optimize multiple objectives with precision. The experimental campaign yields scale-up fabrication recipes, producing transparent conductive thin films with averaged conductivity exceeding 4500 S/cm. Feature importance analysis and morphological characterizations reveal key design factors. This work signifies a significant step towards transforming the manufacturing of electronic polymers, highlighting the potential of AI-driven automation in material science.

Wang, Chengshi [Argonne National Laboratory (ANL),↗

DESI Spectroscopy of HETDEX Emission-line Candidates. I. Line Discrimination Validation

The Hobby–Eberly Dark Energy Experiment (HETDEX) is an untargeted spectroscopic galaxy survey that uses Lyα-emitting galaxies (LAEs) as tracers of 1.9 < z < 3.5 large-scale structure. Most detections consist of a single emission line, whose identity is inferred via a Bayesian analysis of ancillary data. To determine the accuracy of these line identifications, HETDEX detections were observed with the Dark Energy Spectroscopic Instrument (DESI). In two DESI pointings, high-confidence spectroscopic redshifts are obtained for 1157 sources, including 982 LAEs. The DESI spectra are used to evaluate the accuracy of the HETDEX object classifications and tune the methodology to achieve the HETDEX science requirement of ≲2% contamination of the LAE sample by low-redshift emission-line galaxies, while still assigning 96% of the true Lyα emission sample with the correct spectroscopic redshift. We compare emission-line measurements between the two experiments assuming a simple Gaussian line fitting model. Fitted values for the central wavelength of the emission line, the measured line flux, and line widths are consistent between the surveys within uncertainties. Derived spectroscopic redshifts, from the two classification pipelines, when both agree as an LAE classification, are consistent to within $\langle$Δz/(1 + z)$\rangle$ = 6.9 × 10 −5 with an rms scatter of 3.3 × 10 −4 . Data are available at https://data.desi.lbl.gov/desi/public/dr1/vac/dr1/hetdex.

79 ASTRONOMY AND ASTROPHYSICS↗

Developing Asparagaceae1726: An Asparagaceae‐specific probe set targeting 1726 loci for Hyb‐Seq and phylogenomics in the family

Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.

Plant Sciences↗

Integration of equitable resilience metrics into climate-informed electric utility planning processes: phase one

Working together, Sandia National Laboratories, Southern California Edison (SCE) - an Investor-Owned Utility (IOU) - and the California Public Utilities Commission (CPUC) are studying how electric utilities can use equity and resilience metrics to help inform the prioritization and sequencing of resilience-driven infrastructure investments. To this end, this project evaluated “Social Burden,” an equitable resilience metric which measures the potential impact of disruptions in access to non-electric critical services on people and estimates community resilience to these disruptions. The Social Burden was expanded to incorporate SCE’s existing equity metric and applied to evaluate the potential impacts from a range of climate-informed hypothetical outage scenarios developed under SCE’s 2022 Climate Adaptation Vulnerability Assessment. One baseline (“blue-sky”) state and eight different outage scenarios were evaluated to measure the potential impacts of the outages on non-electric infrastructure, critical services, and people. Key findings include: 1) the Social Burden framework is flexible enough to adapt to and build upon existing utility equity and/or resilience metrics, 2) Social Burden results highlight the high degree of non-electric service redundancy within the SCE service area with most (6/8) hypothetical outage scenarios predicted to increase people’s Social Burden by less than 10%; however, 3) access to critical services and people’s ability to obtain them is unequal and spatially clustered, meaning that there are some hypothetical outage scenarios (2/8) that will exert a higher toll on communities directly experiencing the outage as well as some nearby communities with pre-existing vulnerabilities. The report concludes with recommendations for potential use cases of the expanded Social Burden metric and identifies priority follow-on work. Potential use cases may include incorporating equity into IOU’s prioritization of climate resilience investments. Additionally, Social Burden analysis may provide additional data and insights to augment grid planning, potentially by identifying additional needs and/or prioritizing previously identified needs.

24 POWER TRANSMISSION AND DISTRIBUTION↗