Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Strong Lens Discoveries in DESI Legacy Imaging Surveys DR10 with Two Deep Learning Architectures

Abstract We have conducted a search for strong gravitational lensing systems in the Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Surveys Data Release 10 (DR10). This paper is the fourth in a series of searches. This is the first catalog of lens candidates covering nearly the entirety of the extragalactic sky south of declination δ ≈ +32 ∘ , all observed by DECam, covering ∼14,000 deg 2 . We impose a z -band magnitude cut of <20 in AB magnitude. We deploy a residual neural network and EfficientNet as an ensemble trained on a compilation of known lensing systems and high-grade candidates as well as nonlenses in the same footprint. The predictions from these two base models are aggregated using a meta-learner. After applying our ensemble to the survey data, we exclude known candidates and systems, and use our own visual inspection portal to rank images in the top 0.01 percentile of all neural network recommendations. We have found 811 lens candidates, five of which are confirmed through Euclid Quick Data Release (Q1). These include 484 new candidates in the Legacy Surveys DR9 footprint, all parts of which have been searched for strong lenses at least once before, either by our group or others. Combining the discoveries from this work with those from the first three papers in this series (335, 1210, and 1512), we have discovered a total of 3868 new candidates in the DESI Legacy Surveys.

Inchausti, Jose Carlos [University of San Francisc↗

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi↗

Accelerating discoveries at DIII-D with the Integrated Research Infrastructure

DIII-D research is being accelerated by leveraging high performance computing (HPC) and data resources available through the National Energy Research Scientific Computing Center (NERSC) Superfacility initiative. As part of this initiative, a high-resolution, fully automated, whole discharge kinetic equilibrium reconstruction workflow was developed that runs at the NERSC for most DIII-D shots in under 20 min. This has eliminated a long-standing research barrier and opened the door to more sophisticated analyses, including plasma transport and stability. These capabilities would benefit from being automated and executed within the larger Department of Energy Advanced Scientific Computing Research program’s Integrated Research Infrastructure (IRI) framework. The goal of IRI is to empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. For transport, we are looking at producing flux matched profiles and also using particle tracing to predict fast ion heat deposition from neutral beam injection before a shot takes place. Our starting point for evaluating plasma stability focuses on the pedestal limits that must be navigated to achieve better confinement. This information is meant to help operators run more effective experiments, so it needs to be available rapidly inside the DIII-D control room. So far this has been achieved by ensuring the data is available with existing tools, but as more novel results are produced new visualization tools must be developed. In addition, all of the high-quality data we have generated has been collected into databases that can unlock even deeper insights. This has already been leveraged for model and code validation studies as well as for developing AI/ML surrogates. The workflows developed for this project are intended to serve as prototypes that can be replicated on other experiments and can be run to provide timely and essential information for ITER, as well as next stage fusion power plants.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Toward Accelerating Discovery via Physics-Driven and Interactive Multifidelity Bayesian Optimization

Both computational and experimental material discovery bring forth the challenge of exploring multidimensional and often nondifferentiable parameter spaces, such as phase diagrams of Hamiltonians with multiple interactions, composition spaces of combinatorial libraries, processing spaces, and molecular embedding spaces. Often these systems are expensive or time consuming to evaluate a single instance, and hence classical approaches based on exhaustive grid or random search are too data intensive. This resulted in strong interest toward active learning methods such as Bayesian optimization (BO) where the adaptive exploration occurs based on human learning (discovery) objective. However, classical BO is based on a predefined optimization target, and policies balancing exploration and exploitation are purely data driven. In practical settings, the domain expert can pose prior knowledge of the system in the form of partially known physics laws and exploration policies often vary during the experiment. Here, we propose an interactive workflow building on multifidelity BO (MFBO), starting with classical (data-driven) MFBO, then expand to a proposed structured (physics-driven) structured MFBO (sMFBO), and finally extend it to allow human-in-the-loop interactive interactive MFBO (iMFBO) workflows for adaptive and domain expert aligned exploration. These approaches are demonstrated over highly nonsmooth multifidelity simulation data generated from an Ising model, considering spin–spin interaction as parameter space, lattice sizes as fidelity spaces, and the objective as maximizing heat capacity. Detailed analysis and comparison show the impact of physics knowledge injection and real-time human decisions for improved exploration with increased alignment to ground truth. Here, the associated notebooks allow to reproduce the reported analyses and apply them to other systems.

97 MATHEMATICS AND COMPUTING↗

Discovery of additional ancient genome duplications in yeasts

Whole-genome duplication (WGD) has had profound macroevolutionary impacts on diverse lineages, preceding adaptive radiations in vertebrates, teleost fish, and angiosperms. In contrast to the many known ancient WGDs in animals, and especially plants, we are aware of evidence for only four WGDs in fungi. The oldest of these occurred ∼100 million years ago (mya) and is shared by ∼60 extant Saccharomycetales species, including the baker’s yeast Saccharomyces cerevisiae. Notably, this is the only known ancient WGD event in the yeast subphylum Saccharomycotina. The dearth of ancient WGD events in fungi remains a mystery. Some studies have suggested that fungal lineages that experience chromosome and genome duplication quickly go extinct, leaving no trace in the genomic record, while others contend that the lack of known WGDs is due to an absence of data. Under the second hypothesis, additional sampling and deeper sequencing of fungal genomes should lead to the discovery of more WGD events. Coupling hundreds of recently published genomes from nearly every described Saccharomycotina species, with three additional long-read assemblies, we discovered three novel WGD events. Although the functions of retained duplicate genes originating from these events are broad, they bear similarities to the well-known WGD that occurred in the Saccharomycetales. In conclusion, our results suggest that WGD may be a more common evolutionary force in fungi than previously believed.

convergent evolution↗

The Vera C. Rubin Observatory Data Preview 1

We present Rubin Data Preview 1 (DP1), the first data from the National Science Foundation–Department of Energy Vera C. Rubin Observatory, comprising raw and calibrated single-epoch images, coadds, difference images, detection catalogs, and ancillary data products. DP1 is based on 1792 optical–near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera (LSSTComCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile in late 2024. DP1 covers ∼15 deg 2 distributed across seven roughly equal-sized noncontiguous fields, each independently observed in six broad photometric bands, ugrizy. The median FWHM of the point-spread function across all bands is approximately 1"14, with the sharpest images reaching about 0." 58. The 5σ point-source depths for coadded images in the deepest field, the Extended Chandra Deep Field South, are u = 24.55, g = 26.18, r = 25.96, i = 25.71, z = 25.07, and y = 23.1. Other fields are no more than 2.2 mag shallower in any band, where they have nonzero coverage. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band in coadds, and 431 solar system objects, of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and is available to Vera C. Rubin Observatory data rights holders via the Rubin Science Platform, a cloud-based environment for the analysis of petascale astronomical data. While small compared to future LSST releases, its high quality and diversity of data support a broad range of early science investigations ahead of full operations in 2026.

Ground-based astronomy↗

Structural constraint integration in a generative model for the discovery of quantum materials

Billions of organic molecules have been computationally generated, yet functional inorganic materials remain scarce due to limited data and structural complexity. Here, in this work, we introduce Structural Constraint Integration in a GENerative model (SCIGEN), a framework that enforces geometric constraints, such as honeycomb and kagome lattices, within diffusion-based generative models to discover stable quantum materials candidates. SCIGEN enables conditional sampling from the original distribution, preserving output validity while guiding structural motifs. This approach generates ten million inorganic compounds with Archimedean and Lieb lattices, over 10% of which pass multistage stability screening. High-throughput density functional theory calculations on 26,000 candidates shows over 95% convergence and 53% structural stability. A graph neural network classifier detects magnetic ordering in 41% of relaxed structures. Furthermore, we synthesize and characterize two predicted materials, TiPd 0.22 Bi 0.88 and Ti 0.5 Pd 1.5 Sb, which display paramagnetic and diamagnetic behaviour, respectively. Our results indicate that SCIGEN provides a scalable path for generating quantum materials guided by lattice geometry.

36 MATERIALS SCIENCE↗

2025 Workshop on Envisioning Frontiers in AI and Computing for Biological Research: Position Papers

This workshop aims to identify key research directions for transforming biology using artificial intelligence (AI), machine learning (ML) and computational methods to facilitate the discovery of new behaviors, mechanisms, and designs of biological processes relevant to DOE missions, underpinning a broader U.S. bioeconomy. By developing novel AI/ML technologies to analyze and interpret complex biological data, researchers can organize and simulate biological processes at various scales as well as advance predictive understanding and manipulation of biological systems. This integration of computation, experimentation, and next-generation experimental technologies can lead to discoveries in new biological behaviors and mechanisms relevant to DOE missions. The focus is on how advanced computational and mathematical methods can impact this mission by exploring digital twins, foundation models, automated laboratory experiments, modeling of complex living systems, and data-driven approaches for the biodesign of plants and microbial systems. While data management is important, it is not the primary focus of this workshop, which will assess the current state, trends, and AI/ML challenges at the interface between biology and computational science to identify opportunities for high-impact research at their intersection. The goal is to define research needs and opportunities that align with biological sciences, computational sciences, and applied mathematics research.

59 BASIC BIOLOGICAL SCIENCES↗

RatXcan: A framework for cross-species integration of genome-wide association and gene expression data

Genome-wide association studies (GWAS) have implicated specific alleles and genes as risk factors for numerous complex traits. However, translating GWAS results into biologically and therapeutically meaningful discoveries remains extremely challenging. Most GWAS results identify noncoding regions of the genome, suggesting that differences in gene regulation are the major driver of trait variability. To better integrate GWAS results with gene regulatory polymorphisms, we previously developed PrediXcan (also known as “transcriptome-wide association studies” orTWAS), which maps SNPs to predicted gene expression using GWAS data. In this study, we developed RatXcan, a framework that extends this methodology to outbred heterogeneous stock (HS) rats. RatXcan accounts for the close familial relationships among HS rats by modeling the relatedness with a random effect that encodes the genetic relatedness. RatXcan also corrects for polygenic-driven inflation because of the equivalence between a relatedness random effect and the infinitesimal polygenic model. To develop RatXcan, we trained transcript predictors for 8,934 genes using reference genotype and expression data from five rat brain regions. We found that the cis genetic architecture of gene expression in both rats and humans was sparse and similar across brain tissues. We tested the association between predicted expression in rats and two example traits (body length and BMI) using phenotype and genotype data from 5,401 densely genotyped HS rats and identified a significant enrichment between the genes associated with rat and human body length and BMI. Thus, RatXcan represents a valuable tool for identifying the relationship between gene expression and phenotypes across species and paves the way to explore shared biological mechanisms of complex traits.

Genetics & Heredity↗

Protocol for applying a network-enabled gene discovery pipeline to non-model plant species

Identifying upstream regulators of key genes is essential for understanding gene regulatory mechanisms and translating these insights into functional targets. Here, we present a protocol for applying the network-enabled gene discovery pipeline (NEEDLE) to non-model plant species. We describe steps for environment setup, data preparation, computational analysis, expected outputs, and parameter considerations. NEEDLE integrates RNA sequencing (RNA-seq) processing, weighted gene co-expression analysis (WGCNA), Gene Network Inference with Ensemble of trees (GENIE3), and promoter conservation analysis to prioritize candidate transcriptional regulators.

Plant Sciences↗

Superlative mechanical energy absorbing efficiency discovered through self-driving lab-human partnership

Energy absorbing efficiency is a key determinant of a structure’s ability to provide mechanical protection and is defined by the amount of energy that can be absorbed prior to stresses increasing to a level that damages the system to be protected. Here, we explore the energy absorbing efficiency of additively manufactured polymer structures by using a self-driving lab (SDL) to perform >25,000 physical experiments on generalized cylindrical shells. We use a human-SDL collaborative approach where experiments are selected from over trillions of candidates in an 11-dimensional parameter space using Bayesian optimization and then automatically performed while the human team monitors progress to periodically modify aspects of the system. The result of this human-SDL campaign is the discovery of a structure with a 75.2% energy absorbing efficiency and a library of experimental data that reveals transferable principles for designing tough structures.

42 ENGINEERING↗

Electronic structure simulations in the cloud computing environment

The transformative impact of modern computational paradigms and technologies, such as high-performance computing, quantum computing, and cloud computing, has opened up profound new opportunities for scientific simulations. Scalable computational chemistry is one beneficiary of this technological progress. The main focus of this paper is on the performance of various quantum chemical formulations, ranging from low-order methods to high-accuracy approaches, implemented in different computational chemistry packages, such as NWChem, NWChemEx, SPEC, ExaChem, and FLOSIC codes on the Azure Quantum Element (AQE) Microsoft cloud services. We pay particular attention to the intricate workflows for performing composite chemistry simulations, associated data curation, and mechanisms for accuracy assessment, as defined by the enabling cloud Computational Chemistry as a Service (CCaaS). Our focus also extends to Arrows' automated workflow for high throughput simulations. Finally, we provide a perspective on the role of cloud computing in supporting the mission of leadership computational facilities (LCFs).

computational chemistry, electronic structure, Clo↗

VISIONARY: Virtual Intelligence System for Optimizing Novel Analytical Research Yields

VISIONARY is an AI system that accelerates energy materials discovery by automatically generating hypotheses about structure-property relationships. It analyzes patterns in materials data, identifies promising correlations, and proposes testable scientific hypotheses without human intervention. By streamlining this reasoning process, VISIONARY helps researchers efficiently identify candidate materials with desired properties, significantly speeding up the materials development pipeline for energy applications. During the project, we developed a standalone application. The application uses a combination of papers provided by the user and data collected from FutureHouse’s dataset to build an understanding of the background that the user wants to explore for the hypothesis.

36 MATERIALS SCIENCE↗

Transforming Energy Through Computational Excellence: A View From NREL

At the National Renewable Energy Laboratory (NREL)—a U.S. Department of Energy laboratory—computational science, high-performance computing, applied mathematics, advanced computer science, visualization, and data play a pivotal role in advancing energy abundance, affordability, security, and reliability. From fundamental scientifc discovery to systems engineering and analysis, NREL researchers tackle market-relevant challenges to develop solutions for an independent energy system that is reliable, resilient and secure. Collaborative partnerships with industry, government, and academia ensure that our research remains cutting edge, impactful, applicable, and aligned with real-world energy needs. This special issue of Computing in Science & Engineering highlights exemplary NREL projects where computational tools and methodologies drive discovery and accelerate innovation in scalable and integrated energy systems. The featured articles explore the role of computational modeling, high-performance computing, generative AI, and adaptive computing in advancing independent energy solutions, optimizing sustainability research, and enhancing decision-making for energy solutions using a broad mix of energy technologies. Here, these contributions demonstrate how NREL’s computational research bridges the gap between theoretical advancements and practical implementation, emphasizing interdisciplinary collaboration and a commitment to innovation, with a focus on translating computational excellence into real-world impact, thus accelerate progress toward national energy goals. By showcasing cutting-edge research at the intersection of computational science and energy systems, this issue aims to inspire and inform researchers, practitioners, and policymakers dedicated to shaping a more reliable energy future.

97 MATHEMATICS AND COMPUTING↗

Accelerating the discovery of low-energy structure configurations: A computational approach that integrates first-principles calculations, Monte Carlo sampling, and Machine Learning

Finding Minimum Energy Configurations (MECs) is essential in fields such as physics, chemistry, and materials science, as they represent the most stable states of the systems. In particular, identifying such MECs in multi-component alloys considered candidate PFMs is key because it determines the most stable arrangement of atoms within the alloy, directly influencing its phase stability, structural integrity, and thermo-mechanical properties. However, since the search space grows exponentially with the number of atoms considered, obtaining such MECs using computationally expensive first-principles DFT calculations often results in a cumbersome task. To escape the above compromise between physical fidelity and computational efficiency, we have developed a novel physics-based data-driven approach that combines Monte Carlo sampling, first-principles DFT calculations, and Machine Learning to accelerate the discovery of MECs in multi-component alloys. More specifically, we have leveraged well-established Cluster Expansion (CE) techniques with Local Outlier Factor models to establish strategies that enhance the reliability of the CE method. In this work, we demonstrated the capabilities of the proposed approach for the particular case of a tungsten-based quaternary high-entropy alloy. However, the method is applicable to other types of alloys and enables a wide range of applications.

36 MATERIALS SCIENCE↗

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is that the amount of equipment reliability (ER) data being continuously generated are extremely large. These data elements come in different forms: textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) and they provide system engineers with valuable insights and information regarding the discovery of anomalous behaviors or degradation trends, the identification of the possible causes behind such behaviors and trends, and the prediction of their direct consequences. This paper directly targets the generation of knowledge from ER data by putting “data into context”. Here, we employ model-based system engineering (MBSE) models of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by identifying first which elements of the developed MBSE elements they are referring to. This task is much harder for textual data since the information contained in issue or maintenance reports needs to “be understood” by a computational tool. Here we called this process “knowledge extraction” where our methods to extract knowledge from textual data. Lastly, once numeric and textual ER data elements have been processed and “understood”, we discover possible cause-effect relations among them. This is performed by observing if a logical connection through the MBSE models exists, and if there is a temporal relation among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 MATHEMATICS AND COMPUTING↗

Invariant discovery of features across multiple length scales: Applications in microscopy and autonomous materials characterization

Physical imaging is a foundational characterization method in areas from condensed matter physics and chemistry to astronomy and spans length scales from atomic to universe. Images encapsulate crucial data regarding atomic bonding, materials microstructures, and dynamic phenomena such as microstructural evolution and turbulence, among other phenomena. The challenge lies in effectively extracting and interpreting this information. Variational Autoencoders (VAEs) have emerged as powerful tools for identifying the underlying factors of variation in image data, providing a systematic approach to distilling meaningful patterns from complex data sets. However, a significant hurdle in their application is the definition and selection of appropriate descriptors reflecting local structures. Here, we introduce the scale-invariant VAE approach (SI-VAE) based on the progressive training of the VAE with the descriptors sampled at different length scales. The SI-VAE allows the discovery of the length scale-dependent factors of variation in the system. Here, we illustrate this approach using the ferroelectric domain images and generalize it to the movies of the electron-beam induced phenomena in graphene and topography evolution across combinatorial libraries. This approach can further be used to initialize the decision making in automated experiments including structure–property discovery and can be applied across a broad range of imaging methods. This approach is universal and can be applied to any spatially resolved data including both experimental imaging studies and simulations, and can be particularly useful for exploration of phenomena such as turbulence and scale-invariant transformation fronts.

36 MATERIALS SCIENCE↗

A hidden cysteine in Fis1 targeted to prevent excessive mitochondrial fission and dysfunction under oxidative stress

Fis1-mediated mitochondrial localization of Drp1 and excessive mitochondrial fission occur in human pathologies associated with oxidative stress. However, it is not known how Fis1 detects oxidative stress and what structural changes in Fis1 enable mitochondrial recruitment of Drp1. We find that conformational change involving α1 helix in Fis1 exposes its only cysteine, Cys41. In the presence of oxidative stress, the exposed Cys41 in activated Fis1 forms a disulfide bridge and the Fis1 covalent homodimers cause increased mitochondrial fission through increased Drp1 recruitment to mitochondria. Our discovery of a small molecule, SP11, that binds only to activated Fis1 by engaging Cys41, and data from genetically engineered cell lines lacking Cys41 strongly suggest a role of Fis1 homodimerization in Drp1 recruitment to mitochondria and excessive mitochondrial fission. The structure of activated Fis1-SP11 complex further confirms these insights related to Cys41 being the sensor for oxidative stress. Importantly, SP11 preserves mitochondrial integrity and function in cells during oxidative stress and thus may serve as a candidate molecule for the development of treatment for diseases with underlying Fis1-mediated mitochondrial fragmentation and dysfunction.

59 BASIC BIOLOGICAL SCIENCES↗