Search NASA⌕ Search

SEARCH · Search NASA

Results for “data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Network performance analysis for HPC datacenters (net_perf) v1.0

The software has two main features: (1) identify data movement trends in HPC data centers that use network flow monitoring (2) analyze the performance of individual data flows under the existing data movement management strategy and identify performance bottlenecks that impede timely data availability for science workflows. Its main advantage is that it is tailored for HPC network traffic by considering HPC data movement management intricacies.

Giannakou, Anna↗

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Measurement of the free neutron lifetime in a magneto-gravitational trap with in situ detection

Here, in this study, we publish three years of data from the UCNτ experiment performed at the Los Alamos Ultracold Neutron Facility at the Los Alamos Neutron Science Center. These data are in addition to our previously published data. Our goals in this paper are to better understand and quantify systematic uncertainties and to improve the lifetime statistical precision. We previously reported a value from our 2017–2018 data for the neutron lifetime of 877.75 ± 0.28 (statistical) +0.22–0.16 (systematic) s. We have collected an additional three years of data reported here for the first time. When all the data from UCNτ are averaged for 2017, 2018, 2020, 2021, and 2022, we report an updated value for the lifetime of 877.83 ± 0.22 (statistical)+0.20–0.17 (systematic) s. We utilized improved monitor detectors, reduced our correction due to UCN upscattering on residual gas, and employed four different UCN detector geometries both to reduce the correction required for rate dependence and to explore potential contributions due to phase space evolution.

Cabibbo-Kobayashi-Maskawa matrix↗

Elucidating the formation mechanisms of zeolites using data-driven modelling and in-situ characterization

Zeolites are the main solid catalysts used by the chemical industry. The use of zeolites in separations and as shape selective catalysts requires control of the width and connectivity of their pores. 235 distinct zeolite frameworks have been synthesized to date, of over 2 million that have been proposed. Recent work indicates that the limitation is in large part kinetic: new synthetic pathways are required to access new zeolites. Organic cations are used to direct the synthesis towards specific zeolites. However, the molecular mechanisms by which cations direct the nucleation towards specific zeolites is not known. Elucidating these mechanisms is key to realize new zeolites for catalysis and separations, and is the focus of this project. This project developed and implemented a synergistic, data-driven computational and experimental approach to resolve the molecular pathways of nucleation, growth, and polymorph selection of zeolites and the role of organic cations in directing their formation. The project developed computationally efficient and accurate models for the study of the nucleation and growth of pure silica zeolites in molecular simulations, using machine learning with data from experiments. Simulations with these models were integrated with scanning tunneling electron microscopy, computer vision, and deep learning to unveil the molecular pathways of formation of a zeolite. Of particular interest in this project was to elucidate the role of amorphous precursors in the nucleation of the zeolite. Previous experiments indicate that zeolites are born within non-crystalline aggregates in which the silicates and organic cations have local and medium range order similar to that of the zeolite. The organic cations that direct the formation of zeolites and those that direct the formation of ordered mesoporous silicas are similar. We hypothesized that the frustrated attraction that for large organic cations leads to the formation of stable mesophases that direct the synthesis of mesoporous silicas, could promote the formation of metastable mesophases that can assist in the nucleation and polymorph selection of zeolites. The simulations resolved how structure directing agents build crystalline order and showed that mesoscopic pre-ordering occurs is not required to facilitate the nucleation of zeolites, because the synthesis occurs at high driving forces, where the barriers for nucleation are negligible. This project unveiled that polymorph selection in zeolite synthesis occurs after nucleation, opening a distinct area of control through the kinetics of growth and not through nucleation barriers.

36 MATERIALS SCIENCE↗

2024 OES-Environmental 2024 State of the Science Report, Chapter 8: Marine Renewable Energy Data and Information Systems

As the marine renewable energy (MRE) sector grows, large amounts of environmental and technical data and information are being collected. When these data and information are openly available, they can be used to guide research and development, inform responsible siting and consenting of projects, and increase stakeholder understanding through transparency. For example, quality environmental data collected during the siting, consenting, construction, operation, and decommissioning of MRE projects can all play key roles in better characterizing baseline conditions, developing effective monitoring and mitigation strategies, and retiring environmental risks through data transferability (see Chapter 6). Ensuring that these data and information are easily discoverable and accessible will help the MRE sector make informed decisions and coexist in an increasingly busy ocean environment.

16 TIDAL AND WAVE POWER↗

Integrating science for water security governance

Hydrological extremes are intensifying globally, increasing the complexity of decisions required to ensure water security. Advances in hydrological science, modeling, and data systems have expanded the technical frontier of water research, yet uptake of scientific insights in policy and management decisions remains limited. This persistent science–policy gap is not primarily a failure of knowledge generation or robustness, but an institutional challenge shaped by how scientific and governance systems are organized, coordinated, and connected to support the effective use of scientific knowledge. These challenges are particularly pronounced in multi-level and transboundary water governance, where decisions span jurisdictions and require coordination across institutional and political boundaries. We synthesize research at the science–policy interface and evidence from water security initiatives to show how institutional arrangements, scientific tool development, and research practices enable or constrain the sustained use of scientific knowledge in water-security governance processes. Building on these insights, we develop ‘shared decision infrastructure’ as a framing to describe how scientific knowledge is embedded within the institutional, relational, and procedural arrangements that connect science to decision-making processes over time. We translate this framing into a practical intervention roadmap centered on institutional design, tool translation, sustained co-production, and outcome-oriented evaluation to support the integration of science into ongoing governance processes. By positioning science as shared decision infrastructure, the roadmap clarifies how researchers can design scientific efforts that support more coordinated, accountable, and adaptive water security decisions amid deepening uncertainty.

M whitney, Kristen [NASA Goddard Space Flight Cent↗

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.↗

The Early Data Release of the Dark Energy Spectroscopic Instrument

The Dark Energy Spectroscopic Instrument (DESI) completed its 5 month Survey Validation in 2021 May. Spectra of stellar and extragalactic targets from Survey Validation constitute the first major data sample from the DESI survey. This paper describes the public release of those spectra, the catalogs of derived properties, and the intermediate data products. In total, the public release includes good-quality spectral information from 466,447 objects targeted as part of the Milky Way Survey, 428,758 as part of the Bright Galaxy Survey, 227,318 as part of the Luminous Red Galaxy sample, 437,664 as part of the Emission Line Galaxy sample, and 76,079 as part of the Quasar sample. In addition, the release includes spectral information from 137,148 objects that expand the scope beyond the primary samples as part of a series of secondary programs. Here, we describe the spectral data, data quality, data products, Large-Scale Structure science catalogs, access to the data, and references that provide relevant background to using these spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

An open source knowledge graph ecosystem for the life sciences

Translational research requires data at multiple scales of biological organization. Advancements in sequencing and multi-omics technologies have increased the availability of these data, but researchers face significant integration challenges. Knowledge graphs (KGs) are used to model complex phenomena, and methods exist to construct them automatically. However, tackling complex biomedical integration problems requires flexibility in the way knowledge is modeled. Moreover, existing KG construction methods provide robust tooling at the cost of fixed or limited choices among knowledge representation models. PheKnowLator (Phenotype Knowledge Translator) is a semantic ecosystem for automating the FAIR (Findable, Accessible, Interoperable, and Reusable) construction of ontologically grounded KGs with fully customizable knowledge representation. The ecosystem includes KG construction resources (e.g., data preparation APIs), analysis tools (e.g., SPARQL endpoint resources and abstraction algorithms), and benchmarks (e.g., prebuilt KGs). We evaluated the ecosystem by systematically comparing it to existing open-source KG construction methods and by analyzing its computational performance when used to construct 12 different large-scale KGs. With flexible knowledge representation, PheKnowLator enables fully customizable KGs without compromising performance or usability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING↗

Ecobuoys for Scalable Oceanography

An approach to scalable surface-drifting buoys is needed to enable the high spatial and temporal resolution of oceanographic data that the science and meteorological communities are asking for. With the number of active buoys predicted to increase by a factor of 100 or more, the impact on the environment becomes even more important. Here, we present a pathway to a scalable and sustainable generation of buoys. We identify the main criteria to be used when developing such buoys to be low cost, with reliable data and neutral or even positive environmental impact. For each buoy subsystem—hull, electronics, energy generation and storage, sensors, and communication system—cutting-edge technological solutions are presented, many of them from emerging research in marine or other disciplines. We then assess the potential solutions against the design criteria and plot a path toward small, environmentally friendly, low-cost, and low-power buoys.

54 ENVIRONMENTAL SCIENCES↗

Transforming Energy Through Computational Excellence: A View From NREL

At the National Renewable Energy Laboratory (NREL)—a U.S. Department of Energy laboratory—computational science, high-performance computing, applied mathematics, advanced computer science, visualization, and data play a pivotal role in advancing energy abundance, affordability, security, and reliability. From fundamental scientifc discovery to systems engineering and analysis, NREL researchers tackle market-relevant challenges to develop solutions for an independent energy system that is reliable, resilient and secure. Collaborative partnerships with industry, government, and academia ensure that our research remains cutting edge, impactful, applicable, and aligned with real-world energy needs. This special issue of Computing in Science & Engineering highlights exemplary NREL projects where computational tools and methodologies drive discovery and accelerate innovation in scalable and integrated energy systems. The featured articles explore the role of computational modeling, high-performance computing, generative AI, and adaptive computing in advancing independent energy solutions, optimizing sustainability research, and enhancing decision-making for energy solutions using a broad mix of energy technologies. Here, these contributions demonstrate how NREL’s computational research bridges the gap between theoretical advancements and practical implementation, emphasizing interdisciplinary collaboration and a commitment to innovation, with a focus on translating computational excellence into real-world impact, thus accelerate progress toward national energy goals. By showcasing cutting-edge research at the intersection of computational science and energy systems, this issue aims to inspire and inform researchers, practitioners, and policymakers dedicated to shaping a more reliable energy future.

97 MATHEMATICS AND COMPUTING↗

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj↗

Toward a microscopic picture of hadronization and multi-parton processes

This project advanced the understanding of how quarks and gluons produced in high-energy collisions transform into the hadrons observed in particle detectors, a fundamental process known as quantum chromodynamics (QCD) hadronization. By combining theoretical calculations, quantum simulation methods, and modern AI techniques, the research developed new tools to study multi-parton dynamics and nonperturbative effects that are essential for interpreting data from current and future nuclear physics experiments. Key outcomes include new theoretical frameworks for jet and hadron measurements, pioneering quantum simulation algorithms for real-time dynamics in field theories, and the development of advanced machine-learning models, such as diffusion models and explainable classifiers, to simulate and analyze collider events. These results are directly relevant to experiments at Jefferson Lab, Brookhaven National Laboratory, and the future Electron-Ion Collider, and they also have a broader impact in areas such as quantum information science and data-driven modeling of complex systems. The project supported the training of graduate students and postdoctoral fellows and contributed to the broader scientific community through publications, workshops, and collaborative activities. Overall, this work provides new insights into the microscopic mechanisms of hadron formation and establishes a foundation for future studies at the intersection of nuclear physics, artificial intelligence, and quantum computing.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Response of Subsurface Nitrogen-Cycling Microbial Communities to Environmental Fluctuations (Final Technical Report)

Riparian floodplains are dynamic ecosystems linking terrestrial and riverine systems. These floodplains experience hydrological shifts such as changes in water table height, flooding, and drought and can be ‘hotspots’ of biogeochemical cycling due to shifting sediment moisture (and saturation) and subsurface exchanges of water, nutrients, and other compounds across different sediment layers. Subsurface microbial communities are the primary drivers of biogeochemical processes in floodplains, and thus their structure and function can directly influence both surface and groundwater quality. The microbial nitrogen (N) cycle is particularly important in floodplains as it affects nutrient availability and removal. Two functional guilds of chemoautotrophic (i.e. CO2-fixing) microorganisms are responsible for the first oxidative step of the N cycle, nitrification: ammonia-oxidizing archaea (AOA) and bacteria (AOB) catalyze the oxidation of ammonia to nitrite, while nitrite-oxidizing bacteria (NOB) oxidize nitrite to nitrate. Despite the critical role nitrification plays in N-cycling in both terrestrial and aquatic ecosystems, our understanding of the diversity, ecophysiology, and activity of nitrifying organisms in subsurface floodplain soils/sediments is extremely limited. To help address this critical knowledge gap, the overarching goal of this project was to determine how shifts in key environmental parameters and gradients impact microbial N-cycling communities/processes, with particular emphasis on nitrification, within hydrologically-variable floodplain sediments in the Wind River Basin near Riverton, Wyoming. The three specific objectives of this project were to: (1) to associate in situ environmental drivers of N cycling with distinct functional guilds; (2) determine the guild response to variation in key ecosystem drivers; and (3) develop a dynamic ecosystem model of the microbial N cycle with the Riverton subsurface using community genomic and biogeochemical data collected in the first two objectives. Over the course of this project, we employed both 16S rRNA gene amplicon sequencing and genome-resolved metagenomics to examine the phylogenetic diversity and metabolic potential of subsurface nitrifier communities within 68 samples collected across multiple sites, depths, and time points within the Riverton floodplain, allowing for both spatial and temporal investigations at different scales. This project benefitted tremendously from recent advances in high-throughput sequencing technologies coupled with dramatic improvements in the computational tools and algorithms available for analyzing such large, complex genomic datasets. By pairing these cutting-edge genomic approaches with depth-resolved sampling and detailed geochemical analyses of the Riverton floodplain, we have gained novel insights into the structure and function of subsurface nitrifier communities in relation to both hydrology and biogeochemistry. This project resulted in the most detailed and comprehensive characterization of N-cycling floodplain microbial communities to date and will hopefully inspire and pave the way for future studies using similar approaches in other floodplains. Indeed, such information is critical for understanding subsurface biogeochemical cycling and how elemental stores are altered from perturbations initiated by the water cycle within floodplains. Finally, because of the terrestrial-aquatic nature of the Riverton floodplain, results from this project are also of relevance to disciplines such as soil science, estuarine science, limnology & oceanography, biogeochemistry, geobiology, environmental engineering, as well as genomics and data science.

54 ENVIRONMENTAL SCIENCES↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

RTN-011: Rubin Observatory Plans for an Early Science Program

This document outlines Rubin Observatory's plans for a dedicated \emph{Early Science Program} to enable high-impact science prior to the first annual data release of the Legacy Survey of Space and Time (LSST). Components of the Early Science Program include releasing science-grade commissioning data products via a series of ``Data Previews,'' ramping up of the transient alert stream during commissioning, implementing a program of incremental template generation to augment alert production in the early phases of the survey, and the first LSST Data Release, DR1, based on the first 6 months of data from the LSST. A detailed breakdown of which data products can be expected when is provided. The Rubin Operations team is working closely with the science community to optimize the Early Science Program for the time-domain and solar system science achievable in the first year of operations. This is a living document; both it and the Early Science Program will continue to evolve over the course of commissioning and pre-operations in response to the state of the as-built system and to community guidance.

79 ASTRONOMY AND ASTROPHYSICS↗