Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Airborne imaging spectroscopy surveys of Arctic and boreal Alaska and northwestern Canada 2017–2023

Since 2015, NASA’s Arctic Boreal Vulnerability Experiment (ABoVE) has investigated how climate change impacts the vulnerability and/or resilience of the permafrost-affected ecosystems of Alaska and northwestern Canada. ABoVE conducted extensive surveys with the Next Generation Airborne Visible/Infrared Imaging Spectrometer (AVIRIS-NG) during 2017, 2018, 2019, and 2022 and with AVIRIS-3 in 2023 to characterize tundra, taiga, peatlands, and wetlands in unprecedented detail. The ABoVE AVIRIS dataset comprises ~1700 individual flight lines covering ~120,000 km 2 with nominal 5 m × 5 m spatial resolution. Data include individual transects to capture important gradients like the tundra-taiga ecotone and maps of up to 10,000 km 2 for key study areas like the Mackenzie Delta. The ABoVE AVIRIS surveys enable diverse ecosystem science, provide crucial benchmark data for validating retrievals from the PACE, PRISMA, and EnMAP satellite sensors and help prepare for the SBG and CHIME missions. This paper guides interested researchers to fully explore the ABoVE AVIRIS spectral imagery and complements our guide to the ABoVE airborne synthetic aperture radar surveys.

Miller, Charles E. [California Institute of Techno↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗

Urbanization and malaria have a contextual relationship in endemic areas: A temporal and spatial study in Ghana

In West Africa, malaria is one of the leading causes of disease-induced deaths. Existing studies indicate that as urbanization increases, there is corresponding decrease in malaria prevalence. However, in malaria-endemic areas, the prevalence in some rural areas is sometimes lower than in some peri-urban and urban areas. Therefore, the relationship between the degree of urbanization, the impact of living in urban areas, and the prevalence of malaria remains unclear. This study explores this association in Ghana, using epidemiological data at the district level (2015–2018) and data on health, hygiene, and education. We applied a multilevel model and time series decomposition to understand the epidemiological pattern of malaria in Ghana. Then we classified the districts of Ghana into rural, peri-urban, and urban areas using administratively defined urbanization, total built areas, and built intensity. We converted the prevalence time series into cross-sectional data for each district by extracting features from the data. To predict the determinant most impacting according to the degree of urbanization, we used a cluster-specific random forest. We find that prevalence is impacted by seasonality, but the trend of the seasonal signature is not noticeable in urban and peri-urban areas. While urban districts have a slightly lower prevalence, there are still pockets with higher rates within these regions. These areas of high prevalence are linked to proximity to water bodies and waterways, but the rise in these same variables is not associated with the increase of prevalence in peri-urban areas. The increase in nightlight reflectance in rural areas is associated with an increased prevalence. We conclude that urbanization is not the main factor driving the decline in malaria. However, the data indicate that understanding and managing malaria prevalence in urbanization will necessitate a focus on these contextual factors. Finally, we design an interactive tool, ’malDecision’ that allows data-supported decision-making.

60 APPLIED LIFE SCIENCES↗

RTN-045: Guidelines for User Tutorials

This document defines the guidelines, principles, and formats for user-facing tutorials that demonstrate how to use the Rubin Science Platform (RSP) to analyze data from the Legacy Survey of Space and Time (LSST). All Rubin staff and the broader science community should use these guidelines when contributing to the sets of Jupyter Notebook or documentation-based tutorials maintained by the Rubin Community Science team (CST).

79 ASTRONOMY AND ASTROPHYSICS↗

Characterization of Non-Science Grade DESI CCDs and Creating Additional ICARUS Monitoring Metrics for Data Quality

DESI, otherwise known as the Dark Energy Spectroscopic Instrument, is an astronomical project measuring millions of optical spectra to characterize Dark Energy. The survey has been running for 3 years. My work delves into the data taking/analysis side of the DESI CCDs (Charge-Coupled Devices) which will help in swapping out broken CCDs on the instrument. Additionally, I worked on making new metrics for ICARUS, Imaging Cosmic and Rare Underground Signals for their monitoring of run data. This paper will talk about the types of characterizations done for DESI and additional metrics for data quality and tools used for ICARUS monitoring.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

TDCOSMO - XVII. New time delays in 22 lensed quasars from optical monitoring with the ESO-VST 2.6m and MPG 2.2m telescopes

We present new time delays, the main ingredient of time delay cosmography, for 22 lensed quasars resulting from high-cadence r-band monitoring on the 2.6 m ESO VLT Survey Telescope and Max-Planck-Gesellschaft 2.2 m telescope. Each lensed quasar was typically monitored for one to four seasons, often shared between the two telescopes to mitigate the interruptions forced by the COVID-19 pandemic. The sample of targets consists of 19 quadruply and 3 doubly imaged quasars, which received a total of 1918 hours of on-sky time split into 21 581 wide-field frames, each 320 seconds long. In a given field, the 5-σ depth of the combined exposures typically reaches the 27th magnitude, while that of single visits is 24.5 mag – similar to the expected depth of the upcoming Vera-Rubin LSST. The fluxes of the different lensed images of the targets were reliably de-blended, providing not only light curves with photometric precision down to the photon noise limit, but also high-resolution models of the targets whose features and astrometry were systematically confirmed in Hubble Space Telescope imaging. This was made possible thanks to a new photometric pipeline, lightcurver, and the forward modelling method STARRED. Finally, the time delays between pairs of curves and their uncertainties were estimated, taking into account the degeneracy due to microlensing, and for the first time the full covariance matrices of the delay pairs are provided. Of note, this survey, with 13 square degrees, has applications beyond that of time delays, such as the study of the structure function of the multiple high-redshift quasars present in the footprint at a new high in terms of both depth and frequency. The reduced images will be available through the European Southern Observatory Science Portal.Key words: methods: data analysis / surveys / distance scale

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Convergence in simulating global soil organic carbon by structurally different models after data assimilation

Abstract Current biogeochemical models produce carbon–climate feedback projections with large uncertainties, often attributed to their structural differences when simulating soil organic carbon (SOC) dynamics worldwide. However, choices of model parameter values that quantify the strength and represent properties of different soil carbon cycle processes could also contribute to model simulation uncertainties. Here, we demonstrate the critical role of using common observational data in reducing model uncertainty in estimates of global SOC storage. Two structurally different models featuring distinctive carbon pools, decomposition kinetics, and carbon transfer pathways simulate opposite global SOC distributions with their customary parameter values yet converge to similar results after being informed by the same global SOC database using a data assimilation approach. The converged spatial SOC simulations result from similar simulations in key model components such as carbon transfer efficiency, baseline decomposition rate, and environmental effects on carbon fluxes by these two models after data assimilation. Moreover, data assimilation results suggest equally effective simulations of SOC using models following either first‐order or Michaelis–Menten kinetics at the global scale. Nevertheless, a wider range of data with high‐quality control and assurance are needed to further constrain SOC dynamics simulations and reduce unconstrained parameters. New sets of data, such as microbial genomics‐function relationships, may also suggest novel structures to account for in future model development. Overall, our results highlight the importance of observational data in informing model development and constraining model predictions.

54 ENVIRONMENTAL SCIENCES↗

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗

Millimeter-wave observations of Euclid Deep Field South using the South Pole Telescope: A data release of temperature maps and catalogs

Context. The South Pole Telescope third-generation camera (SPT-3G) has observed over 10,000 square degrees of sky at 95, 150, and 220 GHz (3.3, 2.0, 1.4 mm, respectively) and will significantly overlap the ongoing 14,000 square-degree Euclid Wide Survey. The Euclid collaboration recently released Euclid Deep Field South (EDF-S) observations of 23 square degrees at wide field depths in the first quick data release (Q1). Aims. With the goal of releasing complementary millimeter-wave data and encouraging legacy science, we performed dedicated observations of a 57-square-degree field overlapping the EDF-S. Methods. The observing time totaled 20 days, and we reached noise depths of 4.3, 3.8, and 13.2 $μ$K-arcmin at 95, 150, and 220 GHz, respectively. Results. In this work we present the temperature maps and two catalogs constructed from these data. The emissive source catalog contains 601 objects (334 inside EDF-S) with 54% synchrotron-dominated sources and 46% thermal dust emission-dominated sources. The 5$σ$ detection thresholds are 1.7, 2.0, and 6.5 mJy in the three bands. The cluster catalog contains 217 cluster candidates (121 inside EDF-S) with median mass $M_{500c}=2.12 \times 10^{14} M_{\odot}/h_{70}$ and median redshift $z$ = 0.70, corresponding to an order-of-magnitude improvement in cluster density over previous tSZ-selected catalogs in this region (3.81 clusters per square degree). Conclusions. The overlap between SPT and Euclid data will enable a range of multiwavelength studies of the aforementioned source populations. This work serves as the first step toward joint projects between SPT and Euclid and provides a rich dataset containing information on galaxies, clusters, and their environments.

Archipley, M. [Chicago U., Astron. Astrophys. Ctr.↗

Predicting Critical Transitions in Multiscale Data

Predicting the dynamics of complex nonlinear systems remains a challenging problem both in dynamical systems theory as well as real world science and engineering applications. Data-driven methods utilizing the latest advances in machine learning (ML) provide a promising new paradigm for this task. Our work centered on Reservoir Computing (RC), which has shown itself to be capable of skillfully predicting chaotic dynamics in multiscale systems. In the first part of the work, the focus is on how to improve predictions of critical transitions in a class of slow-fast metastable systems in which the equations are known. An additional goal was to determine whether a relationship exists between RC and Koopman operator theory, to improve the efficiency and broaden the applicability of the approach. In the second part of this work, a variation on the RC model known as Reconstructive Reservoir Computing (RRC) is applied to real-world data to identify anomalies.

97 MATHEMATICS AND COMPUTING↗

CROCUS Tipping Bucket Rain Gauge Data from Argonne Deployable Mast Deployed at Argonne National Laboratory During Urban Flooding Campaign

The Tipping Bucket Rain Gauge (TBRG) dataset contains data from a non-heated Met One 12-inch tipping bucket rain gauge that was mounted on the Argonne Deployable Mast (ADM). The ADM is a rapid deployable meteorological trailer that can be outfitted with instrumentation to measure urban heat island effects, urban flooding or urban flux measurements. During the urban flooding field campaign, the ADM was outfitted with multiple precipitation measurement systems, including the TBRG. This dataset contains one minute measurements for precipitation accumulation during the ADM's deployment at the Argonne Testbed for Multiscale Observational Science (ATMOS) site. These data are helpful for identifying periods of precipitation, leading to potential flooding. TBRGs can be used to validate optical rain gauge data and disdrometer data collected during the CROCUS urban flooding campaign. Data were collected at ATMOS, a 20-acre prairie site at Argonne National Laboratory in Lemont, Illinois. The data is presented as daily NetCDF (.nc) files, each containing approximately 24 hours of observations. Files follow the naming convention of: the project (CROCUS), location (ADM-atmos), instrument name (tbrg), data level (raw, a1), and date (year, month, day). The NetCDF format can be accessed using common scientific software such as Python using xarray, netCDF4 or act-doe.

1-min Precipitation Accumulation↗

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-Driven Approach for Controlled Icosahedral Boron- Rich Compound Growth

This final technical report summarizes the research accomplishments and research highlights at the end of the funding period. This project aimed to leverage existing and new computational data produced from first-principles and molecular dynamics simulations to understand the thermodynamic, mechanical, and electronic properties of icosahedral boron compounds. The goal is to achieve targeted material properties by controlling the synthesis routes of these boron-rich compounds.

36 MATERIALS SCIENCE↗

CMIP7 Data Request: atmosphere priorities and opportunities

This paper presents a comprehensive overview of the Coupled Model Intercomparison Project Phase 7 (CMIP7) request for data unlocking key research avenues in atmospheric science and provides justification for the resources needed to produce this data. Topics within the CMIP7 Atmosphere Theme centre around processes and feedbacks in atmospheric science such as clouds, aerosols and atmospheric chemistry, atmospheric circulation, temperature variability and extremes, radiative forcings, and Earth system model evaluation. These topics are summarised in this paper as scientific “opportunities” which will be realised through CMIP7 experiments and Earth system model outputs. These opportunities were submitted by a thematic group of atmospheric science community representatives combined with an extended consultation process. The production of these variables will close key gaps and uncertainties identified during previous rounds of CMIP, and will be broadly used by scientific, policy, governmental, industry, and other communities that rely on climate model projections for research and decision making, including supporting the 7th Intergovernmental Panel on Climate Change Assessment Report (AR7). As an author group, we also reflect on the process used to collate this data request and make recommendations to future CMIP governance on implementing a consultation on this scale in the future.

58 GEOSCIENCES↗

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

Building morphologies of the USA structures database; a gauntlet feature set

In recent years there has been a proliferation of methods and data to extract building footprints from satellite imagery. However there has been very little effort to provide additional insight about these buildings beyond their spatial location and shape. Features derived from their geometries can be used to better characterize these buildings which are critical for further research and development. In this work a set of 65 unique features for every building for more than 131 million buildings covering the US has been developed. This rich feature dataset will enable researchers, policymakers and various agencies to derive additional building characteristics like height, occupancy type, and help to gain valuable and new insights of the built environment.

Environmental sciences↗