Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Carbon Storage Technical Viability Approach (CS TVA) Matrix

The Carbon Storage Technical Viability Approach (CS TVA) Matrix is a knowledge framework developed to outline the information needed for geologic carbon storage. The CS TVA Matrix contains 5 categories, 14 sub-categories, and 47 components. This framework can be leveraged to assess the availability of data and information needed for a carbon storage project. The information categories of the matrix are tied to a list of required data using weighted mapping, published herein.

carbon storage↗

Large reductions in Permian Basin methane intensity shown in multi-year comparison of aerially-visible methane emissions

Spanning the US states of Texas and New Mexico, the Permian Basin has been a hotspot of methane emissions from oil and natural gas activity 1–4, although studies disagree over the magnitude of these emissions. The most comprehensive measurement campaigns published were conducted in 2019 1–5 .There have been large changes in the energy industry since then, including in the prices of oil and gas, both state and federal regulatory environments, investor and activist pressure over methane emissions, and the adoption of new technologies and policies by energy operators. Understanding how any or all of these might influence methane emissions is important for policy makers, oil and gas operators, and other stakeholders. We characterize the time evolution of Permian Basin methane emissions using a series of comprehensive aerial surveys conducted every year from 2020-2023 and compare them to the 2019 results cited above. To maintain comparability, all the data sets are from surveys using Insight M point source methane sensing technology. The scope of these surveys expanded over time: from 33-46% of wells, oil production, and gas production in 2020 to 60-65% in 2021, to 84% of wells and over 90% of both oil and gas production in 2023. These surveys by Insight M also include hundreds of gas processing plants and compressor stations as well as 1000s of km of gathering and transmission pipelines. Considering only the aerially detected portion of emissions (typically the majority of the total in such surveys 4), we find reductions of more than 70% in methane emissions intensity compared to the 2019 New Mexico-only Insight M survey, with variation depending on the year 3,4. Notably, although sources below 100 kg/hr contributed less than 10% of aerially measured emissions the 2019 New Mexico survey 3, these smaller sources constitute a larger proportion of total aerially measured emissions (although not the majority) in 2020-2023. Production facilities and gathering pipelines are responsible for the larges shares of total emissions, followed by compressor stations and gas processing plants. Permian methane emissions were also measured in a comprehensive 2019 Permian-wide survey by the Carbon Mapper team 2. That analysis led to a lower total emissions estimate at the time 4. These new Insight M-based emission rates are still roughly 30-70% lower than the aerially measured portion of the 2019 Carbon Mapper-based estimates 4. Further work is needed to harmonize these surveys in space and time to create the most intercomparable numbers possible 5. Additional analysis is needed to compare our findings to the more spatially constrained 2020, 2021, and 2023 Carbon Mapper surveys in the Permian 4,6. The evidence is strong from these two survey teams that emissions intensity has declined significantly since 2019. Reasons for this trend are currently unclear but point to possible success of emissions control programs. Future work investigating frequency, source, and operator-specific intensities could provide insights into the causes of this promising trend.

methane, oil and gas, data science, remote sensing↗

Integrated Direct Air Capture and H₂-Free CO₂ Valorization

This project advances fundamental understanding of a novel integrated direct air capture (DAC) and CO₂ conversion process that valorizes atmospheric CO₂ without external H₂. The research encompasses four critical components: (1) design of task-specific ionic liquids for efficient CO₂ capture under ambient conditions, (2) development of H₂-free tandem catalytic systems using ethane as a reductant, (3) advanced operando characterization to elucidate capture and conversion mechanisms, and (4) data science-driven predictive computation to accelerate material discovery. Over the project period, we developed five high-performance DAC sorbent systems—including CaO/superbase ionic liquid composites, Ni-MOF/Ionic Liquid (IL) hybrids, fluorinated covalent organic frameworks with ion-pair functional groups, defect-engineered UiO-66, and a validated kinetic model for humid-condition operation, achieving CO₂ capacities up to 1.86 mmol/g at 400 ppm with excellent cycling stability. For H₂-free conversion, we constructed atomically synergistic Zn–O–Cr binuclear catalytic sites that achieve 100% ethylene selectivity, ~9.6% ethane conversion, and 99% CO₂ utilization in equimolar co-conversion of ethane and CO₂. We further demonstrated downstream valorization pathways converting CO and C₂H₄ into polyketones and C₃ chemicals. These advances strengthen the scientific foundation for producing value-added materials from ambient CO₂.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

String Data 2023 (Conference)

The annual String Data conferences have become the flagship annual meeting for the subfield at the interface of formal high energy theory, pure mathematics, and machine learning. String Data 2023 featured invited plenary talks by leading researchers in addition to a parallel session. The funds helped mitigate conference planning and provided support to young researchers.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Analysis of Bis(trifluoromethylsulfonyl)imide Interactions with Metal Cations Through a Chemical Informatics Approach

Nominally weakly coordinating anions are useful for modulating the solubility and chemical properties of metal complexes, but identification and analysis of the systematics of the interactions of anions with cationic metal complexes has not received the attention it deserves. Here, a chemical informatics approach is demonstrated for identifying and quantitatively analyzing the ways that the bis(trifluoromethylsulfonyl)imide anion (TFSI) can interact with metal-containing species. An open access computer program (PyCIFTer) was developed to facilitate large-scale structural analysis of TFSI-containing species by utilization of experimental atomic coordinate data from single-crystal X-ray diffraction (XRD) studies obtained from the Cambridge Structural Database (CSD). PyCIFTer establishes a three-dimensional vector space from the raw atomic coordinates, generating acyclic, undirected graphs that are used to rapidly analyze the structural properties (bond lengths and angles) of TFSI in individual structures in sequential/batch fashion. The structures are sorted by PyCIFTer into groups based on pre-set and chemically sensible criteria, affording a comprehensive and systematic view of TFSI structural chemistry. This approach avoids tedious one-at-a-time interrogation of structures, a prospect unreasonable in this case, and many others of contemporary chemical relevance; there were over 1500 structures in the CSD containing TFSI as of November 2024. The results demonstrate that TFSI only rarely binds to cations in the solid state, favoring the formation of species in which TFSI is found in cations’ outer coordination spheres. The prospect of applying PyCIFTer to other moieties is also discussed. PyCIFTer is also schematically compared to the commercial CSD Python application programming interface (API). Taken together, this work demonstrates the usefulness of modular workflows for sequential/batch analysis of structural data from XRD, an approach that appears poised to accelerate the translation of legacy structural results into new chemical insights and hypotheses.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Testing the LSST Difference Image Analysis Pipeline Using Synthetic Source Injection Analysis

Abstract We evaluate the performance of the Legacy Survey of Space and Time Science Pipelines Difference Image Analysis (DIA) on simulated images. By adding synthetic sources to galaxies on images, we trace the recovery of injected synthetic sources to evaluate the pipeline on images from the Dark Energy Science Collaboration Data Challenge 2. The pipeline performs well, with efficiency and flux accuracy consistent with the signal-to-noise ratio of the input images. We explore different spatial degrees of freedom for the Alard–Lupton polynomial-Gaussian image subtraction kernel and analyze for trade-offs in efficiency versus artifact rate. Increasing the kernel spatial degrees of freedom reduces the artifact rate without loss of efficiency. The flux measurements with different kernel spatial degrees of freedom are consistent. We also here provide a set of DIA flags that substantially filter out artifacts from the DIA source table. We explore the morphology and possible origins of the observed remaining subtraction artifacts and suggest that given the complexity of these artifact origins, a convolution kernel with a set of flexible bases with spatial variation may be needed to yield further improvements.

Liu, S. (ORCID:0000000244612143)↗

A Generalized approach to the operationalization of Software Quality Models

Comprehensive measures of quality are a research imperative, yet the development of software quality models is a wicked problem. Definitive solutions do not exist and quality is subjective at its most abstract. Definitional measures of quality are contingent on a domain, and even within a domain, the choice of representative characteristics to decompose quality is subjective. Thus, the operationalization of quality models brings even more challenges. A promising approach to quality modeling is the use of hierarchies to represent characteristics, where lower levels of the hierarchy represent concepts closer to real-world observations. Building upon prior hierarchical modeling approaches, we developed the Platform for Investigative software Quality Understanding and Evaluation (PIQUE). PIQUE surmounts several quality modeling challenges because it allows modelers to instantiate abstract hierarchical models in any domain by leveraging organizational tools tailored to their specific contexts. Here, we introduce PIQUE; exemplify its utility with two practical use cases; address challenges associated with parameterizing a PIQUE model; and describe algorithmic techniques that tackle normalization, aggregation, and interpolation of measurements.

Data aggregation↗

Understanding and Modeling Pooled Rideshare Acceptance: Influential Factors, Preferred User Experiences, and Implications

Ridesharing allows people to share a vehicle with others traveling in the same direction, which can reduce costs and traffic congestion. Pooled rideshare (PR) services, such as UberX Share and Lyft Shared, offer an economical and environmentally friendly alternative by matching passengers traveling similar routes. However, despite these benefits, PR adoption remains low due to concerns about safety, privacy, and convenience. This research explores the factors influencing PR adoption and provides recommendations to improve user acceptance. A nationwide survey of 5,385 participants across the U.S. was conducted to understand why people choose or avoid PR. The study identified five key factors influencing PR consideration: safety, service experience, privacy, traffic/environment, and time/cost. Additional research examined ways to optimize PR experiences by identifying four critical factors: comfort/ease of use, convenience, vehicle technology/accessibility, and passenger safety. To measure the impact of these factors, a statistical model called the Pooled Rideshare Acceptance Model (PRAM) was developed, providing insights into how each element influences PR adoption. Further analysis using the Pooled Rideshare Acceptance Model Multigroup Analyses (PRAMMA) revealed how demographic characteristics such as age, gender, income, and past rideshare experience shape PR perceptions. Some key findings from the multigroup analyses showed that younger users valued technological features and environmental benefits, while older users prioritized reliability and service transparency. Additionally, privacy concerns were more significant for female users, while convenience was critical for higher-income groups. These results emphasize that a 'onesize-fits-all' approach to PR service design is not effective, highlighting the need for tailored strategies to address different user segments. Further, workshops were conducted with researchers and students to translate the findings into real-world solutions. These workshops and 3 all the statistical analyses led to the development of 95 actionable recommendations. The recommendations focus on key areas such as safety, service reliability, user education, and accessibility, offering tangible improvements to PR services. The insights from this study provide valuable guidance for policymakers, transportation network companies (TNCs), and researchers aiming to make PR services safer, more accessible, and widely accepted. By addressing user concerns, PR can become a more viable transportation option, supporting sustainable urban mobility and reducing reliance on private vehicles. Additionally, these findings emphasize the importance of user-centric service design in encouraging broader PR adoption. Future research should explore evolving trends in PR preferences, technological advancements, and policy changes to ensure continued improvements. By implementing these recommendations, PR services can better align with user expectations, enhance trust in shared mobility, and contribute to a more efficient transportation ecosystem.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Solar Photovoltaic (PV) Module Facts and Trends

Unprecedented growth of solar PV has led to growing concerns about PV module toxicity and potential environmental and human health impacts. This factsheet provides objective, science-based data to help address these concerns and empower communities with the resources they need to make solar PV decisions.

14 SOLAR ENERGY↗

SITCOMTN-154: Initial studies of photometric redshifts with LSSTComCam from DP1

This technote holds reports based on the first analyses of the Data Preview 1 (DP1) data by the Science Unit for photometric redshifts. Although photometric redshifts are not an official DP1 data product, the "Photo-z Science Unit" generated photo-z estimates for every galaxy in DP1 using the available multi-band imaging on a best-effort basis. This work included developing training and test datasets by matching DP1 data to high-quality reference redshifts obtained with spectroscopy, Grism data, and multi-band photometry. The Science Unit used the RAIL software package to make photometric redshift estimates using eight different algorithms, developed simple scientific performance metrics, used those metrics to explore how the performance of the algorithms varied with configuration changes, derived more optimized configurations of the algorithms and tested the performance of those configurations. This work, the resulting data products and expected data distribution mechanism are all described there.

79 ASTRONOMY AND ASTROPHYSICS↗

Preparation of the Multi-Site Data Processing at the Vera C. Rubin Observatory

The Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST) Camera is scheduled to start taking data in the summer of 2025. The Data Release Production will run the LSST Science Pipe software at data facilities in the US, France and the UK. The LSST Science Pipeline consists of complex directed acyclic graphs (DAGs) of tasks. Rubin will use the Production and Distributed Analysis (PanDA) workflow and workload management system to orchestrate this complex workflow and the distribution of workloads to the data facilities. When run end-to-end by a team of data production staff, this processing (the Science Pipelines, distributed by the workflow and workload management system) is referred to as a 'campaign'. This paper describes the central services and data facility specific services that support this multi-site data process model, including the service deployment infrastructure, the workload and workflow system, the Campaign Management tools, and connection to Rubin Data Management. This paper will also mention the experience of processing the Rubin Commissioning Camera data. All these are part of the effort to scale up the processing capabilities for the expected very large data volume from the LSST Camera.

Yang, Wei [SLAC]↗

Raw Data

This dataset contains high-frequency (10Hz) data from the GX5-45 IMU on the Barge Science vans. The data are all raw binary files.

17 WIND ENERGY↗

CROCUS Urban Canyons - Space Science and Engineering Center (SPARC) Doppler lidar data

This is the netCDF format output from the Halo Photonics Streamline XR Doppler lidar that was deployed next to the Space Science and Engineering Center (SPARC) trailer at the University of Illnois-Chicago greenhouse parking lot during CROCUS Urban Canyons. The purpose of collecting this dataset is to provide vertical and horizontal wind profiles for studying the characteristics of turbulence over the urban canyon of Chicago. This data contains the radial velocity, intensity, and backscatter from the vertical profile, range height indicator, and sector scans that were performed over both Intensive Operating Period 1 and 2 of CROCUS Urban Canyons. There are four different types of files: * The Range Height Indicator (RHI) files contain scans that are along a constant azimuth, spanning the entire hemisphere of elevation values above the surface. * The Velocity Azimuth Display (VAD) files contain the raw radial velocity data from the 6-beam, 60 degree scans. * The User1 files contain stacked Plan Position Indicator scans over a 45 degree quadrant over downtown Chicago. * The Stare files contain vertically pointing scans. These are standard netCDF files that can be opened using xarray. The VAD scans can be processed from their raw radial velocities to horizontal wind speeds with the Atmospheric data Community Toolkit (https://arm-doe.github.io/ACT/).

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC WINDS↗

New approaches to Bayesian uncertainty quantification for Nuclear Science (Final Technical Report)

Inverse problems play a central role in experimentation and theory/data comparisons for many areas of modern Nuclear Physics (NP) and High-Energy Physics (HEP). Bayes’s Theorem is a powerful tool for solving Inverse Problems, providing conceptually transparent and unbiased constraints on theoretical parameters and their uncertainties (“Bayesian Inference”) and enabling the quantification of agreement or tension between models and data. However, analyses based on Bayesian Inference are often challenging for NP and HEP applications, either because of the large number of parameters in the problem, the high computational cost, or both. We propose a multi-institutional collaboration to develop and deploy novel Bayesian analysis tools that advance the scientific scope of a broad range of current and future NP experiments. This project brings together NP domain scientists working on several high-profile NP projects for which new, high-performance Bayesian Uncertainty Quantification (“Bayesian UQ”) methods are essential to carry out the science, and data scientists who are developing state-of-the-art methods applicable to these problems. The NP projects in this proposal comprise measurements of the mass and fundamental nature of the neutrino; study of the Quark-Gluon Plasma that filled the early universe; and mapping of natural and anthropogenic radiation environments. While these NP projects have very different scientific goals, with datasets and analysis approaches that differ significantly, they share common requirements for improving computationally intensive Bayesian analyses using advanced Machine Learning algorithms and will benefit strongly from a coherent effort to develop general solutions. This proposal brings together these projects and forefront ML-based data science algorithms to develop such general solutions. The methods developed in this project will also be more widely applicable, thereby advancing science in the larger Nuclear Physics portfolio.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗