Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Observation impacts in the lower troposphere and challenges of Planetary Boundary Layer data assimilation

The Goddard Earth Observing System (GEOS) developed by the NASA Global Modeling and Assimilation Office assimilates a wide range of observations to support various NASA Earth Science missions. To set the stage for follow-on Planetary Boundary Layer (PBL) science and prepare for future observing systems of the next decade, we have assessed the effectiveness of the use of existing observing systems in the lower troposphere in GEOS, and analyzed model responses to the incremental analysis update (IAU) forcing. With a better understanding of the GEOS data assimilation algorithms in the PBL, we have developed strategies for improved PBL data assimilation in GEOS. The strategies to enhance data usages in both the data assimilation system and forecast model will be presented, and the utilization of PBL height data from multiple observing systems will be discussed as well.

Yanqiu Zhu↗

Addressing the Big-Earth-Data Variety Challenge with the Hierarchical Triangular Mesh

We have implemented an updated Hierarchical Triangular Mesh (HTM) as the basis for a unified data model and an indexing scheme for geoscience data to address the variety challenge of Big Earth Data. We observe that, in the absence of variety, the volume challenge of Big Data is relatively easily addressable with parallel processing. The more important challenge in achieving optimal value with a Big Data solution for Earth Science (ES) data analysis, however, is being able to achieve good scalability with variety. With HTM unifying at least the three popular data models, i.e. Grid, Swath, and Point, used by current ES data products, data preparation time for integrative analysis of diverse datasets can be drastically reduced and better variety scaling can be achieved. In addition, since HTM is also an indexing scheme, when it is used to index all ES datasets, data placement alignment (or co-location) on the shared nothing architecture, which most Big Data systems are based on, is guaranteed and better performance is ensured. Moreover, our updated HTM encoding turns most geospatial set operations into integer interval operations, gaining further performance advantages.

SciDB↗

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Data items will occasionally be removed from OSM if they are misidentified, if they no longer exist, if they are duplicates of another item, or similar. For that reason, updated versions of this database may not contain all data center locations included in previous versions. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Artificial Neural Network (ANN) Surface Longwave and Shortwave Fluxes Trained on CERES Observations

The Clouds and Earth’s Radiant Energy System (CERES) project provides satellite-based observations of the radiative fluxes and clouds systems. CERES climate quality data products typically take several months of calibration and validation before release to the public. The Fast Longwave and Shortwave Radiative Flux (FLASHFlux) data product was developed to provide key data for the applied sciences and educational users within a week of observation. FLASHFlux achieves this by using simplified calibration, an operational meteorological product from Global Modeling and Assimilation Office (GMAO), and its own surface parameterizations model. The CERES FLASHFlux provides two data products: 1) an hourly Level 2 Single Scanner Footprint (SSF) data separately for Terra and NOAA-20 observations, and 2) a daily Level 3 Time Interpolated and Spatially Averaged (TISA) 1o x 1o gridded data that combines Terra and NOAA-20 observations. Currently, FLASHFlux uses the Langley Parameterized Shortwave Algorithm (LPSA) and Langley Parameterized Longwave Algorithm (LPLA) to derive its surface fluxes (Kratz et al., 2010; Gupta et al, 2001). A new Machine Learning (ML) based approach using Artificial Neural Networks to derive Surface Longwave (LW) & Shortwave (SW) fluxes based on training data from the CERES Clouds Radiative Swath (CRS) product is being investigated to replace LPSA and LPLA in the SSF surface flux products. One of the biggest hurdles in training ML model is model fitting. To overcome the problem of overfitting we use feature engineering that helps in finding the important feature and remove features that are irrelevant to the model. In our training we employed the Leave-One-Feature-Out Importance (LOFO) to evaluate the significance of each feature in our training. We intercompare ANN fluxes against surface fluxes produced from the Fu-Liou model in CRS and the LPSA/LPLA in FLASHFlux SSF. Furthermore, we validated ANN derived fluxes to the Baseline Surface Radiation Network (BSRN).

P C Sawaengphokhai↗

Envisioning U.S. Climate Predictions and Projections to Meet New Challenges

In the face of a changing climate, the understanding, predictions, and projections of natural and human systems are increasingly crucial to prepare and cope with extremes and cascading hazards, determine unexpected feedbacks and potential tipping points, inform long-term adaptation strategies, and guide mitigation approaches. Increasingly complex socio-economic systems require enhanced predictive information to support advanced practices. Such new predictive challenges drive the need to fully capitalize on ambitious scientific and technological opportunities. These include the unrealized potential for very high-resolution modeling of global-to-local Earth system processes across timescales, reduction of model biases, enhanced integration of human systems and the Earth Systems, better quantification of predictability and uncertainties; expedited science-to-service pathways, and co-production of actionable information with stakeholders. Enabling technological opportunities include exascale computing, advanced data storage, novel observations and powerful data analytics, including artificial intelligence and machine learning. Looking to generate community discussions on how to accelerate progress on U.S. climate predictions and projections, representatives of Federally-funded U.S. modeling groups outline here perspectives on a six-pillar national approach grounded in climate science that builds on the strengths of the U.S. modeling community and agency goals. This calls for an unprecedented level of coordination to capitalize on transformative opportunities, augmenting and complementing current modeling center capabilities and plans to support agency missions. Tangible outcomes include projections with horizontal spatial resolutions finer than 10 km, representing extremes and associated risks in greater detail, reduced model errors, better predictability estimates, and more customized projections to support next generation climate services.

54 ENVIRONMENTAL SCIENCES↗

NASA Technical Management Report (533Q)

The objective of this task is analytical support of the NASA Satellite Laser Ranging (SLR) program in the areas of SLR data analysis, software development, assessment of SLR station performance, development of improved models for atmospheric propagation and interpretation of station calibration techniques, and science coordination and analysis functions for the NASA led Central Bureau of the International Laser Ranging Service (ILRS). The contractor shall in each year of the five year contract: (1) Provide software development and analysis support to the NASA SLR program and the ILRS. Attend and make analysis reports at the monthly meetings of the Central Bureau of the ILRS covering data received during the previous period. Provide support to the Analysis Working Group of the ILRS including special tiger teams that are established to handle unique analysis problems. Support the updating of the SLR Bibliography contained on the ILRS web site; (2) Perform special assessments of SLR station performance from available data to determine unique biases and technical problems at the station; (3) Develop improvements to models of atmospheric propagation and for handling pre- and post-pass calibration data provided by global network stations; (4) Provide review presentation of overall ILRS network data results at one major scientific meeting per year; (5) Contribute to and support the publication of NASA SLR and ILRS reports highlighting the results of SLR analysis activity.

Klosko, S. M.↗

Strawman payload data for science and applications space platforms

The need for a free flying science and applications space platform to host compatible long duration experiment groupings in Earth orbit is discussed. Experiment level information on strawman payload models is presented which serves to identify and quantify the requirements for the space platform system. A description data base on the strawman payload model is presented along with experiment level and group level summaries. Payloads identified in the strawman model include the disciplines of resources observations and environmental observations.

Source record↗

Coupling of the Inner Magnetosphere with the Underlying Atmosphere and Ionosphere

The following is a final report summarizing our very successful inner magnetosphere research program through which we have made significant contributions to: (1) research through data analysis, modeling and participation in community-wide campaigns, (2) the development of the space science discipline through leadership in national and international campaigns, service on steering committees, review panels and the development and maintenance of campaign and community web sites, (3) education and human resources by the participation of graduate, undergraduate and high school students in our research programs and (4) outreach through development of web-based materials and interactive games. We describe each of these activities below.

Sharber, James M.↗

Drought Prediction for Socio-Cultural Stability Project

The primary objective of this project is to answer the question: "Can existing, linked infrastructures be used to predict the onset of drought months in advance?" Based on our work, the answer to this question is "yes" with the qualifiers that skill depends on both lead-time and location, and especially with the associated teleconnections (e.g., ENSO, Indian Ocean Dipole) active in a given region season. As part of this work, we successfully developed a prototype drought early warning system based on existing/mature NASA Earth science components including the Goddard Earth Observing System Data Assimilation System Version 5 (GEOS-5) forecasting model, the Land Information System (LIS) land data assimilation software framework, the Catchment Land Surface Model (CLSM), remotely sensed terrestrial water storage from the Gravity Recovery and Climate Experiment (GRACE) and remotely sensed soil moisture products from the Aqua/Advanced Microwave Scanning Radiometer - EOS (AMSR-E). We focused on a single drought year - 2011 - during which major agricultural droughts occurred with devastating impacts in the Texas-Mexico region of North America (TEXMEX) and the Horn of Africa (HOA). Our results demonstrate that GEOS-5 precipitation forecasts show skill globally at 1-month lead, and can show up to 3 months skill regionally in the TEXMEX and HOA areas. Our results also demonstrate that the CLSM soil moisture percentiles are a goof indicator of drought, as compared to the North American Drought Monitor of TEXMEX and a combination of Famine Early Warning Systems Network (FEWS NET) data and Moderate Resolution Imaging Spectrometer (MODIS)'s Normalizing Difference Vegetation Index (NDVI) anomalies over HOA. The data assimilation experiments produced mixed results. GRACE terrestrial water storage (TWS) assimilation was found to significantly improve soil moisture and evapotransportation, as well as drought monitoring via soil moisture percentiles, while AMSR-E soil moisture assimilation produced marginal benefits. We carried out 1-3 month lead-time forecast experiments using GEOS-5 forecasts as input to LIS/CLSM. Based on these forecast experiments, we find that the expected skill in GEOS-5 forecasts from 1-3 months is present in the soil moisture percentiles used to indicate drought. In the case of the HOA drought, the failure of the long rains in April appears in the February 1, March 1 and April 1 initialized forecasts, suggesting that for this case, drought forecasting would have provided some advance warning about the drought conditions observed in 2011. Three key recommendations for follow-up work include: (1) carry out a comprehensive analysis of droughts observed over the entire period of record for GEOS-5 forecasts; (2) continue to analyze the GEOS-5 forecasts in HOA stratifying by anomalies in long and short rains; and (3) continue to include GRACE TWS, Soil Moisture/Ocean Salinity (SMOS) and the upcoming NASA Soil Moisture Active/Passive (SMAP) soil moisture products in a routine activity building on this prototype to further quantify the benefits for drought assessment and prediction.

Peters-Lidard, Christa↗

A Machine Learning Approach to Predict Martensitic Transition Temperatures for Shape Memory Alloys

Shape memory alloys (SMAs) are a unique class of materials with several remarkable properties including shape recovery, superelasticity, etc. Especially important for many NASA applications is the ability to tune the martensitic phase transition temperature by varying the alloy composition. Nickel-titanium (NiTi) based alloys are the most widely studied of this class, with compositions involving ternary, quaternary, or higher additions being considered. Over the past several years, a significant database of SMA properties has been assembled by NASA researchers. Such a database is ideal for data science-based approaches including machine learning. We present results from a developed machine learning model capable of accurately predicting the transition temperature of SMAs across a wide range of compositions. Our model has the added benefit of interpretability and even provides confidence intervals for our predictions. This model will make rapid screening and design of new SMA materials possible. Predictions from the machine learning model can be validated by empirical and/or atomistic scale modeling.

Shreyas Honrao↗

Senteniel-6 Radio Occultation Product Released by NASA GES DISC to Supplement Satellite Remote Sensing Datasets for PBL Study

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) curates hyperspectral atmospheric sounder remote sensing and numerical model reanalysis datasets which have been utilized in the Planetary Boundary Layer (PBL) research and applications. The hyperspectral sounder remote-sensing datasets include the Atmospheric Infrared Sounder (AIRS) on the Aqua satellite to the Cross-track Infrared Sounder (CrIS) on Suomi--National Polar- orbiting Partnership (NPP) and National Oceanic and Atmospheric Administration -20 (N NOAA-20)/ Joint Polar-orbiting Satellite System -1 (JPSS-1). The Modern-Era Retrospective analysis for Research and Applications Version 2 (MERRA-2) global reanalysis product provides a data record commencing in 1980. The sounder remote sensing and reanalysis datasets include temperature, water vapor, and trace gas profile down to the PBL, and also have a derived PBL height as well. A nearly 10-year (June 2006 to December 2015) seasonal and annual PBL height climatology dataset from COSMIC Global Navigation Satellite System (GNSS) radio occultation (RO) measurement is also available from the GES DISC. In collaboration with Sentinel-6 Project, the GES DISC is implementing curation activities for GNSS RO products from the Sentinel-6A/Sentinel-6 Michael Freilich satellite launched on November 21, 2020. Sentinel-6A RO products provide refractivity, temperature, and humidity profile with finer vertical resolution, leveraging PBL research and application as a supplement to the hyperspectral sounder remote sensing and reanalysis products. The public release of Sentinel- 6A RO products is scheduled for mid-October of 2021. In this presentation, we will introduce all Senitnel-6A products and services, and demonstrate use cases studying the PBL by combining these products with other GES DISC archived data products.

Feng Ding↗

Impact of Canadian Wildfires on Mid Atlantic’s Region Air Quality: An Analysis Using ASDC Data

Wildfires pose a growing concern in North America due to their harmful impacts on air quality and public health, with increased wildfire activity in recent years leading to widespread smoke plumes that can transcend borders. The exposure of New York City (NYC), the most populous city in North America, to Canadian wildfire smoke highlights the substantial implications for public health and urban environments. To better understand the impact of Canadian wildfires on air quality in NYC, satellite data from the NASA Atmospheric Science Data Center (ASDC) at Langley Research Center, along with ground-based measurements and atmospheric modeling results, are analyzed. We examine concentrations of atmospheric aerosols—particularly PM2.5 particulate matter originating from Canadian wildfires—their dispersion patterns, and the duration and intensity of smoke events impacting NYC. Data from multiple satellites, such as those from the Earth Polychromatic Imaging Camera (EPIC), are synergistically used to identify regions affected by wildfires and estimate aerosol loading. Ground-based measurements, including data from air quality monitoring stations, provide localized information for validation and calibration purposes. The findings of this study contribute to our understanding of the impact of Canadian wildfires on NYC's air quality and emphasize the importance of monitoring and prediction of transboundary smoke events using data synthesized from multiple sources, such as those provided by the ASDC. This information is crucial for policymakers, public health officials, and residents in affected areas to develop effective strategies for mitigating the health risks associated with wildfire smoke and improving air quality during wildfire seasons. The utilization of ASDC data in this research highlights the critical role of atmospheric remote sensing in addressing the challenges posed by wildfires and their consequences on regional scales.

Ingrid Garcia-Solera↗

Analyzing the Impact of Canadian Wildfires on Air Quality in the U.S. Mid-Atlantic: with Data and Tools from NASA’s Atmospheric Sciences Data Center

Wildfires pose a growing concern in North America due to their harmful impacts on air quality and public health, with increased wildfire activity in recent years leading to widespread smoke plumes that can transcend borders. The exposure of New York City (NYC), the most populous city in North America, to Canadian wildfire smoke highlights the substantial implications for public health and urban environments. To better understand the impact of Canadian wildfires on air quality in NYC, satellite data from the NASA Atmospheric Science Data Center (ASDC) at Langley Research Center, along with ground-based measurements and atmospheric modeling results, are analyzed. We examine concentrations of atmospheric aerosols—particularly PM2.5 particulate matter originating from Canadian wildfires—their dispersion patterns, and the duration and intensity of smoke events impacting NYC. Data from multiple satellites, such as those from the Earth Polychromatic Imaging Camera (EPIC), are synergistically used to identify regions affected by wildfires and estimate aerosol loading. Ground-based measurements, including data from air quality monitoring stations, provide localized information for validation and calibration purposes. The findings of this study contribute to our understanding of the impact of Canadian wildfires on NYC's air quality and emphasize the importance of monitoring and prediction of transboundary smoke events using data synthesized from multiple sources, such as those provided by the ASDC. This information is crucial for policymakers, public health officials, and residents in affected areas to develop effective strategies for mitigating the health risks associated with wildfire smoke and improving air quality during wildfire seasons. The utilization of ASDC data in this research highlights the critical role of atmospheric remote sensing in addressing the challenges posed by wildfires and their consequences on regional scales.

Ingrid Garcia-Solera↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

Preface to the Special Issue on Modeling and Data Analysis Methods for the SMILE mission

The SMILE (Solar wind Magnetosphere Ionosphere Link Explorer) project (http://www.nssc.cas.cn/smile/, https://www.cosmos.esa.int/web/smile/mission) is a joint spacecraft mission of the European Space Agency (ESA) and the Chinese Academy of Sciences (CAS) with an expected launch in 2025. SMILE aims to study the global interactions of solar wind–magnetosphere–ionosphere innovatively by imaging the Earth’s magnetosheath and cusps in soft X-rays and the northern auroral region in ultraviolet (UV) while simultaneously measuring plasma and magnetic field parameters in the solar wind and magnetosheath along a highly-elliptical and highly-inclined orbit. This special issue is composed of 22 articles, presenting recent progress in modeling and data analysis techniques developed for the SMILE mission. In this preface, we categorize the articles into the following seven topics and provide brief summaries: (1) instrument descriptions of the Soft X-ray Imager (SXI), (2) numerical modeling of the X-ray signals, (3) data processing of the X-ray images, (4) boundary tracing methods from the simulated images, (5) physical phenomena and a mission concept related to the scientific goals of SMILE-SXI, (6) studies of the aurora, and (7) ground-based support for SMILE.

SMILE↗

Evolution of the Earth Observing System (EOS) Data and Information System (EOSDIS)

One of the strategic goals of the U.S. National Aeronautics and Space Administration (NASA) is to "Develop a balanced overall program of science, exploration, and aeronautics consistent with the redirection of the human spaceflight program to focus on exploration". An important sub-goal of this goal is to "Study Earth from space to advance scientific understanding and meet societal needs." NASA meets this subgoal in partnership with other U.S. agencies and international organizations through its Earth science program. A major component of NASA s Earth science program is the Earth Observing System (EOS). The EOS program was started in 1990 with the primary purpose of modeling global climate change. This program consists of a set of space-borne instruments, science teams, and a data system. The instruments are designed to obtain highly accurate, frequent and global measurements of geophysical properties of land, oceans and atmosphere. The science teams are responsible for designing the instruments as well as scientific algorithms to derive information from the instrument measurements. The data system, called the EOS Data and Information System (EOSDIS), produces data products using those algorithms as well as archives and distributes such products. The first of the EOS instruments were launched in November 1997 on the Japanese satellite called the Tropical Rainfall Measuring Mission (TRMM) and the last, on the U.S. satellite Aura, were launched in July 2004. The instrument science teams have been active since the inception of the program in 1990 and have participation from Brazil, Canada, France, Japan, Netherlands, United Kingdom and U.S. The development of EOSDIS was initiated in 1990, and this data system has been serving the user community since 1994. The purpose of this chapter is to discuss the history and evolution of EOSDIS since its beginnings to the present and indicate how it continues to evolve into the future. this chapter is organized as follows. Sect. 7.2 provides a discussion of EOSDIS, its elements and their functions. Sect. 7.3 provides details regarding the move towards more distributed systems for supporting both the core and community needs to be served by NASA Earth science data systems. Sect. 7.4 discusses the use of standards and interfaces and their importance in EOSDIS. Sect. 7.5 provides details about the EOSDIS Evolution Study. Sect. 7.6 presents the implementation of the EOSDIS Evolution plan. Sect. 7.7 briefly outlines the progress that the implementation has made towards the 2015 Vision, followed by a summary in Sect. 7.8.

Ramapriyan, Hampapuram K.↗