Search NASA⌕ Search

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

WET Water Resources: A Google Earth Engine Python API Tool to Automate Wetland Extent Mapping Using Radar Satellite Sensors for Wetland Management and Monitoring

Wetland ecosystems are annually or seasonally wet transition zones between land and water. They provide a range of ecosystem services such as water filtration, flood mitigation, and carbon sequestration, as well as hosting biodiversity hotspots. Although they fulfill fundamental physical and natural processes, wetland extent and health are threatened by anthropogenic influences related to urbanization, population increase, pollution, and climate change. Recognizing the need to quantitatively monitor changes in these recently threatened ecosystems in a timely and cost-effective way, we developed a Google Earth Engine (GEE) Python API tool for automated wetland extent mapping using optical and radar satellite sensors that can be applied globally. The tool will significantly improve wetland change analysis and monitoring as the optical and SAR data proves high resolution (5-10 m) imagery, and SAR data is unaffected by cloud cover and light availability (day vs. night), which are common limitations for other remotely sensed sensors. The tool utilizes Copernicus Sentinel-1 C-band and NISAR L-band synthetic aperture radar (SAR) imagery. During image preprocessing, we applied a MODIS snow mask product to mask global snow coverage, which would affect land classification sensitivity. Calibration and validation were conducted through a historical change and sensitivity analysis of the Sudd watershed located in central Sudan. The tool was the first of its kind, as it enables NISAR data processing through an open-source GEE repository, further expanding and improving the utility of NASA Earth observations and contributing to NASA Open Science initiatives. We anticipate the tool will be used by researchers and practitioners interested in wetland monitoring and management..

Lori Berberian↗

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of machine-learning (ML) models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a user-friendly app that has been created to deploy trained machine-learning (ML) models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of ML application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). The app gets user input for a turbofan design, preprocesses the input data, and deploys trained ML models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The ML predictive models were built by employing supervised deep-learning and K-nearest neighbor regression algorithms to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these ML models using the app shows that Aero-Engines AI is an easy-touse and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage. Current version of the app focuses on the performance prediction of conventional turbofans. However, the scope of the app can easily be expanded to include other engine types (such as turboshaft and hybrid-electric systems) after their ML models are developed. Overall, the use of a machine-learning app for aircraft engine concept assessment represents a promising area of development in aircraft engine conceptual design.

machine learning↗

WET Water Resources: A Google Earth Engine Python API Tool to Automate Wetland Extent Mapping Using Radar Satellite Sensors for Wetland Management and Monitoring

Wetland ecosystems are annually or seasonally wet transition zones between land and water. They provide a range of ecosystem services such as water filtration, flood mitigation, and carbon sequestration, as well as hosting biodiversity hotspots. Although they fulfill fundamental physical and natural processes, wetland extent and health are threatened by anthropogenic influences related to urbanization, population increase, pollution, and climate change. Recognizing the need to quantitatively monitor changes in these recently threatened ecosystems in a timely and cost-effective way, we developed a Google Earth Engine (GEE) Python API tool for automated wetland extent mapping using optical and radar satellite sensors that can be applied globally. The tool will significantly improve wetland change analysis and monitoring as SAR data provides high resolution (5-10 m) imagery, unaffected by cloud cover and light availability (day vs. night), common limitations for other remotely sensed sensors. The tool utilizes Copernicus Sentinel-1 C-band and NISAR L-band (once operational and available on the GEE repository) synthetic aperture radar (SAR) imagery. During image preprocessing, we applied a Terra Moderate Resolution Imaging Spectroradiometer (MODIS) snow product to determine regional snow coverage, which affects land classification sensitivity. Calibration and validation were conducted through a historical change and sensitivity analysis of the Sudd wetland located in central Sudan. The tool was the first of its kind, as it enables NISAR data processing through an open-source GEE repository, further expanding and improving the utility of NASA Earth observations and contributing to NASA Open Science initiatives. We anticipate the tool will be used by researchers and practitioners interested in wetland monitoring and management.

Inundation↗

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of trained machine-learning models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a Windows app that has been created to deploy trained machine-learning models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of machine-learning application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). Current version of the app focuses on the performance prediction of conventional turbofans. The app gets user input for a turbofan design, preprocesses the input data, and deploys trained machine-learning models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The machine-learning predictive models were built by employing supervised deep-learning algorithm to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these machine-learning models using the app shows that Aero-Engines AI is an easy-to-use and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage.

machine learning↗

Prediction of Aircraft Estimated Time of Arrival Using A Supervised Learning Approach

We present a novel data-driven approach for prediction of the estimated time of arrival (ETA) of aircraft in the terminal area via the implementation of a Random Forest regression model. The model uses data fused from a number of sources (flight track, weather, flight plan information, etc.) and provides predictions for the remaining flight time for aircraft landing at Dallas/Fort Worth (DFW) International Airport. The predictions are made when the aircraft is at a distance of 200-miles from the airport. The results show that the model is able to predict estimated time of arrival to within ± 5 min for 90% of the flights in the test data with the mean absolute error being lower at 145 seconds. This paper covers the entire pipeline of data collection, preprocessing, setup and training of the ML model, and the results obtained for DFW.

Machine learning↗

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of machine-learning (ML) models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a user-friendly app that has been created to deploy trained machine-learning (ML) models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of ML application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). The app gets user input for a turbofan design, preprocesses the input data, and deploys trained ML models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The ML predictive models were built by employing supervised deep-learning and K-nearest neighbor regression algorithms to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these ML models using the app shows that Aero-Engines AI is an easy-touse and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage. Current version of the app focuses on the performance prediction of conventional turbofans. However, the scope of the app can easily be easily expanded to include other engine types (such as turboshaft and hybrid-electric systems) after their ML models are developed. Overall, the use of a machine-learning app for aircraft engine concept assessment represents a promising area of development in aircraft engine conceptual design.

machine learning↗

Commercial Smallsat Data Acquisition Program On-ramp #2 Airbus U.S. Synthetic Aperture Radar (SAR) Evaluation Report

In 2017, NASA’s Earth Science Division (ESD) launched the Private-Sector Small Constellation Satellite Data Product Pilot, now referred to as the Commercial Smallsat Data Acquisition (CSDA) program. The objective of CSDA is to identify, evaluate, and acquire commercial remote sensing data that support NASA’s Earth science research and application activities. The Pilot successfully concluded in early 2020, when CSDA transitioned into a sustained program with on-ramping opportunities for new vendors as the industry emerges with new candidates and capabilities. In October 2019, a Request for Information (RFI) seeking capability statements from parties interested in providing data from spaceborne platforms was released for the CSDA on-ramp #2 evaluations. To be responsive to the RFI, the commercial satellite constellations had to consist of three or more operating spacecraft actively collecting data in a non-geostationary orbit with full latitudinal coverage and be U.S. companies. Two vendors responded to the RFI and were evaluated by a committee composed of NASA ESD leadership, program managers, and scientists. Both vendors satisfied the RFI requirements and were asked to respond to a Request for Proposal (RFP). After review of the proposals, NASA entered into a Blanket Purchase Agreement (BPA) with Airbus Defense and Space GEO, Inc. (Airbus) U.S. in September 2021 and with BlackSky Geospatial Solutions, Inc. (BlackSky) in November 2021. In this report, CSDA provides an evaluation of the usefulness of data provided by the Airbus U.S. Synthetic Aperture Radar (SAR) satellite constellation, consisting of TerraSAR-X (launched in 2007), TanDEM-X (launched in 2010), and PAZ (launched in 2018), for advancing NASA’s Earth system science research and applications. The evaluation of the BlackSky commercial data will be provided in a separate report. To conduct the Airbus evaluation, NASA’s ESD augmented 13 existing research projects that could potentially benefit from, and had the expertise to evaluate, the commercial data being considered for longer-term purchase. Investigators from NASA’s Research and Analysis Program science focus areas and from NASA’s Applied Sciences Program elements participated in the evaluation. A summary of the research areas evaluated by the Principal Investigator (PI) teams is presented in Figure 3. CSDA also funded a dedicated activity to evaluate the satellite data quality (calibration and geolocation) independently by assessing the accuracy of data from Airbus. Evaluation activities were carried out by the selected PIs from December 7, 2022, to December 7, 2023. Delivery of datasets requested by the researchers began in January 2023. The vendors were evaluated on the accessibility of data, accuracy and completeness of metadata, and promptness and quality of user support services. Datasets purchased during the evaluation have been archived by NASA and will be made available to current and future government-funded researchers in accordance with the End User License Agreement (EULA). This synthesis report distills and integrates the findings of research reports commissioned by NASA for the Airbus evaluation. This report also includes recommendations that inform the way ahead for the program. The scientific results from the evaluations demonstrated that the commercial data from Airbus were able to advance NASA research and applications. However, the PIs encountered limitations that diminished the usefulness of the data due to the amount of effort that was required to access, preprocess, and analyze these data. One significant issue encountered was the limited spatial and temporal coverage of the data in the Airbus archive that could be used to conduct time series analyses or assessments over large spatial scales. Overall, however, the utility and the quality of the evaluated data outweighed the difficulties encountered, and NASA has concluded that the Airbus SAR data would complement NASA’s existing Earth observation capabilities and Airbus U.S. would qualify to participate in the sustained phase of the program.

Batuhan Osmanoglu↗

Making NASA GES DISC Level 2 Data GIS Analysis Ready

There are many valuable data hosted by NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) for GIS applications in the areas of extreme weather events, climatic anomaly, and public health. However, using NASA Earth Science Data poses some challenges for GIS users. Many of these users are not experts in Earth Observation and have little knowledge about NASA's Earth science data. In the GIS community, GeoTiff is the most widely used raster format, whereas NASA's data is primarily in complex multidimensional netCDF and HDF formats. This complexity makes it difficult for GIS users, especially those who are unfamiliar with these formats. Although GIS software like ArcGIS has made progress in processing multidimensional netCDF data, certain issues still remain, particularly with level 2 data. In this study, we use TROPSpheric Monitoring instrument (TROPOMI) level 2 data as an example to demonstrate how to make such data GIS analysis ready. The process involves: 1. Creating a feature layer from the TROPOMI level 2 data. 2. Converting the feature layer to a gridded raster dataset. 3. Mosaicking gridded raster datasets into a raster dataset covering the entire desired extent. 4. Generating a symbology with a GIBS-specific style that aligns with the visual standards and requirements of GIBS. 5. Publishing image services. By performing these preprocessing and transformation steps, NASA level 2 data can be made compatible and ready for use within GIS software for various spatial analysis and visualization tasks.

Geographic Information System↗

Curating AI-Ready Datasets for Equity and Environmental Justice: A Data-Centric AI Case Study

An equitable and environmentally just community is essentialin order to avoid disproportionate burden borne by vulnerablecommunities. This need becomes pressing in the aftermathof an extreme event such as disaster or hazard when it is diffi-cult for the governing bodies to implement resource allocationas per the need. Artificial Intelligence (AI) algorithms canhelp surface Equity and Environmental Justice (EEJ) issueswhen trained on EEJ datasets. However, curating AI-readyEEJ training datasets is challenging due to differences in fac-tors such as heterogeneity, resolution, modality, and level ofexpertise in labeling. Additionally, EEJ issues involve sensi-tive information where uncertainties and errors could degradethe performance of AI algorithms. For eg. Error in seasonalcrop yield information can highly affect the prediction of an-nual crop yield. To address these challenges, Data-centricAI (DCAI) methods are employed, which enhance AI algo-rithm performance even with limited training samples. DCAIprioritizes data quality, thereby reducing the adverse effectsof uncertainties and errors during the model training process.This research proposes a novel dataset and benchmark for an-alyzing the effect of the Maui Wildfire of 2023 for Equityand Environmental Justice (EEJ) issues. The proposed datasetaligns with the concepts of DCAI such as annotation quality,data preprocessing, privacy, feature engineering, governanceand provenance. We firmly believe that the proposed datasetwould lay a foundation to implement robust and reliable mod-ern AI algorithms for addressing EEJ issues.

Paridhi Parajuli↗

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗

Open Science Approach to Analyze Climate-Crop Relationships in the US Leveraging GES DISC and Galaxy Workflows

Understanding the intricate relationship between climate variability and agricultural production is crucial for ensuring food security. This study investigates the impact of climate parameters, such as temperature, precipitation, and soil moisture, on major US crop yields. Adopting an open science approach, the study analyzes the impact of climate on agricultural production in the United States. The Galaxy workflow engine serves as the primary tool for integrating climate data from the Goddard Earth Sciences Data and Information Services Center (GES DISC), retrieved via the Giovanni system, with yield statistics from the United States Department of Agriculture’s National Agricultural Statistics Service (USDA NASS). Extensions for reading, preprocessing, and analyzing external data have been developed, enabling the creation of workflows within the Galaxy platform. The development of a reproducible workflow allows for the calculation of seasonal climate averages, which are then assessed for their correlation with crop yields. This methodology ensures the replicability of the research, promoting transparency and collaboration in the scientific community. Correlational and regression analyses have been applied to different sub-zones and crops. The findings from this research offer valuable insights into the relationship between climate parameters and crop yields. These insights contribute to a deeper understanding of climate-crop relationships, providing a solid foundation for informed decision-making in the agricultural sector. The high correlation values indicate a significant relationship between climate parameters and crop yields, underscoring the importance of considering climate factors in agricultural planning and policymaking. This research also exemplifies the power of open science in advancing our understanding of complex environmental and agricultural phenomena. By leveraging open data and services, it provides a robust and replicable framework for future studies in this critical field.

Open science↗

Data Democratization: Challenges and Opportunities

Democratizing Earth data is one of the challenges many organizations around the world face in order to maximize the use of their Earth data for research, applications, education, and societal benefits. For example, at the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), over 1600 global and regional datasets in several NASA Earth science focus areas, including atmospheric composition, water and energy cycles, and climate variability, are archived and distributed to the public. Giovanni, the Geospatial Interactive Online Visualization and Analysis Infrastructure, was developed by GES DISC to facilitate data access and exploration, especially for novice users of Earth science. With Giovanni, users can analyze and visualize over 2000 Earth science variables (e.g., precipitation, aerosol, surface wind) without downloading data, software, the expert understanding of data formats and structures, and coding skills, lowering the barrier to data analysis/comparison by preprocessing and accessing to the data. Results of data analysis and visualization can be accessed in several popular formats (e.g., NetCDF, CSV). As a result of Giovanni's efforts, more than 3000 referral papers have been published in various fields. In spite of this, Giovanni is still difficult to use for some users. For instance, if one searches for "precipitation," it will return over 150 related variables. The question is, which one to use? Furthermore, variables from different data providers (e.g., satellites and models) are named differently with different units, further confusing users, especially those outside the communities. Data democratization is complex and multifaceted. Challenges include service and data discovery, user experiences, visualization, data quality, trustworthiness, and more. In this presentation, we will examine Giovanni as an example of challenges and opportunities in developing data democratization services.

data democratization↗

Developing Natural Language Processing and Supervised Learning Techniques to Classify Mars Tasks

As NASA's Human Research Program (HRP) prepares for long-duration Mars missions, understanding astronaut tasks is crucial. This study, conducted at NASA Glenn Research Center (GRC), employed Natural Language Processing (NLP) and machine learning techniques to analyze and classify Mars tasks. A list of 1,058 Mars tasks was provided by HRP experts including binary labeling of 18 Human System Task Categories (HSTCs). We developed an NLP model using Google's BERT language model to capture the semantic and syntactic nuances of these tasks. Supervised training was initially applied to a subset of the NLP-analyzed tasks to assess the model's effectiveness in classifying the remaining tasks. Incorporating HSTC descriptions significantly enhanced the classification accuracy for 9 out of the 18 HSTCs and reduced training time. To address the issue of severe class imbalance in the HSTC data, we introduced innovative weighting and sampling techniques for data augmentation. We then fine-tune BERT to implement a pairwise relatedness scoring method, allowing us to cluster tasks based on their relatedness and similarity, getting a step closer to labeling the tasks without supervision. In this presentation we guide you through data preprocessing, deciphering key syntax components using BERT, and performing supervised classification of the Mars tasks. This work showcases the potential use of advanced NLP techniques to analyze Mars missions to be incorporated into various crew health and performance analyses.

GenAI↗

Exploring Flooded Fraction Prediction through Machine Learning Models Focusing on Medical Infrastructure in the Southeast U.S. Coastal Areas

Rising sea levels due to climate change increasingly threaten medical infrastructure through flooding. This study develops machine learning models to predict flood exposure for 11,508 medical facilities in the southeastern coastal regions of the United States by integrating datasets including meteorological, hydrological, topographic, and geological data, the Natural Risk Index, and historical flood records from NASA, HIFLD, and FEMA. Six regression models, namely Linear Regression, Support Vector Regression, Random Forest, k-Nearest Neighbors, XGBoost, and Artificial Neural Networks, are trained using 16 explanatory variables identified through literature review and correlation analysis. Data preprocessing employs the SMOGN for class imbalance and Winsorization for outliers. Model performance is evaluated using MAE, MSE, and RMSE, with Random Forest and XGBoost models achieving the highest performance (MSE of 2.58e-5 and 3.69e-5, respectively). This multifactorial approach allows the models to capture complex flood-influencing relationships, enhancing adaptability and performance across geographic regions. Future work focuses on expanding across the U.S. and developing a near real-time flood monitoring system.

Jihoon Chung↗

Predicting Team Functioning in Long Term Space Missions Using Acoustic and Linguistic Measures

Maintaining optimal team functioning is critical for long-duration space exploration missions, yet traditional monitoring methods, such as self-reports and wearable sensors, often impose operational burdens or suffer from bias. This paper investigates a non-intrusive speech-based artificial intelligence (AI) framework to predict degradations in team functioning using data from the Human Exploration Research Analog (HERA) of the U.S. National Aeronautics and Space Administration (NASA). Using acoustic features, linguistic descriptors, and semantic embeddings, we evaluate static non-linear and temporal machine learning models to predict both objective (task accuracy) and subjective (self-reported efficacy and cohesion) team functioning outcomes. Results indicate that temporal models outperform static approaches, with prediction of objective task accuracy in Team Interaction Battery (TIB) improving from near chance to 71%. Self-reported outcomes, including team efficacy and cohesion, are predicted more reliably than task performance, achieving balanced accuracies of up to 85.56% and 78.12%, respectively, and are found to be most strongly associated with acoustic features. In a second interdependent task, the MMSEV–EVA, accuracies of up to 78% are achieved using temporal models with acoustic features. Furthermore, incorporating just 1–2 days of team-specific historical data systematically improved performance, and acoustic markers from informal pre-task interactions provided modest predictive gains. Finally, while automated preprocessing yielded viable accuracy, humancorrected data provided moderate performance gains, though transcription error rates did not significantly correlate with model performance. These findings highlight the potential of speech as a passive, high-fidelity monitoring tool for autonomous habitats.

Temporal modeling↗

Herbaceous Feedstock 2022 State of Technology Report

The U.S. Department of Energy promotes production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement in this mission, Idaho National Laboratory completes an annual SOT report for nth-plant and 1st-plant herbaceous biomass feedstock logistics. The purpose of the SOT is to provide the status of feedstock supply system technology development for herbaceous biomass to biofuels relative to technical targets and cost goals from specific design cases, based on data and experimental results. Although conventional feedstock supply systems form the backbone of the emerging biofuels industry, they have limitations that restrict widespread implementation on a national scale. To meet the demands of the future industry, the feedstock supply system must shift from the conventional system to what has been termed “advanced” supply systems. In advanced designs, a distributed network of aggregation and processing centers, termed “depots,” are employed near the points of biomass production (i.e., the field or forest) to reduce feedstock variability and produce feedstocks of a uniform format, moving toward biomass commoditization. The 2022 Herbaceous SOT is part of a vision of achieving an implemented advanced feedstock supply system, which produces a stable, tradable commodity at the decentralized distributed depot. It utilizes feedstock fractionation by incorporating technologies that can separate the biomass into its anatomical fractions (leaves, husks, stems and cobs) to reduce impurities and produce fractions that satisfy downstream quality considerations. By using a series of air classification steps, this strategy can reduce the extrinsic ash in corn stover and produce enriched tissue fractions that can be blended to a conversion specification or converted individually in optimized biochemical conversion campaigns. Additionally, a majority of the leaves (which do not meet the quality specification) are separated out early and can be supplied to alternate markets. The 2022 Herbaceous SOT incorporates an advanced biomass fractionation and processing system to produce pellets enriched tissues from three-pass corn stover. The resulting enriched pellets are delivered to the biorefinery individually where they can be blended to a specification or converted in campaigns where the conditions are optimized for each tissue. Unused fractions can be sent to a a midstream market or to a different conversion process that is better suited to their properties to offset the cost of the delivered feedstock. The main benefits from the proposed system can be summarized as: (1) $6.86/dry ton (2016$) lower cost for the air classification due to elimination of the requirement to discard the high ash lights fraction; (2) $1.56/dry ton lower delivered cost by selling the unsuitable leaf fraction into the feed market as a midstream co-product (assuming a selling price that is 11% higher than their cost of production); (3) 0.98% increase in carbohydrate content (from 60.16% to 61.14%); and (4) 0.97% decrease in ash content (from 6.00% to 5.03%) compared to the 2021 Herbaceous SOT. Overall, the 2022 nth-plant Herbaceous SOT predicts a modeled delivered feedstock cost of $78.64/dry ton (2016$) if it is assumed that the enriched leaf fraction is sold at its production cost; this is a slight increase of $0.43/dry ton increase from the 2021 Herbaceous SOT nth-Supply case cost. The increased cost derived from a $0.38/dry ton increase in transportation and handling cost to procure more biomass (to replace the enriched leaf fraction that was not delivered to the biorefinery. The total preprocessing cost was $0.27/dry ton higher than the 2021 result because of updates to energy consumption, purchasing price and dry matter loss data for the rotary shear ($3.00/dry ton increase) and the pelleting mill ($4.52/dry ton increase). The data utilized were generated in pilot-scale tests in the Biomass Feedstock National User Facility (BFNUF) at INL and at Forest Concepts, including tests for rotary shear and pelleting of the air classified fractions. A greenhouse gas emissions analysis was performed by Argonne National Laboratory using the most up to date version of the Greenhouse Gases, Regulated Emissions, and Energy use in Transportation model (GREET®). The analysis showed an increase of 17.34 kg CO2e/dry ton from the 2021 SOT (67.71 kg CO2e/ton in the 2021 Herbaceous SOT to 85.05 kg CO2e/ton in the 2022 Herbaceous SOT). The net increase is primarily attributed to increased energy consumption in pelleting mill.

09 BIOMASS FUELS↗

Investigating the Effect of Water on the Mechanical Properties of Cellulose from Multiscale Molecular Dynamics Simulations

Classical molecular dynamics (MD) simulations provide insight into the structure and physicochemical properties of materials with atomic resolution. However, the length and time scales accessible to atomistic MD are orders of magnitude smaller than many relevant processes such as the response of a bulk material to experimentally accessible strain rates, which presents challenges when comparing models to experimental measurements. Bottom-up coarse-graining provides a means for systematically mapping atomistic information to lower resolution models to increase the length and time scales achievable by simulation. Cellulose is an abundant carbohydrate biopolymer with applications to many fields of research, such as materials science and renewable energy, due to its desirable mechanical properties and viability for conversion into biofuel. The effect of moisture content on the Young's modulus of cellulose is of special interest due to its native environment often being in the hydrated secondary plant cell wall and the grinding energy requirements for biomass feedstock preprocessing. The current work investigates the effects of water solvent on the Young's modulus of cellulose calculated from coarse-grained MD mechanical stress simulations. The coarse-grained model was parametrized from atomistic MD calculations of cellulose-cellulose potentials of mean force using umbrella sampling techniques under vacuum and solvated conditions. The Young's moduli of the coarse-grained cellulose assemblies parametrized from cellulose in vacuum or solvated in water were computed via mechanical stress simulations to highlight the importance of capturing solvent interactions for modeling the mechanical behavior of cellulose.

BASIC BIOLOGICAL SCIENCES,RADIATION PROTECTION AND↗