Search NASA⌕ Search

SEARCH · Search NASA

Results for “Online algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Use of Semantic Technology to Create Curated Data Albums

One of the continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available online. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the data sets they need can obtain the specific files using these systems. However, in cases where researchers are interested in studying an event of research interest, they must manually assemble a variety of relevant data sets by searching the different distributed data systems. Consequently, there is a need to design and build specialized search and discovery tools in Earth science that can filter through large volumes of distributed online data and information and only aggregate the relevant resources needed to support climatology and case studies. This paper presents a specialized search and discovery tool that automatically creates curated Data Albums. The tool was designed to enable key elements of the search process such as dynamic interaction and sense-making. The tool supports dynamic interaction via different modes of interactivity and visual presentation of information. The compilation of information and data into a Data Album is analogous to a shoebox within the sense-making framework. This tool automates most of the tedious information/data gathering tasks for researchers. Data curation by the tool is achieved via an ontology-based, relevancy ranking algorithm that filters out non-relevant information and data. The curation enables better search results as compared to the simple keyword searches provided by existing data systems in Earth science.

Ramachandran, Rahul↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Northern Great Plains Disasters: Using Earth Observations to Enhance Flood Monitoring on Tribal Lands in the Northern Great Plains

In 2019, the Great Plains experienced unprecedented catastrophic flooding. Large flood events are predicted to increase in frequency and severity, posing risks to communities in this region, particularly Tribal Nations. We used data from Sentinel-1 C-band Synthetic Aperture Radar (C-SAR), imagery from the Sentinel-2 MultiSpectral Instrument (MSI), and digital elevation models (DEMs) from the Shuttle Radar Topography Mission (SRTM) within Google Earth Engine to map historical floods in the region beginning in 2014 with particular attention to the Rosebud Sioux Reservation and the tribal lands of other Great Plains Tribal Water Alliance members. This historical mapping used C-SAR for a combined method approach with a Z-score algorithm in addition to an index for flooded short vegetation. We also developed a flood risk map by weighting different flood predictor variables according to flood risk literature. These variables included soil drainage from the Soil Survey Geographic Database (SSURGO); elevation, slope, and Topographic Wetness Index (TWI) derived from digital elevation models; precipitation from Climate Hazards Group InfraRed Precipitation with Station data (CHIRPS); land cover from the National Land Cover Database (NLDC); and Normalized Difference Vegetation Index (NDVI) derived from Landsat 8 Operational Land Imager (OLI). From the flood extent and risk maps, we identified widespread flooding in short vegetation (including cropland) and noted flood susceptibility in regions exhibiting high social vulnerability and low community resilience (FEMA indices). We created an ArcGIS Online StoryMap to share project background, results, and data. Additionally, we provided a written tutorial so partners may replicate the flood mapping for future flood events.

Anna Ballasiotes↗

Three-dimensional modeling of hyphal fusion, branching, and nutrient transport in filamentous fungi

Fungi exhibit behaviors distinct from other microbes. Filamentous fungi grow by extending complex networks of branched filaments collectively referred to as the mycelium. These networks can expand over large distances and traverse low-nutrient areas by translocating nutrients through the filament network. This spatial characteristic makes filamentous fungi crucial for soil ecosystems, supporting stable microbial communities and promoting plant growth. However, simulating these behaviors is complex. The elongated nature of fungal compartments results in different mechanical interactions compared to the commonly modeled spherical bacteria. These detailed hyphal mechanics require specialized consideration and are often excluded from conventional fungal simulation packages. Additionally, the extensive fungal networks in nature demand computationally intensive simulations, necessitating high-performance algorithms. Therefore, realistic fungi simulations require specialized software. Here, we introduce a fungal modeling expansion to the high-performance biological modelling and interface exchange (bmx) software suite. bmx leverages adaptive mesh refinement in AMReX for chemical diffusion and incorporates a full mechanical model for bacterial cells, accelerated by GPUs. By extending bmx to model filamentous particles, we demonstrate the formation of complex filament networks through interactions like hyphal branching and fusion (anastomosis). We show that the networks produced match real-world fungal structures through various metrics. This work supports computational studies of fungal growth dynamics and can be adapted to investigate the growth of other filamentous structures in biology or materials science. The expanded-BMX package is open-sourced and is available online.

Cell mechanics↗

Diagnosability-Based Sensor Placement through Structural Model Decomposition

Systems health management, and in particular fault diagnosis, is important for ensuring safe, correct, and efficient operation of complex engineering systems. The performance of an online health monitoring system depends critically on the available sensors of the system. However, the set of selected sensors is subject to many constraints, such as cost and weight, and hence, these sensors must be selected judiciously. This paper presents an offline design-time sensor placement approach for complex systems. Our diagnosis method is built upon the analysis of model-based residuals, which are computed using structural model decomposition. Sensor placement in this framework manifests as a residual selection problem, and we aim to find the set of residuals that achieves single-fault diagnosability of the system, uses the minimum number of sensors, and corresponds to the best model decomposition for the best distribution of the diagnosis system. We present a set of algorithms for solving this problem and compare their performance in terms of computational complexity and optimality of solutions. We demonstrate the approach using a benchmark multi-tank system.

Daigle, Matthew↗

Brain-Computer Interfaces for 1-D and 2-D Cursor Control: Designs Using Volitional Control of the EEG Spectrum or Steady-State Visual Evoked Potentials

We have developed and tested two EEG-based brain-computer interfaces (BCI) for users to control a cursor on a computer display. Our system uses an adaptive algorithm, based on kernel partial least squares classification (KPLS), to associate patterns in multichannel EEG frequency spectra with cursor controls. Our first BCI, Target Practice, is a system for one-dimensional device control, in which participants use biofeedback to learn voluntary control of their EEG spectra. Target Practice uses a KF LS classifier to map power spectra of 30-electrode EEG signals to rightward or leftward position of a moving cursor on a computer display. Three subjects learned to control motion of a cursor on a video display in multiple blocks of 60 trials over periods of up to six weeks. The best subject s average skill in correct selection of the cursor direction grew from 58% to 88% after 13 training sessions. Target Practice also implements online control of two artifact sources: a) removal of ocular artifact by linear subtraction of wavelet-smoothed vertical and horizontal EOG signals, b) control of muscle artifact by inhibition of BCI training during periods of relatively high power in the 40-64 Hz band. The second BCI, Think Pointer, is a system for two-dimensional cursor control. Steady-state visual evoked potentials (SSVEP) are triggered by four flickering checkerboard stimuli located in narrow strips at each edge of the display. The user attends to one of the four beacons to initiate motion in the desired direction. The SSVEP signals are recorded from eight electrodes located over the occipital region. A KPLS classifier is individually calibrated to map multichannel frequency bands of the SSVEP signals to right-left or up-down motion of a cursor on a computer display. The display stops moving when the user attends to a central fixation point. As for Target Practice, Think Pointer also implements wavelet-based online removal of ocular artifact; however, in Think Pointer muscle artifact is controlled via adaptive normalization of the SSVEP. Training of the classifier requires about three minutes. We have tested our system in real-time operation in three human subjects. Across subjects and sessions, control accuracy ranged from 80% to 100% correct with lags of 1-5 seconds for movement initiation and turning.

Trejo, Leonard J.↗

Kernelized approaches to streaming compression of scientific data

In this paper three algorithms are developed for the streaming compression of scientific data. The algorithms presented are reliant on the theory of vector-valued reproducing kernel Hilbert spaces and operator valued kernel. Further, the scientific data is modeled as a snapshot of time dependent vector field F(x, t) over a manifold M and the recovery of the data is framed as a learning problem. These processes are then appropriately modified and ana lyzed for the streaming scenario in which data is generated without the ability to revisit past entries.

97 MATHEMATICS AND COMPUTING↗

RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input Preprocessing

Ensuring high-quality recommendations for newly onboarded users requires the continuous retraining of Deep Learning Recommendation Models (DLRMs) with freshly generated data. To serve the online DLRM retraining, existing solutions use hundreds of CPU computing nodes designated for input preprocessing, causing significant power consumption that surpasses even the power usage of GPU trainers. To this end, we propose RAP, an end-to-end DLRM training framework that supports Resource-aware Automated GPU sharing for DLRM input Preprocessing and Training. The core idea of RAP is to accurately capture the remaining GPU computing resources during DLRM training for input preprocessing, achieving superior training efficiency without requiring additional resources. Specifically, RAP utilizes a co-running cost model to efficiently assess the costs of various input preprocessing operations, and it implements a resource-aware horizontal fusion technique that adaptively merges smaller kernels according to GPU availability, circumventing any interference with DLRM training. In addition, RAP leverages a heuristic searching algorithm that jointly optimizes both the input preprocessing graph mapping and the co-running schedule to maximize the end-to-end DLRM training throughput. The comprehensive evaluation shows that RAP achieves 78.3× speedup on average over CPU-based DLRM input preprocessing frameworks. In addition, the end-to-end training throughput of RAP is only 2.04% lower than the ideal case, which has no input preprocessing overhead.

Wang, Zheng↗

Modernizing Mechatronics Course With Quantum Engineering

Mechatronics is the synergistic application of mechanics, electronics, control engineering, and computer science in the development of electromechanical products and systems, through integrated design. This paper proposes to extend the mechatronics course beyond traditional engineering topics, and to modernize the mechatronics instructions with complementary quantum engineering topics. With the recent rapid advances in quantum technologies such as quantum communications, sensing, computers, and algorithms, it is imperative to train the next generation of engineers and prepare them for their future careers in the ever-changing industry in such areas. Furthermore, due to such progress and advances in the fields associated with quantum mechanics, the integration of quantum technologies with classical mechanical systems will be inevitable both in terms of educational and technological standpoints in future. To address the educational needs of the future engineers in such areas of significant importance, quantum entanglement and quantum cryptography experiments, as two fundamental topics in quantum mechanics, are brought into the mechatronics course in an initiative that is reported in this paper. The integrated quantum and mechatronics topics also provides opportunities for open discussions on exploring the interface of quantum technologies and classical engineering systems, which can potentially push the engineering boundaries beyond classical possibilities by accessing the quantum advantages. An innovative online remote demonstration of such quantum experiments are developed and presented to the students. This course has been offered to undergraduate students once with successful results. The students were able to remotely access the experiments, perform the experiments and collect data. The successful result of such quantum experiments is also reflected in a course survey, presented in this paper, even though the quantum mechanics topics offered in this course are unfamiliar to engineering students and hence more challenging. The paper reports, and aims to promote, the integration of selected quantum technology topics with the mechatronics course for training engineering students in this rapidly growing area.

Ghazinejad, Maziar↗

Computationally efficient and error aware surrogate construction for numerical solutions of subsurface flow through porous media

Limiting the injection rate to restrict the pressure below a threshold at a critical location can be an important goal of simulations that model the subsurface pressure between injection and extraction wells. The pressure is approximated by the solution of Darcy’s partial differential equation for a given permeability field. The subsurface permeability is modeled as a random field since it is known only up to statistical properties. This induces uncertainty in the computed pressure. Solving the partial differential equation for an ensemble of random permeability simulations enables estimating a probability distribution for the pressure at the critical location. These simulations are computationally expensive, and practitioners often need rapid online guidance for real-time pressure management. An ensemble of numerical partial differential equation solutions is used to construct a Gaussian process regression model that can quickly predict the pressure at the critical location as a function of the extraction rate and permeability realization. The Gaussian process surrogate analyzes the ensemble of numerical pressure solutions at the critical location as noisy observations of the true pressure solution, enabling robust inference using the conditional Gaussian process distribution. Our first novel contribution is to identify a sampling methodology for the random environment and matching kernel technology for which fitting the Gaussian process regression model scales as O ( n log n ) instead of the typical O ( n 3 ) rate in the number of samples n used to fit the surrogate. The surrogate model allows almost instantaneous predictions for the pressure at the critical location as a function of the extraction rate and permeability realization. Our second contribution is a novel algorithm to calibrate the uncertainty in the surrogate model to the discrepancy between the true pressure solution of Darcy’s equation and the numerical solution. Finally, although our method is derived for building a surrogate for the solution of Darcy’s equation with a random permeability field, the framework broadly applies to solutions of other partial differential equations with random coefficients.

54 ENVIRONMENTAL SCIENCES↗

Statistical Considerations of Data Processing in Giovanni Online Tool

The GES DISC Interactive Online Visualization and Analysis Infrastructure (Giovanni) is a web-based interface for the rapid visualization and analysis of gridded data from a number of remote sensing instruments. The GES DISC currently employs several Giovanni instances to analyze various products, such as Ocean-Giovanni for ocean products from SeaWiFS and MODIS-Aqua; TOMS & OM1 Giovanni for atmospheric chemical trace gases from TOMS and OMI, and MOVAS for aerosols from MODIS, etc. (http://giovanni.gsfc.nasa.gov) Foremost among the Giovanni statistical functions is data averaging. Two aspects of this function are addressed here. The first deals with the accuracy of averaging gridded mapped products vs. averaging from the ungridded Level 2 data. Some mapped products contain mean values only; others contain additional statistics, such as number of pixels (NP) for each grid, standard deviation, etc. Since NP varies spatially and temporally, averaging with or without weighting by NP will be different. In this paper, we address differences of various weighting algorithms for some datasets utilized in Giovanni. The second aspect is related to different averaging methods affecting data quality and interpretation for data with non-normal distribution. The present study demonstrates results of different spatial averaging methods using gridded SeaWiFS Level 3 mapped monthly chlorophyll a data. Spatial averages were calculated using three different methods: arithmetic mean (AVG), geometric mean (GEO), and maximum likelihood estimator (MLE). Biogeochemical data, such as chlorophyll a, are usually considered to have a log-normal distribution. The study determined that differences between methods tend to increase with increasing size of a selected coastal area, with no significant differences in most open oceans. The GEO method consistently produces values lower than AVG and MLE. The AVG method produces values larger than MLE in some cases, but smaller in other cases. Further studies indicated that significant differences between AVG and MLE methods occurred in coastal areas where data have large spatial variations and a log-bimodal distribution instead of log-normal distribution.

Suhung, Shen↗

Modeling for Battery Prognostics

For any battery-powered vehicles (be it unmanned aerial vehicles, small passenger aircraft, or assets in exoplanetary operations) to operate at maximum efficiency and reliability, it is critical to monitor battery health as well performance and to predict end of discharge (EOD) and end of useful life (EOL). To fulfil these needs, it is important to capture the battery's inherent characteristics as well as operational knowledge in the form of models that can be used by monitoring, diagnostic, and prognostic algorithms. Several battery modeling methodologies have been developed in last few years as the understanding of underlying electrochemical mechanics has been advancing. The models can generally be classified as empirical models, electrochemical engineering models, multi-physics models, and molecular/atomist. Empirical models are based on fitting certain functions to past experimental data, without making use of any physicochemical principles. Electrical circuit equivalent models are an example of such empirical models. Electrochemical engineering models are typically continuum models that include electrochemical kinetics and transport phenomena. Each model has its advantages and disadvantages. The former type of model has the advantage of being computationally efficient, but has limited accuracy and robustness, due to the approximations used in developed model, and as a result of such approximations, cannot represent aging well. The latter type of model has the advantage of being very accurate, but is often computationally inefficient, having to solve complex sets of partial differential equations, and thus not suited well for online prognostic applications. In addition both multi-physics and atomist models are computationally expensive hence are even less suited to online application An electrochemistry-based model of Li-ion batteries has been developed, that captures crucial electrochemical processes, captures effects of aging, is computationally efficient, and is of suitable accuracy for reliable EOD prediction in a variety of operational profiles. The model can be considered an electrochemical engineering model, but unlike most such models found in the literature, certain approximations are done that allow to retain computational efficiency for online implementation of the model. Although the focus here is on Li-ion batteries, the model is quite general and can be applied to different chemistries through a change of model parameter values. Progress on model development, providing model validation results and EOD prediction results is being presented.

Prognostics↗

An Online Tool for Preliminary Design and Techno-Economic Analysis of District Geothermal Heating and Cooling Systems

District geothermal heating and cooling systems (DGHCS) have significant benefits for reducing energy consumption as well as building- and grid-level peak electric demand. Currently, no publicly available tools are available to effectively design and conduct techno-economic analysis of DGHCS. GeoWISE was originally developed for preliminary design and techno-economic analysis of geothermal heating and cooling systems in an individual commercial or residential building. This paper introduces recent upgrades of GeoWISE that allow users to design and conduct techno-economic analysis of DGHCS. Several new features are implemented in GeoWISE to allow selection and specification of multiple new or existing buildings. A database of information for over 125 million existing U.S. buildings was used in GeoWISE that allows users easily locate existing buildings of interest based on street addresses, and optionally edit information of the buildings (e.g., footprint, vintage, principal functions, number of floors, window-to-wall ratio). Unique energy simulation models of the selected buildings are then automatically created using the Automatic Building Energy Modeling (AutoBEM) and EnergyPlus simulations are performed to predict thermal loads of the buildings. A simplified DGHCS is then designed and simulated to predict its energy use. A central borehole heat exchanger (BHE) of the DGHCS is sized using the RowWise algorithm of GHEDesigner to meet the thermal loads within user-specified land areas for installing BHE. The upgraded GeoWISE reports the needed capacity of heating and cooling equipment in each building, design of the central BHE, energy consumption reduction, and energy cost saving resulting from using DGHCS compared with conventional HVAC systems. A case study is showcased using the upgraded GeoWISE to design and conduct techno-economic analysis of a simplified DGHCS.

Prem Anand Jayaprabha, Jyothis Anand [ORNL] (ORCID↗

A Satellite-Derived Climate-Quality Data Record of the Clear-Sky Surface Temperature of the Greenland Ice Sheet

We have developed a climate-quality data record of the clear-sky surface temperature of the Greenland Ice Sheet using the Moderate-Resolution Imaging Spectroradiometer (MODIS) ice-surface temperature (1ST) algorithm. A climate-data record (CDR) is a time series of measurements of sufficient length, consistency, and continuity to determine climate variability and change. We present daily and monthly MODIS ISTs of the Greenland Ice Sheet beginning on 1 March 2000 and continuing through 31 December 2010 at 6.25-km spatial resolution on a polar stereographic grid. This record will be elevated in status to a CDR when at least nine more years of data become available either from MODIS Terra or Aqua, or from the Visible Infrared Imager Radiometer Suite (VIIRS) to be launched in October 2011. Our ultimate goal is to develop a CDR that starts in 1981 with the Advanced Very High Resolution (AVHRR) Polar Pathfinder (APP) dataset and continues with MODIS data from 2000 to the present, and into the VIIRS era. Differences in the APP and MODIS cloud masks have so far precluded the current 1ST records from spanning both the APP and MODIS time series in a seamless manner though this will be revisited when the APP dataset has been reprocessed. The complete MODIS 1ST daily and monthly data record is available online.

Hall, Dorothy K.↗

Modeling the Swift BAT Trigger Algorithm with Machine Learning

To draw inferences about gamma-ray burst (GRB) source populations based on Swift observations, it is essential to understand the detection efficiency of the Swift burst alert telescope (BAT). This study considers the problem of modeling the Swift BAT triggering algorithm for long GRBs, a computationally expensive procedure, and models it using machine learning algorithms. A large sample of simulated GRBs from Lien et al. (2014) is used to train various models: random forests, boosted decision trees (with AdaBoost), support vector machines, and artificial neural networks. The best models have accuracies of approximately greater than 97% (approximately less than 3% error), which is a significant improvement on a cut in GRB flux which has an accuracy of 89:6% (10:4% error). These models are then used to measure the detection efficiency of Swift as a function of redshift z, which is used to perform Bayesian parameter estimation on the GRB rate distribution. We find a local GRB rate density of eta(sub 0) approximately 0.48(+0.41/-0.23) Gpc(exp -3) yr(exp -1) with power-law indices of eta(sub 1) approximately 1.7(+0.6/-0.5) and eta(sub 2) approximately -5.9(+5.7/-0.1) for GRBs above and below a break point of z(sub 1) approximately 6.8(+2.8/-3.2). This methodology is able to improve upon earlier studies by more accurately modeling Swift detection and using this for fully Bayesian model fitting. The code used in this is analysis is publicly available online.

gamma rays: general↗

Online Assessment of Satellite-Derived Global Precipitation Products

Precipitation is difficult to measure and predict. Each year droughts and floods cause severe property damages and human casualties around the world. Accurate measurement and forecast are important for mitigation and preparedness efforts. Significant progress has been made over the past decade in satellite precipitation product development. In particular, products' spatial and temporal resolutions as well as timely availability have been improved by blended techniques. Their resulting products are widely used in various research and applications. However biases and uncertainties are common among precipitation products and an obstacle exists in quickly gaining knowledge of product quality, biases and behavior at a local or regional scale, namely user defined areas or points of interest. Current online inter-comparison and validation services have not addressed this issue adequately. To address this issue, we have developed a prototype to inter-compare satellite derived daily products in the TRMM Online Visualization and Analysis System (TOVAS). Despite its limited functionality and datasets, users can use this tool to generate customized plots within the United States for 2005. In addition, users can download customized data for further analysis, e.g. comparing their gauge data. To meet increasing demands, we plan to increase the temporal coverage and expanded the spatial coverage from the United States to the globe. More products have been added as well. In this poster, we present two new tools: Inter-comparison of 3B42RT and 3B42 Inter-comparison of V6 and V7 TRMM L-3 monthly products The future plans include integrating IPWG (International Precipitation Working Group) Validation Algorithms/statistics, allowing users to generate customized plots and data. In addition, we will expand the current daily products to monthly and their climatology products. Whenever the TRMM science team changes their product version number, users would like to know the differences by inter-comparing both versions of TRMM products in their areas of interest. Making this service available to users will help them to better understand associated changes. We plan to implement this inter-comparison in TRMM standard monthly products with the IPWG algorithms. The plans outlined above will complement and accelerate the existing and ongoing validation activities in the community as well as enhance data services for TRMM and the future Global Precipitation Mission (GPM).

Liu, Zhong↗

Aggregation Tool to Create Curated Data albums to Support Disaster Recovery and Response

Despite advances in science and technology of prediction and simulation of natural hazards, losses incurred due to natural disasters keep growing every year. Natural disasters cause more economic losses as compared to anthropogenic disasters. Economic losses due to natural hazards are estimated to be around $6-$10 billion dollars annually for the U.S. and this number keeps increasing every year. This increase has been attributed to population growth and migration to more hazard prone locations such as coasts. As this trend continues, in concert with shifts in weather patterns caused by climate change, it is anticipated that losses associated with natural disasters will keep growing substantially. One of challenges disaster response and recovery analysts face is to quickly find, access and utilize a vast variety of relevant geospatial data collected by different federal agencies such as DoD, NASA, NOAA, EPA, USGS etc. Some examples of these data sets include high spatio-temporal resolution multi/hyperspectral satellite imagery, model prediction outputs from weather models, latest radar scans, measurements from an array of sensor networks such as Integrated Ocean Observing System etc. More often analysts may be familiar with limited, but specific datasets and are often unaware of or unfamiliar with a large quantity of other useful resources. Finding airborne or satellite data useful to a natural disaster event often requires a time consuming search through web pages and data archives. Additional information related to damages, deaths, and injuries requires extensive online searches for news reports and official report summaries. An analyst must also sift through vast amounts of potentially useful digital information captured by the general public such as geo-tagged photos, videos and real time damage updates within twitter feeds. Collecting and aggregating these information fragments can provide useful information in assessing damage in real time and help direct recovery efforts. The search process for the analyst could be made much more efficient and productive if a tool could go beyond a typical search engine and provide not just links to web sites but actual links to specific data relevant to the natural disaster, parse unstructured reports for useful information nuggets, as well as gather other related reports, summaries, news stories, and images. This presentation will describe a semantic aggregation tool developed to address similar problem for Earth Science researchers. This tool provides automated curation, and creates "Data Albums" to support case studies. The generated "Data Albums" are compiled collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; information about the event contained in news reports, and images or videos to supplement research analysis. An ontology-based relevancy-ranking algorithm drives the curation of relevant data sets for a given event. This tool is now being used to generate a catalog of Hurricane Case Studies at Global Hydrology Resource Center (GHRC), one of NASA's Distribute Active Archive Centers. Another instance of the Data Albums tool is currently being created in collaboration with NASA/MSFC's SPoRT Center, which conducts research on unique NASA products and capabilities that can be transitioned to the operational community to solve forecast problems. This new instance focuses on severe weather to support SPoRT researchers in their model evaluation studies

Ramachandran, Rahul↗

Application of ML/AI for Identifying Earth Science Datasets in Research Publications

NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).

Irina Gerasimov↗