Search NASA⌕ Search

SEARCH · Search NASA

Results for “Jupyter notebook”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Jupyter Notebook Environment For Multibody Dynamics

DARTS is a rigid/flexible multibody dynamics toolkit for themodeling and simulation of aerospace and robotic vehicles forengineering applications. In this paper we describe an on-line,browser-based environment using Jupyter notebooks to supporttraining needs for the DARTS software. The suite of curated tutorial notebooks is organized into different topic areas, and intomultiple themes within each topic area. The notebooks within atheme use a progression of examples for users to expand theirunderstanding of the software. The topic areas include one onthe DARTS multibody dynamics software and another one on thetheory underlying the multibody dynamics formulation. We alsodescribe a number of Jupyter extensions that were used - andsome developed in house - to enhance the notebook interface foruse with the dynamics simulation software. One significant extension we implemented allows the embedding of live 3D visualizations within simulation notebooks.

Gaut, Aaron↗

SAGE III/ISS Rapid Data Analysis Through Dashboarding with Jupyter Notebooks

Spaceborne remote sensing observations of Earth’s atmosphere produce significant quantities of data over the life of each mission. In the case of the Stratospheric Aerosol and Gas Experiment III on the International Space Station (SAGE III/ISS) nearly four years of vertical profiles of atmospheric ozone, water vapor, and nitrogen dioxide concentrations as well as aerosol extinction coefficients have been released. The dichotomy of the desire for both long-term trends in the atmospheric state alongside the assessment of short-term impacts of major disruptive events such as volcanic eruptions and pyrocumulus injections requires agile tools to handle these cases in near real-time as new data are produced. The analysis landscape is further complicated by the desire to compare results between the numerous contemporary observations available for a given dataset. The SAGE III/ISS team has developed a suite of tools leveraging modern web-based frameworks allowing members to interact with a dashboard-style interface to load the data record, assess new profiles as they are generated and in ensemble, compare between species, and additionally add in measurements observed by other platforms as necessary. Leveraging a commonly packaged data format of NetCDF alongside the Python Jupyter Notebook framework, the data can be served to interested parties from an analysis server while still runnable on personal systems if required. This presentation illustrates the ecosystem developed by the SAGE III/ISS team, the applicability to measurements made by any limb-observing platform, and the benefit to transforming routine analyses into readily accessible dynamic plots. Frameworks currently exist at larger scales with projects such as GIOVANNI, and this illustration seeks to show that similar frameworks are accessible and possible within the local research environment while simultaneously unloading human processing cycles for more specialized analysis tasks.

Dashboarding↗

GES DISC Data Recipes in Jupyter Notebooks

The Earth Science Data and Information System (ESDIS) Project manages twelve Distributed Active Archive Centers (DAACs) which are geographically dispersed across the United States. The DAACs are responsible for ingesting, processing, archiving, and distributing Earth science data produced from various sources (satellites, aircraft, field measurements, etc.). In response to projections of an exponential increase in data production, there has been a recent effort to prototype various DAAC activities in the cloud computing environment. This, in turn, led to the creation of an initiative, called the Cloud Analysis Toolkit to Enable Earth Science (CATEES), to develop a Python software package in order to transition Earth science data processing to the cloud. This project, in particular, supports CATEES and has two primary goals. One, to transition data recipes created by the Goddard Earth Science Data and Information Service Center (GES DISC) into an interactive and educational environment using JupyterNotebooks. Two, to acclimate Earth scientists to cloud computing. To accomplish these goals, we create JupyterNotebooks to compartmentalize the different steps of data analysis and help users obtain and parse data from the command line. We also develop a Docker container, comprised of Jupyter Notebooks, Python dependencies, and command line tools, and configure it into an easy-to-deploy package. The end result is an end-to-end product that simulates the use case of end users working in the cloud computing environment.

discoverability↗

Connecting Users and Applications with Po.daac Hosted GHRSST Data

The 80+ GHRSST public datasets represent a rich resource for sea surface temperature research and applications given their time series length, resolution, spatial coverage, varying measurement types and processing levels, and availability in the full spectrum of PO.DAAC tools and services ecosystem. The PO.DAAC has created a publicly accessible recipe suite for the user community to perform straightforward yet powerful computations on GHRSST data using python recipes, Jupyter notebooks, R, Matlab, and the NCO programming language. These recipes include numerical computations for regional and global SST trends, anomaly derivations, EOF analysis, climate signal reproduction, and ocean phenology. For example, one recipe reproduces a famous SST based warming figure from the Fourth National Climate Assessment (USA) while another focuses on quantifying the regional changes in ocean SST phenology. Most are python-based while some contain hybrid calls and leverage the NCO programming interface too. All are available on the PO.DAAC user forum (https://podaac.jpl.nasa.gov/forum/) and/or via the open source NASA GitHub repository (https://github.com/nasa/podaac_tools_and_services). Several are available in the Jupyter notebook framework including podaacypy (https://github.com/nasa/podaacpy), a recipe for GHRSST granule metadata discovery and application, and more recently a Jupyter notebook developed to support data analysis and visualization of a cloud-based Zarr formatted Level 4 MUR dataset in the AWS Open Data Registry. Throughout the summer of 2020, the PO.DAAC intends to add and migrate more of its numerical recipes to the Jupyter notebook framework and publish them on its open source GitHub repository.

Gentemann, Chelle↗

A Novel Architecture of JupyterHub on Amazon Elastic Kubernetes Service for Open Data Cube Sandbox

The Open Data Cube (ODC) initiative, with support from the Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has developed a state-of-the-art suite of software tools and products to facilitate the analysis of Earth Observation data. This paper presents a short summary of our novel architecture approach in a project related to the Open Data Cube (ODC) community that provides users with their own ODC sandbox environment. Users can have a sandbox environment all to themselves for the purpose of running Jupyter notebooks that leverage the ODC. This novel architecture layout will remove the necessity of hosting multiple users on a single Jupyter notebook server and provides better management tooling for handling resource usage. In this new layout each user will have their own credentials which will give them access to a personal Jupyter notebook server with access to a fully deployed ODC environment enabling exploration of solutions to problems that can be supported by Earth observation data.

Open Data Cube↗

Automating Surface Attitude Positioning and Pointing Operations for Mars 2020

The Surface Attitude Positioning and Pointing (SAPP) subsystem of the Mars Perseverance rover keeps track of the rover’s position and attitude on the surface of Mars. The SAPP Downlink Engineering Operations team members receive data from the rover on a daily basis. They must interpret the data to make sure the rover is staying safe and to support uplink planning. The SAPP team keeps track of the error growth in the rover’s attitude estimate due to noise in the Rover Inertial Measurement Unit’s (RIMU) gyroscopes used to propagate that attitude estimate whenever the rover is moving. Whenever this error grows to a particular threshold, SAPP is responsible for updating the onboard attitude knowledge using the RIMU’s accelerometers to estimate rover roll and pitch and sun imaging to estimate rover yaw, thereby reducing this attitude estimation error. Accurate attitude estimation is required so that the rover can successfully point its High Gain Antenna (HGA) to receive information from Earth and as a backup to the Mars orbiters used for sending data from the rover to Earth, point instruments on its Remote Sensing Mast (RSM), and support safe movement and placement of instruments by the rover’s ARM relative to the Martian surface. The Mars 2020 Engineering Operations team has been working to increase the operational efficiency of the mission and eventually move to a five-hour timeline for daily operations. In pursuit of this goal, the SAPP Engineering Operations team has automated their downlink process by developing a centralized Jupyter notebook to analyze the data received daily from the rover. The SAPP downlink Jupyter notebook automatically collects the data relevant to the SAPP subsystem and visualizes this information in plots and tables that can be easily read by downlink operators to aid them in assessing the status of the subsystem. Various Application Programming Interfaces (APIs) have been incorporated into the downlink daily notebook to automate the collection and posting of data, such as gathering and posting data products to the cloud. The SAPP team has also developed a SAPP downlink software library that includes functions to aid the notebook in processing data. In addition to assessing the SAPP subsystem on a daily basis, operators need to assess the long-term trending behavior of the subsystem over time. An automated trending process has been developed to collect information from the daily notebooks in order to plot and analyze that data in a centralized place. These daily and trending processes have expedited the SAPP downlink assessment and laid the groundwork to completely automate the SAPP downlink process so that SAPP operators are unnecessary unless something unexpected occurs. This paper will provide an overview of the functions that the SAPP subsystem carries out on a daily basis, and will then dive into the automations that have been developed for daily and trending downlink assessment. An assessment of the downlink efficiency will be provided, along with a summary of lessons learned and work to go. Finally, the authors will discuss how these types of automated spacecraft health assessments could be more broadly used within mission operations.

Zarifian, Anais↗

PACE Water Resources: Demonstrating the Use of NASA's PACE Hyperspectral Ocean Color Instrument Data for Enhanced Coastal Management

This project developed tools to support the future use of Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) hyperspectral imagery in water resource monitoring and research by NASA DEVELOP teams and members of the PACE applications community. We sought to address a need for support in processing and visualizing hyperspectral PACE Ocean Color Instrument (OCI) data among researchers and decision-makers working in coastal water quality management and harmful algal bloom (HAB) monitoring. To supplement the day of simulated PACE imagery available, we used Aqua MODIS earth observations with Level 3 processing from March 2022 to build a Python graphical user interface (GUI) for visualizing ocean biogeochemical parameters relevant to the early detection and monitoring of HABs. We used simulated PACE OCI Level 2 data derived from the Python Top of Atmosphere Simulation Tool (PyTOAST) to build Jupyter Notebooks for band subset and selection. The Level 3 PACE Viewer components support users with quick visualizations as well as the creation of geoTIFFs and time-series. The Level 2 Jupyter Notebooks address users’ concerns over the volume and complexity of hyperspectral imagery. The PACE Viewer is useful for visual inspection and netCDF data processing but should not be used for geospatial analysis. Once PACE launches, this tool will alleviate the technical burdens of working with hyperspectral data and support the early detection and monitoring of HABs using PACE satellite imagery.

Python Top of Atmosphere Simulation Tool↗

Open-source Numerical Modeling of Solidification Cracking Susceptibility: Application to Refractory Alloy Systems

Introduction. Alloys such as aluminum, nickel-base, and austenitic stainless steels are susceptible to solidification cracking during welding and 3D printing. Compositional optimization is one method used to effectively mitigate solidification cracking of those alloy systems. With the surge in hypersonic and in-space propulsion activities, refractory metals (Nb, Mo, Ta, W, and Re) and their alloy derivatives are increasing in importance due to their extreme high melting point and retention of high-temperature strength; however, their chemistry was most typically optimized to promote ductility during mechanical operations such as drawing and forming. Welding of such alloys has been a challenge due to a number of issues including solidification cracking, atmospheric contamination (O, C, and N), as well as a shift in ductile-to-brittle transition to higher temperature following grain growth induced by welding. Compositional optimization of refractory alloys for solidification cracking resistance in particular is desirable as their usage increases with the advent of advanced manufacturing methods such as 3D printing. This work evaluates the effect of compositional variation in refractory metal systems on the solidification cracking susceptibility with the goals of optimizing existing alloys and joining process techniques, and formulating new alloys with increased solidification cracking resistance. Experimental Procedures. A python code was developed in a Jupyter notebook environment (Michael and Sowards, 2023) to facilitate the calculation of crack susceptibility index proposed by Kou (2015). Composition is entered as a single point, or as a 1-D or 2-D array. The notebook calls pycalphad (Otis and Liu, 2017 and Bocklund et al, 2020) to calculate the evolution of fraction solid as a function of temperature (under either Scheil or equilibrium assumptions) and then evaluates steepness of the fraction solid curve near the terminal stage of solidification to predict solidification cracking resistance. Open source thermodynamic databases available at online repositories are used (van de Walle). The process is setup in an automated fashion to generate plots that show variation in solidification cracking susceptibility according to composition on 1-D line plots or 2-D contour plots. The Jupyter notebook and crack susceptibility algorithm was also integrated with a widely used commercial CALPHAD code for validation and alloy exploration. Results and Discussion. The crack susceptibility model was first validated against a series of refractory alloy compositions evaluated in past work which utilized a specialized Varestraint test built inside a vacuum chamber environment (Lessman and Gold, 1971). The alloys tested in the Varestraint apparatus included T-111 (Ta-8W-2Hf), ASTAR-811C (Ta-8W-1Re-0.7Hf-0.025C), FS-85 (Nb-27Ta-10W-1Zr), T-222 (Ta-9.6W-2.4Hf-0.01C), Ta-10W, B-66 (Nb-5Mo-5V-1Zr), and SCb-291 (Nb-10W-10Ta). The initial test of the model showed a strong correlation with empirical Varestraint data, i.e., a Spearman rank correlation between model predictions and hot cracking measurements was observed to be greater than 0.8. Following the validation, a set of refractory metal binary mixtures was investigated to evaluate sensitivity of Nb, Mo, W, and Ta to C, N, and O content. A series of plots were produced that suggest ppmw ranges of C, N, and O where solidification cracking increases significantly and reaches a maximum. Also comparative ranking of each primary refractory metal to each interstitial was produced. For example C produces greater cracking response in Mo whereas O produces greater cracking response in Ta and Nb. Such compositional values have utility in setting limits on pickup of these interstitial elements during welding and printing rather than using a one-size-fits-all approach. Furthermore, the results have use in determining additive powder recycling requirements, which is especially pertinent for refractory metal powders due to their high cost compared to conventional alloys. Another application created thousands of hypothetical alloys within the nominal specified composition range of two widely used refractory alloys C103 (Nb-10Hf-1Ti) and TZM (Mo-0.5Ti-0.1Zr). The cracking index was calculated for the alloys and results were fed into machine learning regression techniques including Multiple Linear Regression, Ridge Regression, and Lasso Regression to determine relative potency each alloying element had on computed solidification cracking index. A series of linear equations were produced that relate composition of C103 and TZM to solidification cracking index. The crack susceptibility of C103 for example is described by an equation of the form: cracking index ~ O + 0.667*C + 0.635*N + 0.00037*Ta – 0.0008*Hf (in wt.%) From that equation, it is clear that O has strong propensity to induce solidification cracking. Interestingly, Hf is shown to reduce calculated cracking response. Finally, realizing the potential of this method to discover new refractory alloy formulations across the period table that have low solidification cracking sensitivity, the code was applied to new untested alloy systems including W-Zr-C, W-Ta-C, and others. Conclusions. In summary, an open source numerical method has been developed using Python code to calculate Kou’s crack susceptibility index. The method was applied to refractory metals which are inherently difficult to study from a weldability testing standpoint since inert shielding gas is not sufficient and welding is typically done in vacuum, especially in light of findings presented here where oxygen has profound influence on solidification cracking. This work revealed the effect of compositional variations on a series of refractory metals and showed the framework defined here will be useful in 1) the development of new alloys that have improved weldability and 3D printability, 2) placing compositional limits on existing alloys, and 3) ensuring adequate controls of manufacturing processes such as 3D printing where powder reuse is critical. Keywords. pycalphad; Python; refractory metals; solidification cracking. References. B. Bocklund et. al. (2020) http://doi.org/10.5281/zenodo.3630657. S. Kou. (2015) https://doi.org/10.1016/j.actamat.2015.01.034. G.G. Lessmann and R.E. Gold. Welding Journal, issue 1, pp. 1-s – 8-s (1971). F.N. Michael and J.W. Sowards. NASA/TM-20230002218 (2023). R. Otis and Z.-K. Liu. (2017) http://doi.org/10.5334/jors.140. A. Van de Wallle et. al. (2018) https://doi.org/10.1016/j.calphad.2018.04.003.

pycalphad↗

Engine Icing Data - An Analytics Approach

Engine icing researchers at the NASA Glenn Research Center use the Escort data acquisition system in the Propulsion Systems Laboratory (PSL) to generate and collect a tremendous amount of data every day. Currently these researchers spend countless hours processing and formatting their data, selecting important variables, and plotting relationships between variables, all by hand, generally analyzing data in a spreadsheet-style program (such as Microsoft Excel). Though spreadsheet-style analysis is familiar and intuitive to many, processing data in spreadsheets is often unreproducible and small mistakes are easily overlooked. Spreadsheet-style analysis is also time inefficient. The same formatting, processing, and plotting procedure has to be repeated for every dataset, which leads to researchers performing the same tedious data munging process over and over instead of making discoveries within their data. This paper documents a data analysis tool written in Python hosted in a Jupyter notebook that vastly simplifies the analysis process. From the file path of any folder containing time series datasets, this tool batch loads every dataset in the folder, processes the datasets in parallel, and ingests them into a widget where users can search for and interactively plot subsets of columns in a number of ways with a click of a button, easily and intuitively comparing their data and discovering interesting dynamics. Furthermore, comparing variables across data sets and integrating video data (while extremely difficult with spreadsheet-style programs) is quite simplified in this tool. This tool has also gathered interest outside the engine icing branch, and will be used by researchers across NASA Glenn Research Center. This project exemplifies the enormous benefit of automating data processing, analysis, and visualization, and will help researchers move from raw data to insight in a much smaller time frame.

Engine Icing↗

OPeNDAP Clients, Aggregation and S3

In this talk, we will discuss our work for testing OPeNDAP client access of data stored in the Amazon S3 cloud storage using a set of common analysis tools including Panoply, Jupyter Notebooks with Python xarray, NCO command line tool package, ArcGIS, and GDAL. We will also discuss our ongoing work on improving performance in Hyrax aggregation functionality.

Amazon S3↗

TPSAS-NF1676L-32493-DND

The Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has supported the Open Data Cube (ODC) initiative to provide a data architecture solution that has value to its global users and increases the impact of EO satellite data. ODC is an open-source platform for processing satellite data. We have developed software products and tools around the core ODC that would help users perform machine learning on EO satellite data. The recent United Nations (UN) Sustainable Development Agenda provides a shared blueprint for peace and prosperity for people and for the planet, considering our current situation and helping to create a plan. The core of this agenda is a set of seventeen Sustainable Development Goals (SDGs), which represent an urgent call for action by all countries - both developed and developing - in a global partnership. The CEOS SEO team has recently developed and released a set of innovative Jupyter notebooks addressing UN SDGs 6.6.1 (spatial extents of water-related ecosystems), 11.3.1 (ratio of land consumption rate to population growth rate), and 15.3.1 (proportion of land that is degraded over total land area). These notebooks empower users by providing features that will assist with streamlining analysis ready data retrieval, processing, and visualization. We have recently incorporated several machine learning techniques in these notebooks. In this paper, we present the lessons learned from our experience on classifying land using supervised and unsupervised machine learning techniques using ODC framework for UN SDGs. We identify the current limitations of ODC to seamlessly support machine learning techniques. We propose features that would help machine learning, specifically within the ODC framework. We propose a thematic indexing/loading of data for both unsupervised learning as well as data annotation/labeling pipeline. Currently, ODC supports machine learning by separating data-management from the analysis process. It works as a mechanism to load cubes of data. ODC does not natively support features that are vital in machine learning such as validation splits, fair/balanced sampling, establishing load size constraints, etc. We believe that our proposed features will empower users by providing features that bring machine learning techniques closed to ODC. Enhancements to ODC to better accommodate machine learning techniques can assist in fulfilling UN SDGs such as 6.3.2, 6.4.2, 6.6.1, 11.3.1, 14.1.1, 15.1.1, 15.3.1, and 15.4.2.

Syed R Rizvi↗

DAAC Collaboration Overview for AGU

The Atmospheric Science Data Center (ASDC), Goddard Earth Sciences Data and Information Systems Center (GES DISC), Socioeconomic Data and Applications Center (SEDAC), Oak Ridge National Laboratory (ORNL), Land Processes DAAC, Alaska Satellite Facility (ASF), as well as NASA Global Imagery Browse Service (GIBS), Earth Science Data Systems (ESDS) Geographic Information Systems Team (EGIST), Earthdata Content Delivery Team, ArcGIS Online Governance Team, and the Systematic Data Transformation ACCESS team have come together to establish the ArcGIS DAAC Collaboration. This coalition of participants will demonstrate their use of the ArcGIS Enterprise to support Earth science research, applied science, and outreach using Earth Observing System Data and Information System (EOSDIS) data. This includes the use of web services to fuse data products across space and time through services (e.g. ArcGIS Image Services, Open Geospatial Consortium (OGC) Web Coverage Service (WCS), and OGC Web Mapping Services (WMS)) for analysis in Jupyter notebooks, desktop tools, and web based applications.

Matthew Steven Tisdale↗

VESIcal Part I: An open-source thermodynamic model engine for mixed volatile (H2O-CO2) solubility in silicate melts

Thermodynamics has been fundamental to the interpretation of geologic data and modeling of geologic systems for decades. However, more recent advancements in computational capabilities and a marked increase in researchers’ accessibility to computing tools has outpaced the functionality and extensibility of currently available modeling tools. Here we present VESIcal (Volatile Equilibria and Saturation Identification calculator): the first comprehensive modeling tool for H 2 O, CO 2 , and mixed (H 2 O-CO 2 ) solubility in silicate melts that: a) allows users access to seven commonly used models, plus easy inter-comparison between models; b) provides universal functionality for all models (e.g., functions for calculating saturation pressures, degassing paths, etc.); c) can process large datasets (1,000’s of samples) automatically; d) can output computed data into an excel spreadsheet for simple post-modeling analysis; e) integrates advanced plotting capabilities directly within the tool; and f) provides all of these within the framework of a python library, making the tool extensible by the user and allowing any of the model functions to be incorporated into any other code capable of calling python. The tool is presented within this manuscript, which is a Jupyter notebook containing worked examples accessible to python users with a range of skill levels. The basic functions of VESIcal can also be access via a web app (https://vesical.anvil.app). The VESIcal python library is open-source and available for download at https://github.com/kaylai/VESIcal.

K. Iacovino↗

GL4U: GeneLab for Colleges and Universities

GeneLab for Colleges and Universities (GL4U) will provide space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab (GL) team will host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – Training of Trainers), in which participants learn to analyze space-relevant omics data hosted on GL. The first bootcamp took place in early June 2021 with about 30 SJSU undergraduate students and covered space biology-specific lectures and hands-on instruction using Jupyter Notebooks (JNs) for RNA sequence (RNAseq) data analysis. All training materials including the enclosed files listed below will be made publicly available on GitHub. RNAseq Bootcamp Lectures (attached in combined file): Introduction to NASA, Space Biology, GeneLab, and the Command Line: NASA_GL_CL_Intro_FINAL.pdf - DRAFT from initial submission NASA_SB_GL_CL_Intro_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version RNAseq and Data Processing Overview: RNAseq_Overview_FINAL.pdf - DRAFT from initial submission RNAseq_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Overview of the Statistics Used for RNAseq Data Analysis: SJSU_Statistics_Intro_Lecture_FINAL.pdf - DRAFT from initial submission Statistics_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Completed JNs in HTML format (attached in combined file): Unix_Intro_JN_06-2021_completed.html R_Intro_JN_06-2021_completed.html RNAseq_fastq_to_counts_JN_06-2021_completed.html RNAseq_DGE_JN_06-2021_completed.html RNAseq Bootcamp Recordings (attached): GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_1_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_2_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_3_of_5.mp4 *There were issues with the part 4 recording so that is not available GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_5_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_1_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_2_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_3_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_4_of_4.mp4

GeneLab↗

GL4U: Training the next generation of bioinformaticians, one omics datatype at a time

Spaceflight modifies gene expression in every organism examined to date, including humans. Understanding how these gene expression changes affect physiology is crucial for the development of countermeasures to enable long-duration manned missions. NASA’s GeneLab project provides researchers open access to multi-omics data, including genetic and gene expression data, from spaceflight experiments that can be mined to understand the effects of spaceflight on biological systems. To ensure new knowledge generation through data re-use, it is important to maximize the number of scientists who utilize GeneLab data. Training students on the GeneLab platform is the best way to create long-term adopters of this NASA database and its tools. Turning students into future instructors and advocates will also accelerate the dissemination of these data and tools to the broader scientific community. Therefore, in collaboration with the GeneLab Educational Working Group (EWG), GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. During the bootcamp, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U direct training pilot program was conducted in June 2021 in collaboration with USRA and San Jose State University (SJSU). During the pilot, SJSU students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrates the capacity of GL4U for training young scientists and encouraging data re-use.

Jonathan Matthew Galazka↗

What (and How) MERRA-2 Reanalysis Data are Used in Applied Sciences

The Modern Era Retrospective-analysis for Research and Applications, Version 2 (MERRA-2) is the global atmospheric data reanalysis for the satellite era produced by NASA’s Global Modeling and Assimilation Office (GMAO), using the Goddard Earth Observing System Model (GEOS)version 5.12.4. The data are officially distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). MERRA-2 data have been widely used by the Earth sciences and application community. Since MERRA-2 data were released in early 2016, the number of registered data users has grown steadily from 1,252 in 2016 to 6477 in 2020. By the end of October 2021, ~16 petabytes (over 360 million files) of data have been distributed to more than 18,900 users. Searching in Google Scholar (https://scholar.google.com/), we have found over 7,000 articles, published between January 2017 and May 2021, involving the use ofMERRA-2 data. The figure shows the numbers for various application areas in which theMERRA-2 data have been used, covering almost all of the application areas defined in NASA Applied Sciences (http://appliedsciences.nasa.gov). The largest number of articles are found in disaster research, with the subcategories ordered in flood, wildfires, hurricanes and cyclones, and other forms of severe weather. In this presentation, we will discuss the preliminary findings from a review of the selected literature that uses MERRA-2 data in applied sciences. The current analytic and interoperable data services at GES DISC are listed, such as the on-the-fly subset and analysis service, NASA Giovanni; THREDDS Data Server(TDS); and Python Jupyter notebooks. In addition, we will introduce two new services for supporting the open sciences: My Dashboard and Related Publications.

data management↗