Search NASA⌕ Search

SEARCH · Search NASA

Results for “big data tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Engine Icing Data - An Analytics Approach

Engine icing researchers at the NASA Glenn Research Center use the Escort data acquisition system in the Propulsion Systems Laboratory (PSL) to generate and collect a tremendous amount of data every day. Currently these researchers spend countless hours processing and formatting their data, selecting important variables, and plotting relationships between variables, all by hand, generally analyzing data in a spreadsheet-style program (such as Microsoft Excel). Though spreadsheet-style analysis is familiar and intuitive to many, processing data in spreadsheets is often unreproducible and small mistakes are easily overlooked. Spreadsheet-style analysis is also time inefficient. The same formatting, processing, and plotting procedure has to be repeated for every dataset, which leads to researchers performing the same tedious data munging process over and over instead of making discoveries within their data. This paper documents a data analysis tool written in Python hosted in a Jupyter notebook that vastly simplifies the analysis process. From the file path of any folder containing time series datasets, this tool batch loads every dataset in the folder, processes the datasets in parallel, and ingests them into a widget where users can search for and interactively plot subsets of columns in a number of ways with a click of a button, easily and intuitively comparing their data and discovering interesting dynamics. Furthermore, comparing variables across data sets and integrating video data (while extremely difficult with spreadsheet-style programs) is quite simplified in this tool. This tool has also gathered interest outside the engine icing branch, and will be used by researchers across NASA Glenn Research Center. This project exemplifies the enormous benefit of automating data processing, analysis, and visualization, and will help researchers move from raw data to insight in a much smaller time frame.

Engine Icing↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗

Spinoff 2005

Topics covered include: Lighting the Way for Quicker, Safer Healing; Discovering New Drugs on the Cellular Level; Hydrogen Sensors Boost Hybrids; Today s Models Losing Gas?; 3-D Highway in the Sky; Popping a Hole in High-Speed Pursuits; Monitoring Wake Vortices for More Efficient Airports; From Rockets to Racecars; All-Terrain Intelligent Robot Braves Battlefront to Save Lives; Keeping the Air Clean and Safe--An Anthrax Smoke Detector; Lightning Often Strikes Twice; Technology That's Ready and Able to Inspect Those Cables; Secure Networks for First Responders and Special Forces; Space Suit Spins; Cooking Dinner at Home--From the Office; Nanoscale Materials Make for Large-Scale Applications; NASA s Growing Commitment: The Space Garden; Bringing Thunder and Lightning Indoors; Forty-Year-Old Foam Springs Back With New Benefits; Experiments With Small Animals Rarely Go This Well; NASA, the Fisherman's Friend; Crystal-Clear Communication a Sweet-Sounding Success; Inertial Motion-Tracking Technology for Virtual 3-D; Then Why Do They Call Earth the Blue Planet?; Valiant 'Zero-Valent' Effort Restores Contaminated Grounds; Harnessing the Power of the Sun; Water and Air Measures That Make 'PureSense'; Remote Sensing for Farmers and Flood Watching; Pesticide-Free Device a Fatal Attraction for Mosquitoes Making the Most of Waste Energy Washing Away the Worries About Germs Celestial Software Scratches More Than the Surface A Search Engine That's Aware of Your Needs Fault-Detection Tool Has Companies 'Mining' Own Business; Software to Manage the Unmanageable; Tracking Electromagnetic Energy With SQUIDs; Taking the Risk Out of Risk Assessment; Satellite and Ground System Solutions at Your Fingertips; Structural Analysis Made 'NESSUSary'; Software of Seismic Proportions Promotes Enjoyable Learning; Making a Reliable Actuator Faster and More Affordable; Cost-Cutting Powdered Lubricant NASA s Radio Frequency Bolt Monitor: A Lifetime of Spinoffs Going End to End to Deliver High-Speed Data; Advanced Joining Technology: Simple, Strong, and Secure; Big Results From a Smaller Gearbox; Low-Pressure Generator Makes Cleanrooms Cleaner; and The Space Laser Business Model.

Source record↗

New generation lidar systems for eye safe full time observations

The traditional lidar over the last thirty years has typically been a big pulse low repetition rate system. Pulse energies are in the 0.1 to 1.0 J range and repetition rates from 0.1 to 10 Hz. While such systems have proven to be good research tools, they have a number of limitations that prevent them from moving beyond lidar research to operational, application oriented instruments. These problems include a lack of eye safety, very low efficiency, poor reliability, lack of ruggedness and high development and operating costs. Recent advances in solid state laser, detectors and data systems have enabled the development of a new generation of lidar technology that meets the need for routine, application oriented instruments. In this paper the new approaches to operational lidar systems will be discussed. Micro pulse lidar (MPL) systems are currently in use, and their technology is highlighted. The basis and current development of continuous wave (CW) lidar and potential of other technical approaches is presented.

Spinhirne, James D.↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

EDX ClaiMM

EDX ClaiMM is a centralized data & analytical platform designed to revolutionize U.S. critical minerals and materials (CMM) activities. By providing a robust digital infrastructure, ClaiMM will accelerate the combination, leveraging, and rapid utilization of vital data, advanced tools, and cutting-edge research advancements in CMM. This adaptive digital research hub connects the CMM community to essential knowledge products and offers access to interoperable datasets, databases, models, software, and tools from the National Energy Technology’s (NETL’s) Energy Data eXchange (EDX) and other authoritative sources, serving both public and private sectors. EDX ClaiMM delivers AI-informed solutions to address fundamental knowledge gaps and fosters the innovation of new techniques for enhanced characterization and recovery of CMMs within the U.S. By leveraging cloud-hosted, scalable digital infrastructure, ClaiMM meets public–private applied energy needs. It equips the CMM community with priority digital resources that harness on-site and cloud compute capabilities, enabling big data storage, advanced processing, analytics, and visualization.

Critical Materials; Critical Minerals; Rare Earth ↗

Rovers as Geological Helpers for Planetary Surface Exploration

Rovers can be used to perform field science on other planetary surfaces and in hostile and dangerous environments on Earth. Rovers are mobility systems for carrying instrumentation to investigate targets of interest and can perform geologic exploration on a distant planet (e.g. Mars) autonomously with periodic command from Earth. For nearby sites (such as the Moon or sites on Earth) rovers can be teleoperated with excellent capabilities. In future human exploration, robotic rovers will assist human explorers as scouts, tool and instrument carriers, and a traverse "buddy". Rovers can be wheeled vehicles, like the Mars Pathfinder Sojourner, or can walk on legs, like the Dante vehicle that was deployed into a volcanic caldera on Mt. Spurr, Alaska. Wheeled rovers can generally traverse slopes as high as 35 degrees, can avoid hazards too big to roll over, and can carry a wide range of instrumentation. More challenging terrain and steeper slopes can be negotiated by walkers. Limitations on rover performance result primarily from the bandwidth and frequency with which data are transmitted, and the accuracy with which the rover can navigate to a new position. Based on communication strategies, power availability, and navigation approach planned or demonstrated for Mars missions to date, rovers on Mars will probably traverse only a few meters per day. Collecting samples, especially if it involves accurate instrument placement, will be a slow process. Using live teleoperation (such as operating a rover on the Moon from Earth) rovers have traversed more than 1 km in an 8 hour period while also performing science operations, and can be moved much faster when the goal is simply to make the distance. I will review the results of field experiments with planetary surface rovers, concentrating on their successful and problematic performance aspects. This paper will be accompanied by a working demonstration of a prototype planetary surface rover.

Stoker, Carol↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

Evaluating the Impact of Power Outages on Occupancy Patterns During the 2021 Texas Power Crisis

Large-scale power outages, such as those caused by extreme weather events, have a big impact on human behavior. A short power outage is merely a nuisance for most, and may not change people's locations. An outage that lasts for a few hours can result in spoiled food and medical supplies, and people will have to restock spoiled items. Long outages result in temperatures outside tolerable levels in homes, and may prompt people to acquire supplies, such as generators and gas, or change location. The long outages during Winter Storm Uri in Texas resulted in millions of dollars in property damage due to freezing pipes. This level of damage is expected to result in a sharp increase in supply runs and contractor activity. In this paper, we present a tool to explore differences in visiting patterns before, during, and after power outages. It allows to compare different points of interest like medical facilities, grocery stores, hardware stores, and other types of businesses.

big data↗

Combining Large Datasets - Cancer Moonshot Task Group Final Summary

In February 2022, President Biden re-ignited the Cancer Moonshot with bold new goals: to reduce the cancer death rate by half within 25 years and improve the lives of people with cancer and cancer survivors. To achieve these ambitious goals, the White House convened the first-ever Cancer Cabinet, bringing together departments and agencies from across the federal government to end cancer as we know it.The Cancer Cabinet convened three task forces and supporting task groups, including the Data and Innovation Task Force, which supported the Cancer Moonshot priority to “Deliver innovation to patients and communities.” In early 2023, the “Combining Large Datasets” (CoLD) Task Group was created within the Data and Innovation Task Force. The scope of the CoLD Task Group was how federal agencies combine large datasets for broad applications across cancer prevention and control, including nutrition, epidemiology, and military/Veteran health. Within this scope, the group sought to better leverage the immense potential of data and power of data tools to increase our understanding of cancer incidence, causes, mortality, treatments, prevention, outcomes, costs, and all other aspects of the burden of cancer.

data integration↗

Interactive Multi-Instrument Database of Solar Flares

The fundamental motivation of the project is that the scientific output of solar research can be greatly enhanced by better exploitation of the existing solar/heliosphere space-data products jointly with ground-based observations. Our primary focus is on developing a specific innovative methodology based on recent advances in "big data" intelligent databases applied to the growing amount of high-spatial and multi-wavelength resolution, high-cadence data from NASA's missions and supporting ground-based observatories. Our flare database is not simply a manually searchable time-based catalog of events or list of web links pointing to data. It is a preprocessed metadata repository enabling fast search and automatic identification of all recorded flares sharing a specifiable set of characteristics, features, and parameters. The result is a new and unique database of solar flares and data search and classification tools for the Heliophysics community, enabling multi-instrument/multi-wavelength investigations of flare physics and supporting further development of flare-prediction methodologies.

Heliophysics↗

Classification of Ion Mobility Data Using the Neural Network Approach

Determination of atmospheric and surface elemental and molecular composition of various solar system bodies is essential to the development of a firm understanding of the origin and evolution of the solar system. Furthermore, such data is needed to address the intriguing question of whether or not life exists or once existed elsewhere in the Solar System. As such, these measurements are among the primary scientific goals of NASA s current and future planetary missions. In recent years, significant progress toward both miniaturization and field portability of in situ analytical separation and detection devices have been made with future planetary explorations in mind. However, despite all these advances, accurate in situ identification of atmospheric and surface compounds remains a big challenge. In response to that we are developing various hardware and software tools which would enable us to uniquely identify species of interest in a complex chemical environment.

Duong, T. A.↗

Using Docker Containers to Extend Reproducibility Architecture for the NASA Earth Exchange (NEX)

NASA Earth Exchange (NEX) is a data, supercomputing and knowledge collaboratory that houses NASA satellite, climate and ancillary data where a focused community can come together to address large-scale challenges in Earth sciences. As NEX has been growing into a petabyte-size platform for analysis, experiments and data production, it has been increasingly important to enable users to easily retrace their steps, identify what datasets were produced by which process chains, and give them ability to readily reproduce their results. This can be a tedious and difficult task even for a small project, but is almost impossible on large processing pipelines. We have developed an initial reproducibility and knowledge capture solution for the NEX, however, if users want to move the code to another system, whether it is their home institution cluster, laptop or the cloud, they have to find, build and install all the required dependencies that would run their code. This can be a very tedious and tricky process and is a big impediment to moving code to data and reproducibility outside the original system. The NEX team has tried to assist users who wanted to move their code into OpenNEX on Amazon cloud by creating custom virtual machines with all the software and dependencies installed, but this, while solving some of the issues, creates a new bottleneck that requires the NEX team to be involved with any new request, updates to virtual machines and general maintenance support. In this presentation, we will describe a solution that integrates NEX and Docker to bridge the gap in code-to-data migration. The core of the solution is saemi-automatic conversion of science codes, tools and services that are already tracked and described in the NEX provenance system, to Docker - an open-source Linux container software. Docker is available on most computer platforms, easy to install and capable of seamlessly creating and/or executing any application packaged in the appropriate format. We believe this is an important step towards seamless process deployment in heterogeneous environments that will enhance community access to NASA data and tools in a scalable way, promote software reuse, and improve reproducibility of scientific results.

earth exchange↗

Distributed Machine Learning Workflow with PanDA and iDDS in LHC ATLAS

Machine Learning (ML) has become one of the important tools for High Energy Physics analysis. As the size of the dataset increases at the Large Hadron Collider (LHC), and at the same time the search spaces become bigger and bigger in order to exploit the physics potentials, more and more computing resources are required for processing these ML tasks. In addition, complex advanced ML workflows are developed in which one task may depend on the results of previous tasks. How to make use of vast distributed CPUs/GPUs in WLCG for these big complex ML tasks has become a popular research area. In this paper, we present our efforts enabling the execution of distributed ML workflows on the Production and Distributed Analysis (PanDA) system and intelligent Data Delivery Service (iDDS). First, we describe how PanDA and iDDS deal with large-scale ML workflows, including the implementation to process workloads on diverse and geographically distributed computing resources. Next, we report real-world use cases, such as HyperParameter Optimization, Monte Carlo Toy confidence limits calculation, and Active Learning. Finally, we conclude with future plans.

97 MATHEMATICS AND COMPUTING↗

SERVIR: Leveraging the Expertise of a Space Agency and a Development Agency to Increase Impact of Earth Observation in the Developing World

SERVIR is a joint initiative of the National Aeronautics and Space Administration (NASA) and the U.S. Agency for International Development (USAID), in collaboration with leading technical organizations around the world-- called SERVIR hubs--that serve and empower developing countries to use satellite data addressing critical challenges in food security and agriculture; water and water-related disasters; land cover, land use and ecosystems; and weather and climate. Over the past fourteen years, the program has worked with stakeholders in 50 countries across the world, partnered with 390 institutions, and generated and shared over 70 products from 27 satellites and sensors. In that process, around 7,400 specialists have been trained in the application of Earth observation data and technology. In its lifetime, SERVIR has been agile and innovative in shifting from what was essentially an incubator for testing and deploying Earth observation science and technology to making co-development the hallmark of its work, exemplified by both South-South and North-South scientific collaborations. SERVIR’s approach has embodied the concept that to solve really big problems, big, creative solutions are needed. SERVIR represents the world working together to address environmental challenges using spaced-based and geospatial technologies. Aligning with the very meaning of SERVIR, i.e. “to serve,” the program continues to be demand-driven in developing and deploying services (versus one-off products) which address development challenges using geospatial tools and Earth observation science. In 2016, as part of SERVIR’s evolution, USAID and NASA released the ‘SERVIR Service Planning Toolkit,’ a guidance document which provides a framework for how geospatial services can be used to tackle development challenges in a sustained manner. Since then, the Service Planning Toolkit’s systematic approach has begun to catch on in other Earth observation efforts. To improve access and use, SERVIR launched a Service Catalogue in February 2019, a searchable collection of demand-driven geospatial services that use Earth observations to support decision making. SERVIR implementing hub partners –include SERVIR-West Africa at the Agrometeorology, Hydrology and Meteorology (AGRHYMET) Regional Center, in Niamey, Niger; SERVIR-Eastern & Southern Africa at the Regional Centre for Mapping of Resources for Development in Nairobi, Kenya; SERVIR-Hindu Kush Himalaya at the International Centre for Integrated Mountain Development in Kathmandu, Nepal; SERVIR-Mekong at the Asian Disaster Preparedness Center in Bangkok, Thailand; and SERVIR’ -Amazonia, at the International Center for Tropical Agriculture (CIAT) in Cali, Colombia.

Searby, Nancy D.↗

Open Source Application of Fusing Aerosol Products from GEO and LEO Satellites

Retrieving aerosol optical depths (AODs) from sun-synchronous polar orbiting (aka low earth orbit, LEO) satellites, such as MODISs, and VIIRSs, OMI, TROPOMI, etc, has become well-established as a tool for extracting information on particulate matter (PM) and related processes in the atmosphere. However, with recently launched geostationary satellites (GEO), such as GOES-16/17/18, and Himawari-8/9, and Meteosat Third Generation (MTG) they provide a much higher temporal resolution (order of 10 minutes), typically an image once or more per hour during daylight compared to LEO once per day. By combining these observations, we may be able to characterize the diurnal cycle of global AOD at the local, regional and global scale. While the science community is still exploring the new data from GEO observations, we have been thinking about how to properly combine/merge/fuse those data considering differences in their spatial and temporal resolutions. However, this poses a “Big Data” challenge. The big data challenge is not just about data storage, but also about data discoverability, and accessibility, and even more, about data migration/mirroring in the cloud-computing environment. This paper is merely showing some of the efforts and approaches we have attempted in fusing six satellites’ Level 2 aerosol data (three are from GEO (GOES-16/17 and Himawari-8), and the other three are from LEO (TERRA/MODIS, AQUA/MODIS, SNPP-VIIRS) from Dark Target (DT) aerosol retrieval algorithm. Having the on-demand capability of fusing remote sensing products onto the desired temporal and spatial domain enables researchers and application practitioners to better manipulate and work with satellite and sensor data. It is our hopeWe hope that by making such an open-source package, and the accompanying functionality, the scientific community will be granted easier access to aerosol data processing resources. The MEaSUREs Program (Making Earth System Data Records for Use in Research Environments) expands our understanding of the Earth's current system through atmospheric and surface measurements. In an effort to aid the scientific research component and improve open source methods, this project developed Python code for fusing six satellite Level 2 aerosol data (three are from geostationary satellites (GEO), and the other three are from low earth orbital satellites (LEO)) from Dark Target Aerosol Retrieval Algorithm.

Jennifer Wei↗

Analyzing a 35-Year Hourly Data Record: Why So Difficult?

At the Goddard Distributed Active Archive Center, we have recently added a 35-Year record of output data from the North American Land Assimilation System (NLDAS) to the Giovanni web-based analysis and visualization tool. Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) offers a variety of data summarization and visualization to users that operate at the data center, obviating the need for users to download and read the data themselves for exploratory data analysis. However, the NLDAS data has proven surprisingly resistant to application of the summarization algorithms. Algorithms that were perfectly happy analyzing 15 years of daily satellite data encountered limitations both at the algorithm and system level for 35 years of hourly data. Failures arose, sometimes unexpectedly, from command line overflows, memory overflows, internal buffer overflows, and time-outs, among others. These serve as an early warning sign for the problems likely to be encountered by the general user community as they try to scale up to Big Data analytics. Indeed, it is likely that more users will seek to perform remote web-based analysis precisely to avoid the issues, or the need to reprogram around them. We will discuss approaches to mitigating the limitations and the implications for data systems serving the user communities that try to scale up their current techniques to analyze Big Data.

computational performance↗

End-to-End Mission Design & Trajectory Optimization

Need: A need exists for a generalized, robust, user-friendly and accessible end-to-end mission design optimization tool. Solution: Our solution to developing this capability was to interface two JSC tools—Copernicus and Genesis. Each of these tools has a specific area of the mission design process that it excels at. By utilizing them both, we can gain performance benefits not seen by either on their own. - Copernicus is a trajectory design and optimization software used for in-space trajectories around multiple bodies. - Genesis is a flight mechanics tool used to model ascent, entry, descent, and landing trajectories around a single planetary body. Year 1 was focused on combining these 2 software packages—allowing Copernicus to incorporate the ascent/descent capabilities of Genesis into the optimization problem—and developing this end-to-end mission design capability. Year 2 we focused on increasing the robustness of this capability by building the initial guess generator (IGG), which produces initial guesses based on simplifying assumptions and the physics of the problem. Year 3 of our project focused on utilizing the end-to-end mission design and optimization capabilities developed in the previous 2 years to analyze specific mission scenarios—scaling up from proof of concept to real analyses—capturing any resulting performance benefits, as well as addressing the Big Data challenges we’re faced with—namely, how we’re going to manage and interpret all the data that’s generated.

Kristin Nichols↗