Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Open-Source Science-led Development of the Atmosphere Observing System (AOS) Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David Giles↗

Open-Source Science-led Development of the AOS Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David M. Giles↗

Stewardship Best Practices for Improved Discovery and Reuse of Heterogeneous and Cross-Disciplinary Earth System Data

Some of the Earth system data products such as those from NASA airborne and field investigations (a.k.a. campaigns), are highly heterogeneous and cross-disciplinary, making the data extremely challenging to manage. For example, airborne and field campaign measurements tend to be sporadic over a period of time, with large gaps. Data products generated are of various processing levels and utilized for a wide range of inter- and cross-disciplinary research and applications. Data and derived products have been historically stored in a variety of domain-specific standard (and some non-standard) formats and in various locations such as NASA Distributed Active Archive Centers (DAACs), NASA airborne science facilities, field archives, or even individual scientists’ computer hard drives. As a result, airborne and field campaign data products have often been managed and represented differently, making it onerous for data users to find, access, and utilize campaign data. Some difficulties in discovering and accessing the campaign data originate from the incomplete data product and contextual metadata that may contain details relevant to the campaign (e.g. campaign acronym and instrument deployment locations), but tend to lack other significant information needed to understand conditions surrounding the data. Such details can be burdensome to locate after the conclusion of a campaign. Utilizing consistent terminology, essential for improved discovery and reuse, is also challenging due to the variety of involved disciplines. To help address the aforementioned challenges faced by many repositories and data managers handling airborne and field data, this presentation will describe stewardship practices developed by the Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) under the NASA’s Earth Science Data systems (ESDS) Program.

best practices↗

Representation of Serendipitous Scientific Data

A computer program defines and implements an innovative kind of data structure than can be used for representing information derived from serendipitous discoveries made via collection of scientific data on long exploratory spacecraft missions. Data structures capable of collecting any kind of data can easily be implemented in advance, but the task of designing a fixed and efficient data structure suitable for processing raw data into useful information and taking advantage of serendipitous scientific discovery is becoming increasingly difficult as missions go deeper into space. The present software eases the task by enabling definition of arbitrarily complex data structures that can adapt at run time as raw data are transformed into other types of information. This software runs on a variety of computers, and can be distributed in either source code or binary code form. It must be run in conjunction with any one of a number of Lisp compilers that are available commercially or as shareware. It has no specific memory requirements and depends upon the other software with which it is used. This program is implemented as a library that is called by, and becomes folded into, the other software with which it is used.

James, Mark↗

Reduction and analysis of photometric data on Comet Halley

The discovery that periodic variations in the brightness of Comet Halley were characterized by two unrelated frequencies implies that the nucleus is in a complex state of rotation. It either nutates as a result of the random addition of small torque perturbations accumulated over many perihelion passages, or the jet activity torques are so strong that it precesses wildly at each perihelion passage. To diagnose the state of nuclear rotation, researchers began a program to acquire photometric time series of the comet as it recedes from the sun. The intention is to observe the decay of the comet's atmosphere and then, when it is unemcumbered by the light of the coma, follow the light variation of the nucleus itself. The latter will be compared with preperihelion time series and the orientation of the nucleus at the time of Vega and Giotto flybys and an accurate rotational ephemeris constructed. Halley was observed on 38 nights during 1987 and approximately 21 nights in 1988. The comet moved from 5 AU to 8.5 AU during this time. The brightness of the coma was found to rapidly decrease in 1988 as the coma and cometary activity collapses. The magnitude in April 1988 was 19 mag (visual) and it is predicted that the nucleus itself will be the major contributor to the brightness in the 1988 and 1989 season.

Belton, Michael J. S.↗

NASA Environmental Justice Data Search Interface Overview

NASA’s Earth Science Division (ESD) is committed to empower Environmental Justice (EJ) communities by expanding awareness, accessibility, and use of Earth science data to enable contributions to Earth science research and applications. To that end, the NASA Earth Science Data Systems (ESDS) Program developed an EJ Data Catalog, a simple guide to NASA datasets and socioeconomic datasets that may be useful in EJ research. The EJ Data Catalog is divided by topics—such as disasters, urban flooding, extreme heat, food availability, water availability, climate, and health and air quality—and possible use cases for each dataset. The new version of the EJ Data Catalog is now integrated into NASA’s Science Discovery Engine (SDE), an open-source science infrastructure to enable collaborative and interdisciplinary science. In this workshop you will learn about NASA’s Equity and Environmental Justice (EEJ) activities and opportunities as well as participate on an interactive live demo of the new Science Discovery Engine for Environmental Justice.

environmental justice↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Material Units, Structures/Landforms, and Stratigraphy for the Global Geologic Map of Ganymede (1:15M)

In the coming year a global geological map of Ganymede will be completed that represents the most recent understanding of the satellite on the basis of Galileo mission results. This contribution builds on important previous accomplishments in the study of Ganymede utilizing Voyager data and incorporates the many new discoveries that were brought about by examination of Galileo data. Material units have been defined, structural landforms have been identified, and an approximate stratigraphy has been determined utilizing a global mosaic of the surface with a nominal resolution of 1 km/pixel assembled by the USGS. This mosaic incorporates the best available Voyager and Galileo regional coverage and high resolution imagery (100-200 m/pixel) of characteristic features and terrain types obtained by the Galileo spacecraft. This map has given us a more complete understanding of: 1) the major geological processes operating on Ganymede, 2) the characteristics of the geological units making up its surface, 3) the stratigraphic relationships of geological units and structures, and 4) the geological history inferred from these relationships. A summary of these efforts is provided here.

Patterson, G. Wesley↗

Data Mining and Complex Problems: Case Study in Composite Materials

Data mining is defined as the discovery of useful, possibly unexpected, patterns and relationships in data using statistical and non-statistical techniques in order to develop schemes for decision and policy making. Data mining can be used to discover the sources and causes of problems in complex systems. In addition, data mining can support simulation strategies by finding the different constants and parameters to be used in the development of simulation models. This paper introduces a framework for data mining and its application to complex problems. To further explain some of the concepts outlined in this paper, the potential application to the NASA Shuttle Reinforced Carbon-Carbon structures and genetic programming is used as an illustration.

Rabelo, Luis↗

The Little Photometer That Could: Technical Challenges and Science Results from the Kepler Mission

The Kepler spacecraft launched on March 7, 2009, initiating NASA's first search for Earth-size planets orbiting Sun-like stars. Since launch, Kepler has announced the discovery of 17 exoplanets, including a system of six transiting a Sun-like star, Kepler-11, and the first confirmed rocky planet, Kepler-10b, with a radius of 1.4 that of Earth. Kepler is proving to be a cornucopia of discoveries: it has identified over 1200 candidate planets based on the first 120 days of observations, including 54 that are in or near the habitable zone of their stars, and 68 that are 1.2 Earth radii or smaller. An astounding 408 of these planetary candidates are found in 170 multiple systems, demonstrating the compactness and flatness of planetary systems composed of small planets. Never before has there been a photometer capable of reaching a precision near 20 ppm in 6.5 hours and capable of conducting nearly continuous and uninterrupted observations for months to years. In addition to exoplanets, Kepler is providing a wealth of astrophysics, and is revolutionizing the field of asteroseismology. Designing and building the Kepler photometer and the software systems that process and analyze the resulting data to make the discoveries presented a daunting set of challenges, including how to manage the large data volume. The challenges continue into flight operations, as the photometer is sensitive to its thermal environment, complicating the task of detecting 84 ppm drops in brightness corresponding to Earth-size planets transiting Sun-like stars.

Kepler Mission↗

The use of radar and LANDSAT data for mineral and petroleum exploration in the Los Andes region, Venezuela

A geological study of a 27,500 sq km area in the Los Andes region of northwestern Venezuela was performed which employed both X-band radar mosaics and computer processed Landsat images. The 3.12 cm wavelength radar data were collected with horizontal-horizontal polarization and 10 meter spatial resolution by an Aeroservices SAR system at an altitude of 12,000 meters. The radar images increased the number of observable suspected fractures by 27 percent over what could be mapped by LANDSAT alone, owing mostly to the cloud cover penetration capabilities of radar. The approximate eight fold greater spatial resolution of the radar images made possible the identification of shorter, narrower fractures than could be detected with LANDSAT data alone, resulting in the discovery of a low relief anticline that could not be observed in LANDSAT data. Exploration targets for petroleum, copper, and uranium were identified for further geophysical work.

Vincent, R. K.↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique signature of visual behavior. This is true even in a restricted domain, such as piloting an aircraft, where patterns of visual signatures might represent activities like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these activities based on their behavioral signatures. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as τ. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration τ, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to counts by multiplying by τ divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We have investigated the method using a set of 10 synthetic "activities," that sample 4 AOIs. Four of these activities visit 3 of the 4 AOIs, with equal probability; as there are four different ways to leave-one- out, there are four such activities. Similarly, there are six different activities that leave-two-out. Sequences of simulated behavior were generated by running each activity for 40 seconds, in sequence, for a total of 6.7 minutes. The figure to the right shows the matrix of chi-square statistics, using a value of 2.8 seconds for τ, corresponding to 10 fixations. Low values (dark) indicate poor evidence for activity differences, while high values (bright) indicate strong evidence. The dark squares along the main diagonal each correspond to the forty second intervals in which the activity was held constant; the 4x4 block at the lower left corresponds to the four leave-one-out activities, while the 6x6 block in the upper right corresponds to the leave-two-out activities. (The anti-diagonal pattern of white squares indicates those activity pairs that share no AOIs.) The chi-square values can be binarized by choosing a particular significance level; we are interested in grouping bins that represent the same activity, effectively accepting the null hypothesis. Therefore, we may adopt a relatively lax criterion; for example, choosing a p-value of 0.2 means that two behaviors that have only a 1-in-5 chance of being produced by a single activity might nevertheless be clustered together. We have explored several methods to perform clustering on the data and solving for the activity probabilities. Greedy methods begin by selecting the time bin that is similar to the most (or least) other bins, and then forming a cluster from it and all other non-discriminable bins. These methods show mediocre performance, as they do not take into account temporal contiguity. Preliminary results indicate that methods that "grow" clusters in time from seed points perform better.

activity analysis↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a piloting an aircraft, where activities with distinct visual signatures might be things like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as delta. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration delta, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to integer counts by multiplying by delta divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. The method has been applied to approximately 100 hours of eye movement data collected from pilots in a high-fidelity B747 flight simulator, and the results have been compared to synthetic data in which the each activity is represented as first-order Markov process with random probabilities assigned to the AOIs. Randomly-generated synthetic activities can require thousands of fixations to be discriminated with statistical significance, while the human data can be clustered using averaging windows of some 10's of seconds, suggesting that the actual activities are much more narrowly focused than random Markov models.

activity analysis↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a pilot flying an airplane; in this case, activities with distinct visual signatures might be things like communicating, navigating, monitoring, etc. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. We compare, not individual fixations, but groups of fixations aggregated over a fixed time interval (Tau). We assume that the environment has been divided into a finite set of discrete areas-of-interest (AOIs). For a given time interval, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector, where N is the number of AOIs. These proportions can be converted to integer counts by multiplying by Tau divided by the average fixation duration, a parameter that we fix at 283 milliseconds. We compare different intervals by computing the chi-squared statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We cluster the intervals, first by merging adjacent intervals that are sufficiently similar, optionally shifting the boundary between non-merged intervals to maximize the difference. Then we compare and cluster non-adjacent intervals. The method is evaluated using synthetic data generated by a hand-crafted set of activities. While the method generally finds more activities than put into the simulation, we have obtained agreement as high as 80 percent between the inferred activity labels and ground truth.

Eye Movements↗

Enabling Future Low-Cost Small Mission Concepts

A SmallSat using a small Radioisotope Power System for deep space destinations could potentially fit into a Discovery class mission cost cap and perform significant science with a timely return of data. Only applicable when the Discovery 12 guidelines were applied. Commonality of hardware and science instruments among identical spacecraft enabled to meet the Discovery Class mission cost cap. Multiple spacecraft shared the costs of the Launch Approval Engineering Process. Assumed a secondary science instrument was contributed. Small RPS could provide small spacecraft with a relatively high power (approx. 60 We) option for missions to deep space destinations (> 10 AU) with multiple science instruments. Study of Centaur mission demonstrated the ability to achieve New Frontiers level science. Multiple spacecraft possible with small RPS, allowing for multiple targets, science from multiple platforms, and/or redundancy.

Lee, Young↗

NASA's EOSDIS Near Term Challenges

Given the long-term requirements, and the rapid pace of information technology and changing expectations of the user community, the ESDIS Project has had to evolve EOSDIS continually over the past three decades. However, many challenges remain. One near-term challenge is the enormous quantity of new data that will need to be managed by the EOSDIS. With the upcoming launch of the latest NASA missions coupled with existing operational missions and field campaigns, EOSDIS can expect to handle as much as 50 petabytes of data per year. In perspective, this is twice the size of the current existing archive, which took over 21 years to collect. Another continuing challenge is the disparate requirements of a diverse science community. Maintaining rigorous long-term data preservation, supporting ease of discovery and access, incorporating user feedback, enabling reanalysis/ reprocessing, and agile integration of new data sources, continue to be the Project's objectives.

Remote Sensing↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗