Search NASASearch

SEARCH · Search NASA

Results for “data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Online and Offline Data Quality Monitoring for the Mu2e Calorimeter

This thesis presents the design, implementation, and validation of a calorimeter Data Quality Monitoring (DQM) toolchain for the Mu2e experiment at Fermilab. Mu2e searches for charged lepton flavor violation via coherent muon-to-electron conversion in the field of an aluminum nucleus, $\mu^- Al \rightarrow e^-Al$, a process whose observation would constitute clear evidence of physics beyond the Standard Model. Achieving target sensitivity requires stringent control of detector performance and data integrity during acquisition, as subtle issues in readout configuration, data formatting, or electronics behavior can compromise reconstruction and bias downstream analyzes. To address these challenges, this work develops a multi-layer DQM approach spanning both raw data validation and reconstructed digi-level diagnostics. At the low level, a fragment analysis component performs word- and bit-field decoding of calorimeter readout blocks, enabling sanity checks of the expected structure and producing detailed error and integrity statistics useful for commissioning and troubleshooting. At the digi level, the CaloDigiDQM analyzer is implemented within the art framework and transforms each CaloDigiCollection into a structured hierarchy of ROOT histograms designed for fast drill-down diagnostics. The module generates coherent monitoring views at global, disk, board, and channel granularity, including occupancy, waveform-derived features (baseline, RMS, peak amplitude and position), and left-right sensor consistency metrics. Detector-aware channel-to-electronics mapping is performed through the conditions system (CaloDAQMap), ensuring that diagnostics remain aligned with hardware identifiers used in operations. For end-to-end testing without reliance on live DAQ data, a synthetic CaloDigi producer is developed to generate realistic waveforms with controlled noise and pulse shapes. The resulting system supports both offline ROOT-file production and online operation, including optional histogram streaming through otsdaq via ots::HistoSender. This toolchain provides a practical and scalable foundation for calorimeter commissioning and stable data collection, enabling early detection of anomalies and reducing operational risk for Mu2e.

Vakulenko, Mark [Drew U.] (ORCID:0009000276197818)

Sensor Management for Applied Research Technologies (SMART)-On Demand Modeling (ODM) Project

NASA requires timely on-demand data and analysis capabilities to enable practical benefits of Earth science observations. However, a significant challenge exists in accessing and integrating data from multiple sensors or platforms to address Earth science problems because of the large data volumes, varying sensor scan characteristics, unique orbital coverage, and the steep learning curve associated with each sensor and data type. The development of sensor web capabilities to autonomously process these data streams (whether real-time or archived) provides an opportunity to overcome these obstacles and facilitate the integration and synthesis of Earth science data and weather model output. A three year project, entitled Sensor Management for Applied Research Technologies (SMART) - On Demand Modeling (ODM), will develop and demonstrate the readiness of Open Geospatial Consortium (OGC) Sensor Web Enablement (SWE) capabilities that integrate both Earth observations and forecast model output into new data acquisition and assimilation strategies. The advancement of SWE-enabled systems (i.e., use of SensorML, sensor planning services - SPS, sensor observation services - SOS, sensor alert services - SAS and common observation model protocols) will have practical and efficient uses in the Earth science community for enhanced data set generation, real-time data assimilation with operational applications, and for autonomous sensor tasking for unique data collection.

Goodman, M.

Automated pipeline processing X-ray diffraction data from dynamic compression experiments on the Extreme Conditions Beamline of PETRA III

Presented and discussed here is the implementation of a software solution that provides prompt X-ray diffraction data analysis during fast dynamic compression experiments conducted within the dynamic diamond anvil cell technique. It includes efficient data collection, streaming of data and metadata to a high-performance cluster (HPC), fast azimuthal data integration on the cluster, and tools for controlling the data processing steps and visualizing the data using the DIOPTAS software package. This data processing pipeline is invaluable for a great number of studies. The potential of the pipeline is illustrated with two examples of data collected on ammonia–water mixtures and multiphase mineral assemblies under high pressure. The pipeline is designed to be generic in nature and could be readily adapted to provide rapid feedback for many other X-ray diffraction techniques, e.g. large-volume press studies, in situ stress/strain studies, phase transformation studies, chemical reactions studied with high-resolution diffraction etc.

97 MATHEMATICS AND COMPUTING

Method for Accessing Distributed Heterogeneous Databases

A scenario of relational, hierarchial, and network data bases is presented and a distributed access view integrated data base system (DAVID) is described for uniformly accessing data bases which are heterogeneous and physically distributed. The DAVID system is based on data base logic so that the relational approach is generalized to the heterogeneous approach. The global data manager is explained as are global data manipulation languages which can operate on all the data bases and can query the data dictionary and the data directory.

Jacobs, B. E.

Laser fringe anemometry for aero engine components

Advances in flow measurement techniques in turbomachinery continue to be paced by the need to obtain detailed data for use in validating numerical predictions of the flowfield and for use in the development of empirical models for those flow features which cannot be readily modelled numerically. The use of laser anemometry in turbomachinery research has grown over the last 14 years in response to these needs. Based on past applications and current developments, this paper reviews the key issues which are involved when considering the application of laser anemometry to the measurement of turbomachinery flowfields. Aspects of laser fringe anemometer optical design which are applicable to turbomachinery research are briefly reviewed. Application problems which are common to both laser fringe anemometry (LFA) and laser transit anemometry (LTA) such as seed particle injection, optical access to the flowfield, and measurement of rotor rotational position are covered. The efficiency of various data acquisition schemes is analyzed and issues related to data integrity and error estimation are addressed. Real-time data analysis techniques aimed at capturing flow physics in real time are discussed. Finally, data reduction and analysis techniques are discussed and illustrated using examples taken from several LFA turbomachinery applications.

Strazisar, A. J.

Laser fringe anemometry for aero engine components

Advances in flow measurement techniques in turbomachinery continue to be paced by the need to obtain detailed data for use in validating numerical predictions of the flowfield and for use in the development of empirical models for those flow features which cannot be readily modelled numerically. The use of laser anemometry in turbomachinery research has grown over the last 14 years in response to these needs. Based on past applications and current developments, the key issues which are involved when considering the application of laser anemometry to the measurement of turbomachinery flowfields are discussed. Aspects of laser fringe anemometer optical design which are applicable to turbomachinery research are briefly reviewed. Application problems which are common to both laser fringe anemometry (LFA) and laser transit anemometry (LTA) such as seed particle injection, optical access to the flowfield, and measurement of rotor rotational position are covered. The efficiency of various data acquisition schemes is analyzed and issues related to data integrity and error estimation are addressed. Real-time data analysis techniques aimed at capturing flow physics in real time are discussed. Finally, data reduction and analysis techniques are discussed and illustrated using examples taken from several LFA turbomachinery applications.

Strazisar, Anthony J.

Data Access Services that Make Remote Sensing Data Easier to Use

This slide presentation reviews some of the processes that NASA uses to make the remote sensing data easy to use over the World Wide Web. This work involves much research into data formats, geolocation structures and quality indicators, often to be followed by coding a preprocessing program. Only then are the data usable within the analysis tool of choice. The Goddard Earth Sciences Data and Information Services Center is deploying a variety of data access services that are designed to dramatically shorten the time consumed in the data preparation step. On-the-fly conversion to the standard network Common Data Form (netCDF) format with Climate-Forecast (CF) conventions imposes a standard coordinate system framework that makes data instantly readable through several tools, such as the Integrated Data Viewer, Gridded Analysis and Display System, Panoply and Ferret. A similar benefit is achieved by serving data through the Open Source Project for a Network Data Access Protocol (OPeNDAP), which also provides subsetting. The Data Quality Screening Service goes a step further in filtering out data points based on quality control flags, based on science team recommendations or user-specified criteria. Further still is the Giovanni online analysis system which goes beyond handling formatting and quality to provide visualization and basic statistics of the data. This general approach of automating the preparation steps has the important added benefit of enabling use of the data by non-human users (i.e., computer programs), which often make sub-optimal use of the available data due to the need to hard-code data preparation on the client side.

Lynnes, Christopher

CrossMP: Enabling Cross-Modality Translation between Single-Cell RNA-Seq and Single-Cell ATAC-Seq through Web-Based Portal

In recent years, there has been a growing interest in profiling multiomic modalities within individual cells simultaneously. One such example is integrating combined single-cell RNA sequencing (scRNA-seq) data and single-cell transposase-accessible chromatin sequencing (scATAC-seq) data. Integrated analysis of diverse modalities has helped researchers make more accurate predictions and gain a more comprehensive understanding than with single-modality analysis. However, generating such multimodal data is technically challenging and expensive, leading to limited availability of single-cell co-assay data. Here, we propose a model for cross-modal prediction between the transcriptome and chromatin profiles in single cells. Our model is based on a deep neural network architecture that learns the latent representations from the source modality and then predicts the target modality. It demonstrates reliable performance in accurately translating between these modalities across multiple paired human scATAC-seq and scRNA-seq datasets. Additionally, we developed CrossMP, a web-based portal allowing researchers to upload their single-cell modality data through an interactive web interface and predict the other type of modality data, using high-performance computing resources plugged at the backend.

59 BASIC BIOLOGICAL SCIENCES

Local effects of partly-cloudy skies on solar and emitted radiation

A computer automated data acquisition system for atmospheric emittance, and global solar, downwelled diffuse solar, and direct solar irradiances is discussed. Hourly-integrated global solar and atmospheric emitted radiances were measured continuously from February 1981 and hourly-integrated diffuse solar and direct solar irradiances were measured continuously from October 1981. One-minute integrated data are available for each of these components from February 1982. The results of the correlation of global insolation with fractional cloud cover for the first year's data set. A February data set, composed of one-minute integrated global insolation and direct solar irradiance, cloud cover fractions, meteorological data from nearby weather stations, and GOES East satellite radiometric data, was collected to test the theoretical model of satellite radiometric data correlation and develop the cloud dependence for the local measurement site.

Whitney, D. A.

Artificial Intelligence Medical Support for Long-Duration Space Missions

We envision an artificial intelligence (AI) based system that will provide support and recommendations to the crew medical officer (CMO) and ground flight surgeon during long-duration space missions. Such a system would be pretrained on the knowledgebase of clinical knowledge on Earth, minimizing the amount of Earth data that needs to be transferred into space. Then during deployment, the system would be constantly refined through active learning from diverse streams of data from sensors in the spacecraft, data collected daily from individual astronauts, and human-in-the-loop feedback from the crew. The model could be interrogated for predictions and recommendations on personalized crew health based on the overall status of the spacecraft, medicinal stores, and status of other crew members. Adaptation techniques would be used to incorporate spaceflight data that have very different distributions from the training data due to the extreme environment. Edge computing and the most advanced neuromorphic processing would enable computation in scenarios with low power and bandwidth, while dimensionality reduction would be employed to ensure that the input data streams from spaceflight are as small as possible. In order to realize this long-term vision, several hardware and software aspects need to be developed and assembled. First, models pretrained on Earth biomedical data would need to be evaluated for predictive accuracy, and the best one selected. That model would need to be adapted to learn from diverse, sparse, and inconsistently measured data streams, as well as human-in-the-loop feedback. A data integration, standardization, and dimensionality reduction methodology would need to be developed to handle all data types and feed them into the model. Once the software and data infrastructure is developed, it would need to be integrated with small footprint compute processors and tested in high-radiation, high-vibration, unregulated temperature situations. As a short-term goal, we recommend to focus on the development of the data and model software structure. Several large language models (LLM) already exist that have been trained on Earth biomedical and clinical knowledgebases, including BioMedLLM, Med-PaLM, SPOKE LLM, and Foresight. These models need to be evaluated for accuracy and the best one chosen for a proof-of-concept structure, while maintaining awareness of the accelerating AI field and incorporating any newly improved model architectures as needed. Then, we recommend to develop a database of synthetic data types to mimic the diverse data streams that are expected in a long-duration space mission. This should include environmental and microbial data from the spacecraft, non-invasive data from wearables and point-of-care devices employed by astronauts, and more invasive molecular and physiological monitoring of clinical and biomarker data from astronauts. The data standardization methodology should be developed, and these data streams used to refine the clinical LLM. Several scenarios should be developed that could plausibly come up in a long-duration space mission, and changes or aberrations introduced to the data at specific times to mimic these scenarios. Then, question and answer tasks should be designed to interrogate the model for predictions and recommendations, with acceptable answers already identified.

Artificial Intelligence

Apollo Lunar Sample Integration into Google Moon: A New Approach to Digitization

The Google Moon Apollo Lunar Sample Data Integration project is part of a larger, LASER-funded 4-year lunar rock photo restoration project by NASA s Acquisition and Curation Office [1]. The objective of this project is to enhance the Apollo mission data already available on Google Moon with information about the lunar samples collected during the Apollo missions. To this end, we have combined rock sample data from various sources, including Curation databases, mission documentation and lunar sample catalogs, with newly available digital photography of rock samples to create a user-friendly, interactive tool for learning about the Apollo Moon samples

Dawson, Melissa D.

Visualizing Geospatial Data through ESRI Story Maps for Earth Science Education: Lessons Learned from My NASA Data

For 20 years My NASA Data (MND) has curated NASA Earth science data and provided the data to educators in engaging learner-centered resources. MND has recently featured story maps as an innovative way to engage students in NASA Earth data. A story map is a cloud-based lesson that engages the learner in interactive geospatial maps using NASA data, and other multimedia content, text, and tasks that can be seamlessly incorporated in classroom instruction. This immersive technology eliminates the need for the user to move among tricky interfaces to access and visualize Earth science data, and no special software is required to be downloaded. Each story map integrates data from different NASA satellite missions, retrieved from Distributed Active Archive Centers (DAACs). Story maps also employ data analysis tools, such as time series options and swipe tools that allow learners to view and analyze relationships between scientific variables. MND has produced 25 story map lesson plans on the topics of air quality, the urban heat island effect, Earth’s energy budget, phytoplankton distribution, hurricane formation, solar eclipses, ocean circulation patterns, sea ice extent, and volcanic eruptions. Nine of them are extended story maps and written in the 5E format, which is internationally recognized as best practice based on how children learn science. Each story map resource is developed by the MND team featuring a GIS programming specialist, a lead scientist, and educational specialist/s to ensure the context, content, and methods are scientifically and educationally sound. The MND story maps are written for middle and high school science teachers and students as they connect with the Earth Systems Science phenomena featured in the Next Generation Science Standards. Each story map includes supporting resources for smooth integration in the classroom. During Fiscal Year 2023, The My NASA Data website received over 100,000 story map engagements during. These metrics highlight the interest in story maps as an Earth Science educational resource.

Desiray Wilson

A One-Stop Web Application for the Mars Team

THe Mars Exploration Rover Collaborative Information Portal (MERCIP) provides a window to all mission events. It supports mission updates, data sharing, and collaboration. The report also discusses: Technology spotlight. Integrating data multiple sources. Securing access for multiple clients. A new information infrastructure.

Laufenberg, Lawrence

Bundle Data Approach at GES DISC Targeting Natural Hazards

Severe natural phenomena such as hurricane, volcano, blizzard, flood and drought have the potential to cause immeasurable property damages, great socioeconomic impact, and tragic loss of human life. From searching to assessing the Big, i.e., massive and heterogeneous scientific data (particularly, satellite and model products) in order to investigate those natural hazards, it has, however, become a daunting task for Earth scientists and applications researchers, especially during recent decades. The NASA Goddard Earth Sciences Data and Information Service Center (GES DISC) has served Big Earth science data, and the pertinent valuable information and services to the aforementioned users of diverse communities for years. In order to help and guide our users to online readily (i.e., with a minimum effort) acquire their requested data from our enormous resource at GES DISC for studying their targeted hazard event, we have thus initiated a Bundle Data approach in 2014, first targeting the hurricane event topic. We have recently worked on new topics such as volcano and blizzard. The bundle data of a specific hazard event is basically a sophisticated integrated data package consisting of a series of proper datasets containing a group of relevant (knowledge--based) data variables readily accessible to users via a system-prearranged table linking those data variables to the proper datasets (URLs). This online approach has been developed by utilizing a few existing data services such as Mirador as search engine; Giovanni for visualization; and OPeNDAP for data access, etc. The online Data Cookbook site at GES DISC is the current host for the bundle data. We are now also planning on developing an Automated Virtual Collection Framework that shall eventually accommodate the bundle data, as well as further improve our management in Big Data.

GES DISC

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics

The Integrated Sensor System Data Enhancement Package

The purpose of the Integrated Sensor System (ISS) Data Enhancement Package (DEP) is to improve the accuracies of the data obtained from the inflight tests performed on aircraft. The DEP is a microprocessor-based, flight-qualified electronics package that assimilates data from a Ring Laser Gyro (RGL) system, a standard NASA air data package, and other inputs. The DEP then processes these inputs in real-time to obtain optimal estimates of the aircraft velocity, attitude, and altitude. These estimates can be passed to the flight crew, downlinked, and/or stored on a mass storage medium. The DEP is now being built for the NASA Dryden Flight Research Center. Completion is anticipated in early 1984. A primary use of the ISS/DEP will be for the collection of quality data for the estimation of aircraft aerodynamic coefficients, including stability derivatives, using system identification methods. Initial anticipated applications will be on the AV-8B, F-14, and X-29 test aircraft.

Trankle, T. L.