Search NASASearch

SEARCH · Search NASA

Results for “data augmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Augmented Reality Data Generation for Training Deep Learning Neural Network

One of the major challenges in deep learning is retrieving sufficiently large labeled training datasets, which can become expensive and time consuming to collect. A unique approach to training segmentation is to use Deep Neural Network (DNN) models with a minimal amount of initial labeled training samples. The procedure involves creating synthetic data and using image registration to calculate affine transformations to apply to the synthetic data. The method takes a small dataset and generates a highquality augmented reality synthetic dataset with strong variance while maintaining consistency with real cases. Results illustrate segmentation improvements in various target features and increased average target confidence.

Torres, Gil

Toward more-robust, AI-enabled subsurface seismic imaging for geotechnical applications

Non-invasive seismic imaging has the potential to cost-effectively evaluate large volumes of subsurface material to inform geotechnical site investigation. However, seismic imaging using full waveform inversion (FWI) requires significant computational time and is dependent on an initial starting model. As a result, FWI has not yet been widely adopted into geotechnical practice. Previous efforts, on relatively simple two-layered models, indicate that data-driven artificial intelligence (AI) models may be as effective as FWI at predicting 2D images of shear wave velocity (V s ). Furthermore, the AI model predictions can be made almost instantaneously after data acquisition and do not require an initial starting model. We examine the generality of these findings by developing a new AI model for subsurface seismic imaging, whereby we make several notable contributions. First, we architect a multimodal AI model that combines time- and frequency-domain representations of the seismic wavefield to predict a 50 m by 20 m subsurface image of V s . Second, we developed a new diverse dataset of 100,000 images with their corresponding seismic wavefields to train the AI model. Third, we propose four physics-informed data augmentations for data-driven seismic imaging. Fourth, we develop two prediction consistency tests to evaluate the model’s performance when the true subsurface is unknown. Our final model, which has been made publicly available, is capable of predicting a subsurface V s image from a single seismic wavefield with an average, mean absolute percent error (MAPE) of 24 %. The predictive model is applied to a field dataset and shown to be consistent with local geology and shear-wave refraction measurements from the same location.

Artificial intelligence

Contrastive learning for robust representations of neutrino data

In neutrino physics, analyses often depend on large simulated datasets, making it essential for models to generalize effectively to real-world detector data. Contrastive learning, a well-established technique in deep learning, offers a promising solution to this challenge. By applying controlled data augmentations to simulated data, contrastive learning enables the extraction of robust and transferable features. This improves the ability of models trained on simulations to adapt to real experimental data distributions. In this paper, we investigate the application of contrastive learning methods in the context of neutrino physics. Through a combination of empirical evaluations and theoretical insights, we demonstrate how contrastive learning enhances model performance and adaptability. Additionally, we compare it to other domain adaptation techniques, highlighting the unique advantages of contrastive learning for this field. Published by the American Physical Society 2025

Wilkinson, Alex (ORCID:0000000253404506)

Navigation for space shuttle approach and landing using an inertial navigation system augmented by data from a precision ranging system or a microwave scan beam landing guidance system

A preliminary study has been made of the navigation performance which might be achieved for the high cross-range space shuttle orbiter during final approach and landing by using an optimally augmented inertial navigation system. Computed navigation accuracies are presented for an on-board inertial navigation system augmented (by means of an optimal filter algorithm) with data from two different ground navigation aids; a precision ranging system and a microwave scanning beam landing guidance system. These results show that augmentation with either type of ground navigation aid is capable of providing a navigation performance at touchdown which should be adequate for the space shuttle. In addition, adequate navigation performance for space shuttle landing is obtainable from the precision ranging system even with a complete dropout of precision range measurements as much as 100 seconds before touchdown.

Mcgee, L. A.

Imbalanced Multi-layer Cloud Classification with Advanced Baseline Imager (ABI) and CloudSat/CALIPSO Data

Clouds at different altitudes play different roles in Earth’s climate. Comprehensive understanding of overlapping clouds is important for climate and weather prediction. The East Pacific region is where El Ni˜no and La Ni˜na originate and where multi-layer clouds frequently occur. The overlap of clouds at different altitudes in this region increases the classification complexity for cloud-based climatological studies. Unlike prior work in cloud layer classification that assumes single layer or two-layer of clouds, in this work, we consider multi-layer cloud classification with 8 cloud-level classes (clear-sky, high, middle, low, high+middle, high+low, middle+low, high+middle+low). We develop and analyze machine learning models on features extracted from satellite images from the East Pacific regions collected by GOES Advanced Baseline Imager (ABI). These are used to classify CloudSat/CALIPSO observed multi-layer clouds. Due to the imbalanced nature of the data, we investigate the adoption of conventional resampling methods, as well as deep learning methods with data augmentation. In our experiments, we utilize the random forest classifier and Multilayer perceptron classifier with data augmentation methods to reduce the class imbalance during training. With these approaches, we achieve a classification accuracy of 83.6% without exploiting any ancillary information.

machine learning

Ground-Based VIS/NIR Reflectance Spectra of 25143 Itokawa: What Hayabusa will See and How Ground-Based Data can Augment Analyses

Planning for the arrival of the Hayabusa spacecraft at asteroid 25143 Itokawa includes consideration of the expected spectral information to be obtained using the AMICA and NIRS instruments. The rotationally-resolved spatial coverage the asteroid we have obtained with ground-based telescopic spectrophotometry in the visible and near-infrared can be utilized here to address expected spacecraft data. We use spectrophotometry to simulate the types of data that Hayabusa will receive with the NIRS and AMICA instruments, and will demonstrate them here. The NIRS will cover a wavelength range from 0.85 m, and have a dispersion per element of 250 Angstroms. Thus, we are limited in coverage of the 1.0 micrometer and 2.0 micrometer mafic silicate absorption features. The ground-based reflectance spectra of Itokawa show a large component of olivine in its surface material, and the 2.0 micrometer feature is shallow. Determining the olivine to pyroxene abundance ratio is critically dependent on the attributes of the 1.0- and 2.0 micrometer features. With a cut-off near 2,1 micrometer the longer edge of the 2.0- feature will not be obtained by NIRS. Reflectance spectra obtained using ground-based telescopes can be used to determine the regional composition around space-based spectral observations, and possibly augment the longer wavelength spectral attributes. Similarly, the shorter wavelength end of the 1.0 micrometer absorption feature will be partially lost to the NIRS. The AMICA filters mimic the ECAS filters, and have wavelength coverage overlapping with the NIRS spectral range. We demonstrate how merging photometry from AMICA will extend the spectral coverage of the NIRS. Lessons learned from earlier spacecraft to asteroids should be considered.

Vilas, Faith

Automated Pneumothorax Diagnosis using Deep Neural Networks

Thoracic ultrasound can provide information leading to rapid diagnosis of pneumothorax with improved accuracy over the standard physical examination and with higher sensitivity than anteroposterior chest radiography. However, the clinical We have Furthermore, remote environments, such as the battlefield or deep-space exploration, may lack expertise for diagnosing developed an automated image interpretation pipeline for the analysis of thoracic ultrasound data and the classification of pneumothorax events to provide decision support in such situations. Our pipeline consists of image preprocessing, data augmentation, and deep learning architectures for medical diagnosis. In this work, we demonstrate that robust, accurate interpretation of chest images and video can be achieved using deep neural networks. A number of novel image processing techniques were employed to achieve this result. Affine transformations were applied for data augmentation. Hyperparameters were optimized for learning rate, dropout regularization, batch size, and epoch iteration by a sequential model-based Bayesian approach. In addition, we utilized pretrained architecturesinterpretation of a patient medical image is highly operator dependent. certain pathologies., applying transfer learning and fine-tuning techniques to fully connected layers. Our pipeline yielded binary classification validation accuracies of 98.3% for M-mode images and 99.8% with B-mode video frames.

US Army collaboration

The interannual variability of polar cap recessions as a measure of Martian climate and weather: Using Earth-based data to augment the time line for the Mars observer mapping mission

The recessions of the polar ice caps are the most visible and most studied indication of seasonal change on Mars. Circumstantial evidence links these recessions to the seasonal cycles of CO2, water, and dust. The possible advent of a planet encircling storm during the Mars Observer (MO) mission will provide a detailed correlation with a cap recession for that one Martian year. That cap recession will then be compared with other storm and nonstorm years. MO data will also provide a stronger link between cap recessions and the water and CO2 cycles. Cap recession variability might also be used to determine the variability of these cycles. After nearly a century of valiant attempts at measuring polar cap recessions, including Mariner 9 and Viking data, MO will provide the first comprehensive dataset. In contrast to MO, the older data are much less detailed and precise and could be forgotten, except that it will still be the only information on interannual variability. By obtaining simultaneous Earth-based observations (including those from Hubble) during the MO mission, direct comparisons can be made between the datasets.

Martin, L. J.

Digital registration of topographic data and satellite MSS data for augmented spectral analysis

Results are presented for a project directed to assess the type and extent of long-term natural and man-made changes in floodplain features of a definite area in the Mississipi River Valley. 1:6000 scale maps were prepared from remotely sensed film/filter photographic coverage of the area at different times to compare changes in such features as channels, backwaters, vegetation as well as changes in the amount, location, and conditions of dredge spoil. Changes observed by this sequential photo comparison include abandonment of agricultural land followed by orderly secondary plant succession, the filling-in of shallows by sediment and the orderly succession of plants from a hydric to a more mesic environment and changes in deltas of small tributary streams in response to agricultural practices occurring at their headwaters. Plant colonization is observed on some spoil, and secondary movement of spoil is revealed in other areas.

Anuta, P. E.

Integration of Optical Coherence Tomography Scan Patterns to Augment Clinical Data Suite

Vision changes identified in long duration spaceflight astronauts has led Space Medicine at NASA to adopt a more comprehensive clinical monitoring protocol. Optical Coherence Tomography (OCT) was recently implemented at NASA, including on board the International Space Station in 2013. NASA is collaborating with Heidelberg Engineering to increase the fidelity of the current OCT data set by integrating the traditional circumpapillary OCT image with radial and horizontal block images at the optic nerve head. The retinal nerve fiber layer was segmented by two experienced individuals. Intra-rater (N=4 subjects and 70 images) and inter-rater (N=4 subjects and 221 images) agreement was performed. The results of this analysis and the potential benefits will be presented.

Mason, S.

A 360-degree and -order model of Venus topography

This report presents the most recent spherical harmonic topography model of Venus developed at Jet Propulsion Laboratory. It was produced by a spherical harmonic analysis of the most complete set of Magellan altimetry data, augmented by Pioneer Venus and Venera data. The harmonic coefficients of the topography were computed to degree and order 360. Compared to previous topography models, this one has the highest correlation with the gravity field of Venus.

Rappaport, Nicole

A 360-Degree and -Order Model of Venus Topography

This report presents the most recent spherical harmonic topography model of Venus developed at Jet Propulsion Laboratory. It was produced by a spherical harmonic analysis of the most complete set of Magellan altimetry data, augmented by Pioneer Venus and Venera data. The harmonic coefficients of the topography were computed to degree and order 360. Compared to previous topography models, this one has the highest correlation with the gravity field of Venus.

Venus venus topgraphy harmonic analysis spherical

NASA Pilot-Engaged Expert Response Using IBM Watson Technology: Prototype Evaluation of Knowledge Retrieval System

NASA Langley Research Center and IBM have been investigating the use of IBM Watson technology in aerospace research and development. One application of Watson technology is the Pilot-Engaged Expert Response (PEER) use case. The PEER system is envisioned as an in-cockpit advisor that will act as a source of situationally-relevant information for pilots and other flight crew members to assist in decision making about real-time events and situations that arise in the course of aircraft operations. PEER will make available vast stores of knowledge and information quickly and directly, putting important informational resources where they are needed most. IBM has worked with NASA to develop an architecture and articulate a roadmap for the development of the PEER system. That vision is built around Watson Discovery Advisor (WDA) software solution, derived from IBM's Jeopardy!-winning automatic question answering system. PEER makes use of WDA's sophisticated question-answering capabilities as its core, adding important User Interface components and other customizations for the cockpit environment, including communication with flight systems and other external data sources. The development plan for PEER includes four development stages, with the current project constituting the first phase. In this project, a prototype instance of PEER was successfully adapted to the aviation domain, enabling users to ask questions about aviation topics and receive useful and accurate answers to these questions. Major tasks accomplished include the development of procedures for domain adaptation through automatic lexicon extraction from domain glossaries; generation of question-answer training data which was used to train the system; and assessment of the effectiveness of domain adaptation, which showed a dramatic improvement in the ability of the PEER system to answer domain-relevant questions. In addition, the vision for the PEER system was pushed forward by the articulation of a plan for the automatic enhancement of question-answering with contextual information. This initial phase focused on two main goals: 1) the targeted domain adaptation of the underlying WDA system to the aviation domain; and, 2) the design of the software systems needed to leverage flight-contextual data. Domain adaptation of the WDA system proceeds via three main activities: Domain data ingestion, lexical customization and model training. A textual corpus consisting of 1,147 individual documents with more than 7.5 million words of text was ingested into the system and this served as the basis of all further development. A domain lexicon of over 3,500 aviation-domain terms was semi-automatically generated from domain documents and used to train the system. In addition, a set of over 500 question-answer (QA) pairs relevant to the PEER use case was developed; these were used to train and assess the system. These important first steps established the basis for the PEER system. In addition, steps were taken towards the integration of the PEER system into the cockpit environment with the development of a functional design for the Contextual Data Augmentation (CDA) subsystem. This subsystem brings to bear contextual data to improve system responses. It has three main submodules: the Contextual Data Collection module, the Contextual Data Selection module, and the Contextual QA Augmentation module. These modules form a processing pipeline that addresses the problems associated with automatically integrating information from external resources into the knowledge-retrieval mechanism.

Machine learning

A Simple Model of Pulsed Ejector Thrust Augmentation

A simple model of thrust augmentation from a pulsed source is described. In the model it is assumed that the flow into the ejector is quasi-steady, and can be calculated using potential flow techniques. The velocity of the flow is related to the speed of the starting vortex ring formed by the jet. The vortex ring properties are obtained from the slug model, knowing the jet diameter, speed and slug length. The model, when combined with experimental results, predicts an optimum ejector radius for thrust augmentation. Data on pulsed ejector performance for comparison with the model was obtained using a shrouded Hartmann-Sprenger tube as the pulsed jet source. A statistical experiment, in which ejector length, diameter, and nose radius were independent parameters, was performed at four different frequencies. These frequencies corresponded to four different slug length to diameter ratios, two below cut-off, and two above. Comparison of the model with the experimental data showed reasonable agreement. Maximum pulsed thrust augmentation is shown to occur for a pulsed source with slug length to diameter ratio equal to the cut-off value.

Wilson, Jack

Information Fusion and Data Analytics for Human Lunar Exploration (CIF REPORT: Detailed PI Write-up)

The Information Fusion & Data Analytics (IFDA) project commenced in FY20, continued through FY21, and its final platform development phase continues in FY22. The objective remains the fusion and rapid accessibility of large quantities of disparate sourced human spaceflight data. IFDA is a platform tailored for NA (S&MA) to develop highly advanced operational data integration and analysis techniques. IFDA leverages the JSC ER7 modeling, simulation,and data fusion capabilities to collect, warehouse, and augment data human exploration data integration and analysis techniques. The IFDA project’s integrated data visualizations have been demonstrated in two validation scenarios in FY21, and provided the architecture and platform basis for development of a full-scale data analysis suite and storage solution useful to all JSC organizations engaged in real time operations and safety tasks. Scenarioand prototypical development including the construction of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, is central to IFDA Phase 3 and provides a demonstrable pathway for the Digital Transformation Program. IFDA Phase 3 is focused on data provider, data utilizer, and SME hands-on workshops that will conclude the Dem / Valphase and deliver a program-ready data integration tool as a product.

information fusion

Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution

When doing multivariate data analysis, one commonobstacle is the presence of incomplete observations, i.e., observationsfor which one or more key fields are blank. Missing datais often countered by deleting entire observations that containmissing data. The negative effects of deleting entire observationsare multiple: deleting observations reduces sample size andcan also result in biased inferences even if data is missing atrandom. In addition, knowledge contained within incompleteobservations is knowledge lost when they are deleted– and theeffort spent collecting that knowledge is effort wasted. Data imputationmethods, or methods of statistically “filling-in” missingdata, can help combat small sample sizes by using the existinginformation in partially complete observations with the end goalof producing less biased and higher confidence inferences. Whena sample from a multivariate normal population is only partiallycomplete, and the missing data meets appropriate assumptions(missing at random), robust data imputation of the missing datacan be implemented with monotone data augmentation (MDA)using the multivariate t distribution.Missing data imputation is applied to data from the NASA InstrumentCost Model (NICM) using the MDA algorithm underthe assumption of having a multivariate t distribution with fixeddegrees of freedom. A sensitivity analysis to the degrees offreedom parameter is presented to demonstrate robustness ofthe multivariate t distribution when dealing with small samplesas compared to the multivariate normal distribution.

DiNicola, Michael

Information Fusion & Analytics for Human Lunar Exploration

The Information Fusion & Data Analytics (IFDA) project commenced in FY20, continued through FY21, and its final platform development phase continues in FY22. The objective remains the fusion and rapid accessibility of large quantities of disparate sourced human spaceflight data. IFDA is a platform tailored for NA (S&MA) to develop highly advanced operational data integration and analysis techniques. IFDA leverages the JSC ER7 modeling, simulation,and data fusion capabilities to collect, warehouse, and augment data human exploration data integration and analysis techniques. The IFDA project’s integrated data visualizations have been demonstrated in two validation scenarios in FY21, and provided the architecture and platform basis for development of a full-scale data analysis suite and storage solution useful to all JSC organizations engaged in real time operations and safety tasks. Scenarioand prototypical development including the construction of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, is central to IFDA Phase 3 and provides a demonstrable pathway for the Digital Transformation Program. IFDA Phase 3 is focused on data provider, data utilizer, and SME hands-on workshops that will conclude the Dem / Valphase and deliver a program-ready data integration tool as a product.

information fusion

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE