Search NASASearch

SEARCH · Search NASA

Results for “Data augmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Comparison of QuikSCAT and GPS-Derived Ocean Surface Winds

The Colorado Center for Astrodynamics has completed a study comparing ocean surface winds derived from GPS bistatic measurements with QuikSCAT wind fields. We have also compiled an extensive database of the bistatic GPS flight data collected by NASA Langley Research Center over the last several years. The GPS data are augmented with coincident data from QuikSCAT, buoys, TOPEX, and ERS.

Axelrad, Penina

Low Latency DESDynI Data Products for Disaster Response, Resource Management and Other Applications

We are developing onboard processor technology targeted at the L-band SAR instrument onboard the planned DESDynI mission to enable formation of SAR images onboard opening possibilities for near-real-time data products to augment full data streams. Several image processing and/or interpretation techniques are being explored as possible direct-broadcast products for use by agencies in need of low-latency data, responsible for disaster mitigation and assessment, resource management, agricultural development, shipping, etc. Data collected through UAVSAR (L-band) serves as surrogate to the future DESDynI instrument. We have explored surface water extent as a tool for flooding response, and disturbance images on polarimetric backscatter of repeat pass imagery potentially useful for structural collapse (earthquake), mud/land/debris-slides etc. We have also explored building vegetation and snow/ice classifiers, via support vector machines utilizing quad-pol backscatter, cross-pol phase, and a number of derivatives (radar vegetation index, dielectric estimates, etc.). We share our qualitative and quantitative results thus far.

applications

Using Concurrent Cardiovascular Information to Augment Survival Time Data from Orthostatic Tilt Tests

Orthostatic Intolerance (OI) is the propensity to develop symptoms of fainting during upright standing. OI is associated with changes in heart rate, blood pressure and other measures of cardiac function. Problem: NASA astronauts have shown increased susceptibility to OI on return from space missions. Current methods for counteracting OI in astronauts include fluid loading and the use of compression garments. Multivariate trajectory spread is greater as OI increases. Pairwise comparisons at the same time within subjects allows incorporation of pass/fail outcomes. Path length, convex hull area, and covariance matrix determinant do well as statistics to summarize this spread Missing data problems Time series analysis need many more time points per OTT session treatment of trend? how incorporate survival information?

Feiveson, Alan H.

A Preliminary Study on the Feasibility of Large Language Models for Detecting Micro-Behaviors Among Team Members in Space Missions

Large-language models (LLMs) have been recently used for spoken language understanding (SLU) to infer meaning and semantics from speech in tasks such as speaker intent and sentiment classification. Due to being trained on large amounts of data, and their ability to understand context and relationships between words, LLMs are competent, enabling them to generalize across tasks without requiring many task-specific training samples. This research examines the feasibility of few-shot learning in LLMs for detecting subtle, brief, and possibly unconscious interactions between team members, called ``micro-behaviors," and provides insights into the appropriate design of LLMs for this task. Our data came from 5 teams participating in a 45-day mission at the US National Aeronautics and Space Administration’s (NASA) Human Exploration Research Analog (HERA). More specifically we used data collected from team interaction battery (TIB) tasks teams performed five times in-mission which comprise an average 1.5 hours of conversation data per day. Micro-behaviors were coded according to an adapted version of Smith & Griffins (2022) theoretical framework in terms of Violation (i.e., presence of valenced behavior, uplifting/positive or discouraging/negative), Intensity (i.e., force of behavior in terms of how uplifting or discouraging is the behavior), and Intent (i.e., motive of the behavior in terms of whether it was deliberate or unintentional). We explore the ability of LLMs to detect the presence and intensity of micro-behaviors. We examine employing and fine-tuning readily available LLMs (i.e., RoBERTa, DistilBERT), as well as prompting state-of-the-art sequence classification models (i.e., Llama-2, Llama-3). In a total of 13,058 conversational turns (17.8% uplifting, 3.3% discouraging, 75.76% neutral, 3.14% nulls), we compute the macro F1-score of the 3-way micro-behavior classification task (i.e., classifying among uplifting, discouraging, and neutral; 33% chance). Results indicate that the RoBERTa model achieves a F1-score of 36.2% (uplift: 43.3% precision (P), 15.1% recall (R); discourage: 20% P, 0.5% R). These results significantly improve when we augment the data via paraphrasing in the RoBERTa model, reaching a 41.2% macro F1-score (uplift: 37.7% P, 86.3% R; discourage: 3.5% P, 1.8% R). Finally, the Llama-2 model with 3-shot prompting yields 38% macro F1-score (uplift: 28.7% P, 20% R; discourage: 7.2% P, 18% R), which is slightly better compared to the RoBERTa model without data augmentation, highlighting the effectiveness of sequence classification models in detecting minority classes with a small sample size. Findings indicate that LLMs hold potential to detect subtle behaviors in conversations, which could be valuable in assessing team behavior in space exploration missions. Future studies will evaluate the performance of different LLM prompting strategies or fine-tuning methods.

Ankush Raut

Probability Bounds Analysis Applied to Multi-Purpose Crew Vehicle Nonlinearity

The Multi-Purpose Crew Vehicle (MPCV) Program Orion vehicle finite element model (FEM) was updated based on a modal test performed by Lockheed Martin. Due to nonlinearity observed in the test results, linear low force level (LL) and high force level (HL) FEMs were developed for use during various Space Launch System (SLS) flight regimes depending on expected forcing levels. Uncertainty models were derived for the combined MPCV and MPCV Stage Adaptor LL and HL Hurty/Craig-Bampton (HCB) components based on the MPCV structural test article Configuration 4 modal test-analysis correlation results. Subsequently, system-level uncertainty quantification analyses were performed using both models for various SLS flight configurations to determine the impact of the nonlinearity on important system metrics. The system metrics included both transfer functions associated with attitude control and dynamic loads associated with aerodynamic buffeting during ascent. In each case, an independent Monte Carlo (MC) analysis was performed, and no attempt was made to combine the results. The Hybrid Parametric Variation (HPV) method was used to develop the LL and HL MPCV HCB uncertainty models. The HPV method provides both parametric and non-parametric components of uncertainty. The non-parametric uncertainty accounts for the difference in model-form between the linearized analytical model and the corresponding linearized component test results in the form of mode shapes and frequencies at that force level. This linear model-form uncertainty is implemented in the HPV method using random matrix theory. However, the HPV uncertainty models developed for the linear LL and HL MPCV components do not account for the nonlinearity in the MPCV. With respect to the linearized models, this nonlinearity is also an uncertainty in model form, but in this case, it must be treated independently as an epistemic uncertainty. It represents a lack of knowledge, in contrast to an aleatory uncertainty due to the randomness of a variable. In the case of an epistemic variable, the true value is unknown, only the interval within which it lies is known. Epistemic uncertainty can be reduced with increased knowledge, while in general, aleatory uncertainty cannot. This work combines the epistemic uncertainty due to the MPCV nonlinearity with the parametric and non-parametric uncertainty within the HPV method using a second order propagation approach. The LL and HL test data is augmented with surrogate test data derived from a nonlinear MPCV representation. The impact of the MPCV nonlinearity on system response statistics is determined using a series of cumulative distribution functions in the form of a horsetail plot, or p-box. This results in an interval of probabilities for a specific response value, or an interval of response values at a specific probability.

Daniel C Kammer

The effect of changing environmental conditions on microwave signatures of forest ecosystems - Preliminary results of the March 1988 Alaskan aircraft SAR experiment

In preparation for the ESA ERS-1 mission, a series of multitemporal, multifrequency, multipolarization aircraft SAR data sets were acquired near Fairbanks in March 1988. P-, L-, and C-band data were acquired with the NASA/JPL Airborne SAR on five different days over a period of two weeks. The airborne data were augmented with intensive ground calibration data as well as detailed simultaneous in situ measurements of the geometric, dielectric, and moisture properties of the snow and forest canopy. During the time period over which the SAR data were collected, the environmental conditions changed significantly; temperatures ranged from unseasonably warm (1 to 9 C) to well below freezing (-8 to -15 C), and the moisture content of the snow and trees changed from a liquid to a frozen state. The SAR data clearly indicate the radar return is sensitive to these changing environmental factors, and preliminary analysis of the L-band SAR data shows a 0.4 to 5.8 dB increase (depending on polarization and canopy type) in the radar cross section of the forest stands under the warm conditions relative to the cold. These SAR observations are consistent with predictions from a theoretical scattering model.

Way, Jobea

Utilizing Earth Observations to Understand Landscape Patterns and Assist in Wildlife Management in Iona National Park, Angola

Following the end of the Angolan Civil War (1975-2002), human habitation in Iona National Park has grown exponentially, as has the livestock population. An ongoing drought beginning in 2017 has brought people, livestock, and wildlife into increasing competition for resources within the park. This study used Earth observation data, primarily Landsat and Sentinel imagery, to examine landscape trends to improve wildlife preservation approaches in Iona National Park, Angola. In collaboration with the NGO African Parks, we developed a robust land use and land cover (LULC) classification model using remote sensing data to augment sparse ground-based data in this arid land region. We used Google Earth Engine and a random forest classifier to map vegetation types, water bodies, and potential wildlife habitats. This analysis resulted in a high spatial resolution LULC time-series between 1984-2023, highlighting critical periods of socioecological change over the past 40 years. These results increased the partner’s ability to make scientifically grounded decisions about resource allocation and conservation priorities. This analysis supports the feasibility of applying remote sensing techniques coupled with machine learning models in dry regions, where standard survey methods are frequently limited by accessibility and resource availability. However, we identified limitations in ground-truth data and the difficulty of recognizing certain vegetation types in arid areas. Despite these limitations, the study demonstrated Earth observations' ability to transform wildlife management techniques in distant and data-scarce locations, providing a reproducible foundation for similar ecosystems around the world.

Emmanuel Aklie

NeMO-Net & Fluid Lensing: The Neural Multi-Modal Observation & Training Network for Global Coral Reef Assessment Using Fluid Lensing Augmentation of NASA EOS Data

We present preliminary results from NASA NeMO-Net, the first neural multi-modal observation and training network for global coral reef assessment. NeMO-Net is an open-source deep convolutional neural network (CNN) and interactive active learning training software in development which will assess the present and past dynamics of coral reef ecosystems. NeMO-Net exploits active learning and data fusion of mm-scale remotely sensed 3D images of coral reefs captured using fluid lensing with the NASA FluidCam instrument, presently the highest-resolution remote sensing benthic imaging technology capable of removing ocean wave distortion, as well as hyperspectral airborne remote sensing data from the ongoing NASA CORAL mission and lower-resolution satellite data to determine coral reef ecosystem makeup globally at unprecedented spatial and temporal scales. Aquatic ecosystems, particularly coral reefs, remain quantitatively misrepresented by low-resolution remote sensing as a result of refractive distortion from ocean waves, optical attenuation, and remoteness. Machine learning classification of coral reefs using FluidCam mm-scale 3D data show that present satellite and airborne remote sensing techniques poorly characterize coral reef percent living cover, morphology type, and species breakdown at the mm, cm, and meter scales. Indeed, current global assessments of coral reef cover and morphology classification based on km-scale satellite data alone can suffer from segmentation errors greater than 40%, capable of change detection only on yearly temporal scales and decameter spatial scales, significantly hindering our understanding of patterns and processes in marine biodiversity at a time when these ecosystems are experiencing unprecedented anthropogenic pressures, ocean acidification, and sea surface temperature rise. NeMO-Net leverages our augmented machine learning algorithm that demonstrates data fusion of regional FluidCam (mm, cm-scale) airborne remote sensing with global low-resolution (m, km-scale) airborne and spaceborne imagery to reduce classification errors up to 80% over regional scales. Such technologies can substantially enhance our ability to assess coral reef ecosystems dynamics.

satellite data

The near-Earth magnetic field at 1980 determined from MAGSAT data

Data from the MAGSAT spacecraft for November 1979 through April 1980 and from 91 magnetic observatories for 1978 through 1982 are used to derive a spherical harmonic model of the Earth's main magnetic field and its secular variation. Constant coefficients are determined through degree and order 13 and secular variation coefficients through degree and order 10. The first degree external terms and corresponding induced internal terms are given as a function of Dst. Preliminary modeling using separate data sets at dawn and dusk local time showed that the dusk data contains a substantial field contribution from the equatorial electrojet current. The final data set is selected first from dawn data and then augmented by dusk data to achieve a good geographic data distribution for each of three time periods: (1) November/December, 1979; (2) January/February; 1980; (3) March/April, 1980. A correction for the effects of the equatorial electrojet is applied to the dusk data utilized. The solution included calculation of fixed biases, or anomalies, for the observation data.

Langel, R. A.

The near-earth magnetic field at 1980 determined from Magsat data

Data from the Magsat spacecraft for November 1979 through April 1980 and from 91 magnetic observatories for 1978 through 1982 are used to derive a spherical harmonic model of the earth's main magnetic field and its secular variation. Constant coefficients are determined through degree and order 13 and secular variation coefficients through degree and order 10. The first degree external terms and corresponding induced internal terms are given as a function of Dst. Preliminary modeling using separate data sets at dawn and dusk local time showed that the dusk data contains a substantial field contribution from the equatorial electrojet current. The final data set is selected first from dawn data and then augmented by dusk data to achieve a good geographic data distribution for each of three time periods: (1) November/December, 1979; (2) January/February, 1980; (3) March/April, 1980. A correction for the effects of the equatorial electrojet is applied to the dusk data utilized. The solution included calculation of fixed biases, or anomalies, for the observation data.

Langel, R. A.

A Martian Telecommunications Network: UHF Relay Support of the Mars Exploration Rovers by the Mars Global Surveyor, Mars Odyssey, and Mars Express Orbiters

NASA and ESA have established an international network of Mars orbiters, outfitted with relay communications payloads, to support robotic exploration of the red planet. Starting in January, 2004, this network has provided the Mars Exploration Rovers with telecommunications relay services, significantly increasing rover engineering and science data return while enhancing mission robustness and operability. Augmenting the data return capabilities of their X-band direct-to-Earth links, the rovers are equipped with UHF transceivers allowing data to be relayed at high rate to the Mars Global Surveyor (MGS), Mars Odyssey, and Mars Express orbiters. As of 21 July, 2004, over 50 Gbits of MER data have been obtained, with nearly 95% of that data returned via the MGS and Odyssey UHF relay paths, allowing a large increase in science return from the Martian surface relative to the X-band direct-to-Earth link. The MGS spacecraft also supported high-rate UHF communications of MER engineering telemetry during the critical period of entry, descent, and landing (EDL), augmenting the very low-rate EDL data collected on the X-band direct-to-Earth link. Through adoption of the new CCSDS Proximity-1 Link Protocol, NASA and ESA have achieved interoperability among these Mars assets, as validated by a successful relay demonstration between Spirit and Mars Express, enabling future interagency cross-support and establishing a truly international relay network at Mars.

telecommunications

Reinforcement Learning Applied to Cognitive Space Communications

The future of space exploration depends on robust, reliable communication systems. As the number of such communication systems increase, automation is fast becoming a requirement to achieve this goal. A reinforcement learning solution can be employed as a possible automation method for such systems. The goal of this study is to build a reinforcement learning algorithm which optimizes data throughput of a single actor. A training environment was created to simulate a link within the NASA Space Communication and Navigation (SCaN) infrastructure, using state of the art simulation tools developed by the SCaN Center for Engineering, Networks, Integration, and Communications (SCENIC) laboratory at NASA Glenn Research Center to obtain the closest possible representation of the real operating environment. Reinforcement learning was then used to train an agent inside this environment to maximize data throughput. The simulation environment contained a single actor in low earth orbit capable of communicating with twenty-five ground stations that compose the Near-Earth Network (NEN). Initial experiments showed promising training results, so additional complexity was added by augmenting simulation data with link fading profiles obtained from real communication events with the International Space Station. A grid search was performed to find the optimal hyperparameters and model architecture for the agent. Using the results of the grid search, an agent was trained on the augmented training data. Testing shows that the agent performs well inside the training environment and can be used as a foundation for future studies with added complexity and eventually tested in the real space environment.

Schubert, Carson D.

Public Health Data Applications Using the CDC Tracking Network: Augmenting Environmental Hazard Information with Lower-latency NASA Data

Exposure to environmental hazards is an important determinant of health, and the frequency and severity of exposures is expected to be impacted by climate change. Through a partnership with the U.S. National Aeronautics and Space Administration, the U.S. Centers for Disease Control and Prevention’s National Environmental Public Health Tracking Network is integrating timely observations and model data of priority environmental hazards into its publicly accessible Data Explorer (https://ephtracking.cdc.gov/DataExplorer/). Newly integrated datasets over the contiguous U.S. (CONUS) include: daily 5-day forecasts of air quality based on the Goddard Earth Observing System Composition Forecast (GEOS-CF), daily historical (1980-present) concentrations of speciated PM2.5 based on the Modern Era Retrospective analysis for Research and Applications, version 2 (MERRA-2), and Moderate Resolution Imaging Spectroradiometer (MODIS) daily near real-time maps of flooding (MCDWD). Data integrated into the CDC Tracking Network are broadly intended to improve community health through action by informing both research and early warning activities, including (1) describing temporal and spatial trends in disease and potential environmental exposures, (2) identifying populations most affected, (3) generating hypotheses about associations between health and environmental exposures, and (4) developing, guiding, and assessing environmental public health policies and interventions aimed at reducing or eliminating health outcomes associated with environmental factors.

air quality

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert

The role of FGGE data in the understanding and prediction of atmospheric planetary waves

The augmented observational data base provided by FGGE offers a unique opportunity to improve the understanding and prediction of atmospheric planetary waves. Substantial progress was made, but a number of problems remain. The research on planetary waves conducted thus far with FGGE data is reviewed. Areas of progress are summarized, and some remaining problems are discussed.

Baker, W. E.

NASA Center for Climate Simulation (NCCS) Presentation

The NASA Center for Climate Simulation (NCCS) offers integrated supercomputing, visualization, and data interaction technologies to enhance NASA's weather and climate prediction capabilities. It serves hundreds of users at NASA Goddard Space Flight Center, as well as other NASA centers, laboratories, and universities across the US. Over the past year, NCCS has continued expanding its data-centric computing environment to meet the increasingly data-intensive challenges of climate science. We doubled our Discover supercomputer's peak performance to more than 800 teraflops by adding 7,680 Intel Xeon Sandy Bridge processor-cores and most recently 240 Intel Xeon Phi Many Integrated Core (MIG) co-processors. A supercomputing-class analysis system named Dali gives users rapid access to their data on Discover and high-performance software including the Ultra-scale Visualization Climate Data Analysis Tools (UV-CDAT), with interfaces from user desktops and a 17- by 6-foot visualization wall. NCCS also is exploring highly efficient climate data services and management with a new MapReduce/Hadoop cluster while augmenting its data distribution to the science community. Using NCCS resources, NASA completed its modeling contributions to the Intergovernmental Panel on Climate Change (IPCG) Fifth Assessment Report this summer as part of the ongoing Coupled Modellntercomparison Project Phase 5 (CMIP5). Ensembles of simulations run on Discover reached back to the year 1000 to test model accuracy and projected climate change through the year 2300 based on four different scenarios of greenhouse gases, aerosols, and land use. The data resulting from several thousand IPCC/CMIP5 simulations, as well as a variety of other simulation, reanalysis, and observationdatasets, are available to scientists and decision makers through an enhanced NCCS Earth System Grid Federation Gateway. Worldwide downloads have totaled over 110 terabytes of data.

Webster, William P.