Massive Data Set Analysis for NASA's Atmospheric Infrared Sounder
No abstract available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
No abstract available
Explore the source record for details and available documents.
The Virtual Observatory (VO) is realizing global electronic integration of astronomy data. One of the long-term goals of the U.S. VO project, the Virtual Astronomical Observatory (VAO), is development of services and protocols that respond to the growing size and complexity of astronomy data sets. This paper describes how VAO staff are active in such development efforts, especially in innovative strategies and techniques that recognize the limited operating budgets likely available to astronomers even as demand increases. The project has a program of professional outreach whereby new services and protocols are evaluated.
This talk discusses a method for creating low-volume versions of massive geophysical data sets that approximately retain high-resolution data structure.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
This paper describes a method for summaraizing data in a way that approximately preserves high-resolution data structure while reducing data volume and maintaining global integrity of very large, remote sensing data sets. The method is under development for one of Terra's instruments, the Multi-angle Imaging SpectroRadiometer (MISR).
The purpose of the workshop was to invite statisticians, applied mathematicians, computer scientists, data system architects, experts in remote sensing technology, and Climate and Earth System scientists to review, discuss, and plan research on issues related to large-scale, efficient analysis of distributed data using spatial statistical methods. Our motivation in organizing this event was to catalyze interchange among experts on the fast-emerging problem of analysis of distributed data. As part of SAMSI's 2017-2018 Program on Mathematical and Statistical Methods for Climate and the Earth System, a Working Group on Remote Sensing was established to address statistical and mathematical research problems in the analysis of remote sensing data. The Working Group has five subgroups: 1) Spatial Retrieval Methodology (the so-called \Spatial-X" subgroup); 2) Spatial Analysis for Hyperspectral Data (the so-called \Spatial-Y" subgroup); 3) Emulators for Complex Forward Models; 4) Optimization for Remote Sensing Retrievals; and 5) Theory of Data Systems (ToDS). The ToDS subgroup spent the first half of this academic year formulating a framework in which to consider the joint problem of a) optimizing statistical methods for environments where data are distributed and too large to move to a central location, and b) the design of data system infrastructures within which to implement those statistical methods. To x ideas, the Workshop focused on spatial statistical methods. To date there are many new spatial statistical methods designed with massive data sets in mind, in the literature. However, very few have been implemented for remote sensing data, and none have been implemented in operational settings like those used by NASA and NOAA. A major impediment to their use in these cases is that the data are not only massive, but are stored in different physical locations. These data must be brought together in some way in order to estimate spatial covariance functions, but moving data to a central location for analysis is tedious at best and impossible at worst. Some remote data reduction is almost certainly necessary, but how much? What are the consequences for inference? The fundamental issue underlying these questions is how to navigate the trade-space between costs and uncertainty in the estimates or inferences that are ultimately produced.
Explore the source record for details and available documents.
The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.
Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.
Impervious surfaces, mainly artificial structures and roads, cover less than 1% of the world's land surface (1.3% over USA). Regardless of the relatively small coverage, impervious surfaces have a significant impact on the environment. They are the main source of the urban heat island effect, and affect not only the energy balance, but also hydrology and carbon cycling, and both land and aquatic ecosystem services. In the last several decades, the pace of converting natural land surface to impervious surfaces has increased. Quantitatively monitoring the growth of impervious surface expansion and associated urbanization has become a priority topic across both the physical and social sciences. The recent availability of consistent, global scale data sets at 30m resolution such as the Global Land Survey from the Landsat satellites provides an unprecedented opportunity to map global impervious cover and urbanization at this resolution for the first time, with unprecedented detail and accuracy. Moreover, the spatial resolution of Landsat is absolutely essential to accurately resolve urban targets such a buildings, roads and parking lots. With long term GLS data now available for the 1975, 1990, 2000, 2005 and 2010 time periods, the land cover/use changes due to urbanization can now be quantified at this spatial scale as well. In the Global Land Survey - Imperviousness Mapping Project (GLS-IMP), we are producing the first global 30 m spatial resolution impervious cover data set. We have processed the GLS 2010 data set to surface reflectance (8500+ TM and ETM+ scenes) and are using a supervised classification method using a regression tree to produce continental scale impervious cover data sets. A very large set of accurate training samples is the key to the supervised classifications and is being derived through the interpretation of high spatial resolution (approx. 2 m or less) commercial satellite data (Quickbird and Worldview2) available to us through the unclassified archive of the National Geospatial Intelligence Agency (NGA). For each continental area several million training pixels are derived by analysts using image segmentation algorithms and tools and then aggregated to the 30m resolution of Landsat. Here we will discuss the production/testing of this massive data set for Europe, North and South America and Africa, including assessments of the 2010 surface reflectance data. This type of analysis is only possible because of the availability of long term 30m data sets from GLS and shows much promise for integration of Landsat 8 data in the future.
Modern astronomical surveys detect asteroids by linking together their appearances across multiple images taken over time. This approach faces limitations in detecting faint asteroids and handling the computational complexity of trajectory linking. Here, we present a novel method that adapts “digital tracking” – traditionally used for short-term linear asteroid motion across images – to work with large-scale synoptic surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time (Rubin/LSST). Our approach combines hundreds of sparse observations of individual asteroids across their non-linear orbital paths to enhance detection sensitivity by several magnitudes. To address the computational challenges of processing massive data sets and dense orbital phase spaces, we developed a specialized high-performance computing architecture. We demonstrate the effectiveness of our method through experiments that take advantage of the extensive computational resources at Lawrence Livermore National Laboratory. This work enables the detection of significantly fainter asteroids in existing and future survey data, potentially increasing the observable asteroid population by orders of magnitude across different orbital families, from near-Earth objects (NEOs) to Kuiper belt objects (KBOs).
Multi-dimensional data contained in very large databases is efficiently and accurately clustered to determine patterns therein and extract useful information from such patterns. Conventional computer processors may be used which have limited memory capacity and conventional operating speed, allowing massive data sets to be processed in a reasonable time and with reasonable computer resources. The clustering process is organized using a clustering feature tree structure wherein each clustering feature comprises the number of data points in the cluster, the linear sum of the data points in the cluster, and the square sum of the data points in the cluster. A dense region of data points is treated collectively as a single cluster, and points in sparsely occupied regions can be treated as outliers and removed from the clustering feature tree. The clustering can be carried out continuously with new data points being received and processed, and with the clustering feature tree being restructured as necessary to accommodate the information from the newly received data points.
There is a wealth of cosmological information encoded in the spatial power spectrum of temperature anisotropies of the cosmic microwave background. The sky, when viewed in the microwave, is very uniform, with a nearly perfect blackbody spectrum at 2.7 degrees. Very small amplitude brightness fluctuations (to one part in a million!!) trace small density perturbations in the early universe (roughly 300,000 years after the Big Bang), which later grow through gravitational instability to the large-scale structure seen in redshift surveys... In this talk, I will discuss a Bayesian formulation of this problem; discuss a Gibbs sampling approach to numerically sampling from the Bayesian posterior, and the application of this approach to the first-year data from the Wilkinson Microwave Anisotropy Probe. I will also comment on recent algorithmic developments for this approach to be tractable for the even more massive data set to be returned from the Planck satellite.
NASA's highly successful Kepler Mission has revolutionized our understanding of the Galaxy. We now know that planets, even Earth-size planets in the habitable zone, are common. With the end of the Kepler Mission we now look to the future with the Transiting Exoplanet Survey Satellite (TESS) which will discover thousands of exoplanets in orbit around the brightest stars in the sky. In a two-year survey, TESS will perform an all-sky search of more than 200,000 stars for temporary drops in brightness caused by planetary transits. With Kepler and TESS, humanity is finally at the verge of studying the masses, sizes, densities, orbits, and atmospheres of a large cohort of small planets, including a sample of rocky worlds in the habitable zones of their host stars which may prove to host life. The massive data sets generated by Kepler and TESS must be meticulously combed for the weakest planetary signals every month. While a daunting and error-prone task for humans, this is an exciting opportunity for the breakthroughs recently seen in machine learning. Specifically, traditional methods for identifying planet transits require extensive data processing pipelines followed by extensive human vetting. This manual process risks loss of information due to the data processing and to inconsistency and biases due to individual human vetters. The latest advancements in machine learning will allow an objective classifier to minimize the losses of information and greatly lessen the burden on the human vetters, in addition to providing assessment of quality and score to each planet candidate, freeing the humans to concentrate on border cases and other more interesting investigations.
In previous papers we've shown how a well known data compression algorithm called Entropy-constrained Vector Quantization ( can be modified to reduce the size and complexity of very large, satellite data sets. In this paper, we descuss how to visualize and understand the content of such reduced data sets.
With the recent surge of success in big-data driven deep learning problems, many of these frameworks focus on the notion of architecture design and utilizing massive databases. However, in some scenarios massive sets of data may be difficult, and in some cases infeasible, to acquire. In this paper we discuss a trajectory-based framework that quickly learns the underlying decision manifold of binary simulation classifications while judiciously selecting exploratory target states to minimize the number of required simulations. Furthermore, we draw particular attention to the simulation prediction application idealized to the case where failures in simulations can be predicted and avoided, providing machine intelligence to novice analysts. We demonstrate this framework in various forms of simulations and discuss its efficacy.