Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Interpretation, Statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Identification of the Hard X-Ray Source Dominating the E>25 keV Emission of the Nearby Galaxy M31

We report the identification of a bright hard X-ray source dominating the M31 bulge above 25 keV from a simultaneous NuSTAR-Swift observation. We find that this source is the counterpart to Swift J0042.6+4112, which was previously detected in the Swift BAT All-sky Hard X-ray Survey. This Swift BAT source had been suggested to be the combined emission from a number of point sources; our new observations have identified a single X-ray source from 0.5 to 50 keV as the counterpart for the first time. In the 0.5-10 keV band, the source had been classified as an X-ray Binary candidate in various Chandra and XMM-Newton studies; however, since it was not clearly associated with Swift J0042.6+4112, the previous E<10 keV observations did not generate much attention. This source has a spectrum with a soft X-ray excess (kT ∼ 0.2 keV) plus a hard spectrum with a power law of G ~ 1 and a cutoff around 15-20 keV, typical of the spectral characteristics of accreting pulsars. Unfortunately, any potential pulsation was undetected in the NuSTAR data, possibly due to insufficient photon statistics. The existing deep HST images exclude high-mass (>3 solar mass) donors at the location of this source. The best interpretation for the nature of this source is an X-ray pulsar with an intermediate-mass (<3 solar mass) companion or a symbiotic X-ray binary. We discuss other possibilities in more detail.

Yukita, M.↗

Analyzing the impact of design factors on solar module thermomechanical durability using interpretable machine learning techniques

Solar modules in utility-scale systems are expected to maintain decades of lifetime to rival conventional energy sources. However, cyclic thermomechanical loading often degrades their long-term performance, highlighting the importance of effective design to mitigate thermal expansion mismatches between module materials. Given the complex composition of solar modules, isolating the impact of individual components on overall durability remains a challenging task. In this work, we analyze a comprehensive data set that comprises bill-of-materials (BOM) and thermal cycling power loss from 251 distinct module designs to identify the predominant design factors and their impacts on the thermomechanical durability of modules. The methodology of our analysis combines machine learning modeling (random forest) and Shapley additive explanation (SHAP) to correlate design factors with power loss and interpret the model’s decision-making. The interpretation reveals that silicon type (monocrystalline or polycrystalline), encapsulant thickness, busbar numbers, and wafer thickness predominantly influence the degradation. With lower power loss of around 0.6% on average in the SHAP analysis, monocrystalline cells present better durability than polycrystalline cells. This finding is further substantiated by statistical testing on our raw data set. The SHAP analysis also demonstrates that while thicker encapsulants lead to reduced power loss, further increasing their thickness over around 0.6 to 0.7 mm does not yield additional benefits, particularly for the front side one. In addition, other important BOM features such as the number of busbars are analyzed. This study provides a blueprint for utilizing explainable machine learning techniques in a complex material system and can potentially guide future research on optimizing the design of solar modules.

14 SOLAR ENERGY↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Magnetospheric chorus - Occurrence patterns and normalized frequency

Over 400 hours of continuous broadband data obtained by the OGO 3 satellite are analyzed to provide a statistically accurate description of band-limited (magnetospheric) chorus. Certain aspects of the chorus frequency distribution are interpreted in terms of a gyroresonant electron feedback model of generation. An example of high chorus activity during an outbound pass through the noon magnetosphere is examined in detail, the spectral complexity of some chorus is illustrated, and the diurnal variation of chorus occurrence is investigated. The frequency and bandwidth distributions of chorus are analyzed. The results indicate that chorus occurrence depends strongly on local time and dipole latitude, the general region of maximum chorus occurrence approximates the previously reported zone of 'hard' electron precipitation, and the normalized chorus frequency is strongly dependent on dipole latitude. It is shown how a change in the curvature of the whistler-mode refractive-index surface affects focusing of radiation along magnetic field lines and how interference can occur between modes with slightly different ray velocities. It is concluded that most magnetospheric chorus consists of rising emissions which are probably generated by gyroresonant electrons slightly off the equator.

Burtis, W. J.↗

Statistical Model Selection for TID Hardness Assurance

Radiation Hardness Assurance (RHA) methodologies against Total Ionizing Dose (TID) degradation impose rigorous statistical treatments for data from a part's Radiation Lot Acceptance Test (RLAT) and/or its historical performance. However, no similar methods exist for using "similarity" data - that is, data for similar parts fabricated in the same process as the part under qualification. This is despite the greater difficulty and potential risk in interpreting of similarity data. In this work, we develop methods to disentangle part-to-part, lot-to-lot and part-type-to-part-type variation. The methods we develop apply not just for qualification decisions, but also for quality control and detection of process changes and other "out-of-family" behavior. We begin by discussing the data used in ·the study and the challenges of developing a statistic providing a meaningful measure of degradation across multiple part types, each with its own performance specifications. We then develop analysis techniques and apply them to the different data sets.

Ladbury, R.↗

Predicting protein functions from redundancies in large-scale protein interaction networks

Interpreting data from large-scale protein interaction experiments has been a challenging task because of the widespread presence of random false positives. Here, we present a network-based statistical algorithm that overcomes this difficulty and allows us to derive functions of unannotated proteins from large-scale interaction data. Our algorithm uses the insight that if two proteins share significantly larger number of common interaction partners than random, they have close functional associations. Analysis of publicly available data from Saccharomyces cerevisiae reveals >2,800 reliable functional associations, 29% of which involve at least one unannotated protein. By further analyzing these associations, we derive tentative functions for 81 unannotated proteins with high certainty. Our method is not overly sensitive to the false positives present in the data. Even after adding 50% randomly generated interactions to the measured data set, we are able to recover almost all (approximately 89%) of the original associations.

Proteins/chemistry/metabolism↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

GeoDash: Assisting Visual Image Interpretation in Collect Earth Online by Leveraging Big Data on Google Earth Engine

Collect Earth Online (CEO) is a free and open online implementation of the FAO Collect Earth system for collaboratively collecting environmental data through the visual interpretation of Earth observation imagery. The primary collection mechanism in CEO is human interpretation of land surface characteristics in imagery served via Web Map Services (WMS). However, interpreters may not have enough contextual information to classify samples by only viewing the imagery served via WMS, be they high resolution or otherwise. To assist in the interpretation and collection processes in CEO, SERVIR, a joint NASA-USAID initiative that brings Earth observations to improve environmental decision making in developing countries, developed the GeoDash system, an embedded and critical component of CEO. GeoDash leverages Google Earth Engine (GEE) by allowing users to set up custom browser-based widgets that pull from GEE's massive public data catalog. These widgets can be quick looks of other satellite imagery, time series graphs of environmental variables, and statistics panels of the same. Users can customize widgets with any of GEE's image collections, such as the historical Landsat collection with data available since the 1970s, select date ranges, image stretch parameters, graph characteristics, and create custom layouts, all on-the-fly to support plot interpretation in CEO. This presentation focuses on the implementation and potential applications, including the back-end links to GEE and the user interface with custom widget building. GeoDash takes large data volumes and condenses them into meaningful, relevant information for interpreters. While designed initially with national and global forest resource assessments in mind, the system will complement disaster assessments, agriculture management, project monitoring and evaluation, and more.

SERVI↗

Evaluation of several classification schemes for mapping forest cover types in Michigan

Landsat MSS data were evaluated for mapping forest cover types in the northern Lower Peninsula of Michigan. The study examined seasonal variations, interpretation procedures and vegetation composition/distribution and their effect on overall classification accuracy and ability to identify individual pine species. Photographic images were used for visual interpretations while digital analysis was performed using a common (ERDAS) microcomputer image processing system. The classification schemes were evaluated using contingency tables and were ranked using the KAPPA statistic. The various classification schemes were ranked differentially according to study site location. Visual interpretation procedures ranked best, or least accurate, depending on the spatial distribution and complexity of the forest cover. Supervised classification techniques were more accurate than unsupervised clustering over all sites and seasons. Maximum likelihood classification of June data was superior to any digital classification technique of February data. The study indicates that classification accuracy is more dependent on the composition and distribution of forests in the northern lower Peninsula of Michigan than on the selection of a particular classification scheme.

Hudson, W. D.↗

A comparison of unsupervised classification procedures on LANDSAT MSS data for an area of complex surface conditions in Basilicata, Southern Italy

Two unsupervised classification procedures were applied to ratioed and unratioed LANDSAT multispectral scanner data of an area of spatially complex vegetation and terrain. An objective accuracy assessment was undertaken on each classification and comparison was made of the classification accuracies. The two unsupervised procedures use the same clustering algorithm. By on procedure the entire area is clustered and by the other a representative sample of the area is clustered and the resulting statistics are extrapolated to the remaining area using a maximum likelihood classifier. Explanation is given of the major steps in the classification procedures including image preprocessing; classification; interpretation of cluster classes; and accuracy assessment. Of the four classifications undertaken, the monocluster block approach on the unratioed data gave the highest accuracy of 80% for five coarse cover classes. This accuracy was increased to 84% by applying a 3 x 3 contextual filter to the classified image. A detailed description and partial explanation is provided for the major misclassification. The classification of the unratioed data produced higher percentage accuracies than for the ratioed data and the monocluster block approach gave higher accuracies than clustering the entire area. The moncluster block approach was additionally the most economical in terms of computing time.

Justice, C.↗

Characterization of Venera 15/16 Geologic Units from Pioneer Venus Reflectivity and Roughness Data

Geologic units have been defined for the surface of Venus from Venera 15/16 image data. A characterization of these geologic units is carried out using information on surface properties derived from Pioneer Venus (PV) reflectivity and rms slope data. The geologic context provided by Venera 15/16 units allows additional, more specific interpretations of surface radar properties to be made. Characterization of Venera units results in the definition of four groups of Venera units: (1) smooth rocky units, 2) rough rocky units, (3) rough high dielectric units, and (4) diffusely scattering units. On the basis of correlations of surface morphology to spatial and statistical distributions in rms slope and reflectivity data, we test models for the origin of the surface properties of some units. We conclude that plains and tectonic units can be contrasted in terms of the average roughness of the surface and that tectonic deformation appears to roughen the surface at 0.5- to 10-m and 5- to 50-cm scales. This tectonic weathering process appears to dominate the erosional regime of Venus. Unlike Earth or Mars, production and transport of soils dominates only a small portion (less than or equal to 5%) of the surface. Some of the Venera units display distinctive spatial and statistical distributions of PV radar data. In particular, apparent low reflectivity in the tesserae appears to be caused by small (5-50 cm) rock fragments on the surface which cause diffuse scattering at Pioneer Venus wavelengths. Analysis of models for the formation of these fragments suggests that they are due to the pervasive deformation undergone by the tesserae. Finally, aspects of this study have been used to extend results of Venera image data analysis southward of 30 deg. N lat, resulting in it prediction of the distribution of tessera. Such results can aid in Magellan investigations.

Bindschadler, D. L.↗

Polarization and the envelopes of B(e) supergiants in the Magellanic Clouds

We report optical linear polarization observations of nine B(e) supergiants in the Magellanic Clouds. Several of them have large intrinsic polarizations. The data are consistent with nonspherically symmetric envelopes, around the B(e) stars, that possess a range of intrinsic polarizations. Comparison of the polarimetric data with the viewing aspect, inferred from the spectroscopy of the individual objects, agrees with this interpretation. The polarization correlates best with the infrared excess due to dust, suggesting the latter as the cause of most of the polarization. We cannot find, using the available data, a statistically significant difference between the polarization distributions of Galactic and Magellanic B(e) stars. Comparison of our data with previous polarimetric data, recovered from the literature for four stars, seems to indicate that the envelopes are stable.

Magalhaes, A. M.↗

Strategies and Technologies for In Situ Mineralogical Investigations on Mars

Surface landers on Mars (Viking and Pathfinder) have not revealed satisfying answers to the mineralogy and lithology of the planet's surface. In part, this results from their prime directives: Viking focused on exobiology, Pathfinder focused on technology demonstration. The analytical instruments on board the landers made admirable attempts to extract the mineralogy and geology of Mars, as did countless modeling efforts after the missions. Here we suggest a framework for elucidating martian, or any other planetary geology, through an approach that defines (a) type of information required, (b) explorational strategy harmonious with acquisition of these data, (c) interpretation approach to the data, (d) compatible mission architecture, (e) instrumentation for interrogating rocks and soil. (a) Data required: The composition of a planet is ordered at scales ranging from molecules to minerals to rocks, and from geological units to provinces to planetary-scale systems. The largest ordering that in situ compositional instruments can attempt to interrogate is rock type "aggregate" information. This is what the geologist attempts to identify first. From this, mineralogy can be either directly seen or inferred. From mineralogy can be determined elemental abundances and perhaps the state of the compounds as being crystalline or amorphous. Knowledge of rock type and mineralogy is critical for elucidating geologic process. Mars landers acquired extremely valuable elemental data, but attempted to move from elements to aggregates, but this can only be done by making many assumptions and sometimes giant leaps of faith. Data we believe essential are elements, minerals, degree of ordering of compounds, and the aggregate or rock type that these materials compose. (b) Explorational strategy: A lander should function as a surrogate geologist. Of the total landscape, a geologist sees much, but gives detailed attention to an infinitesimally small amount of what is seen. To acquire samples worth detailed scrutiny, as many samples as possible need examining at a cursory or reconnaissance level. A representative, statistically-meaningful sample number cannot be overemphasized. This maxim still applies to geological exploration of our own planet of which we have abundant knowledge. Analysis of many samples mandates low-power consumption per sample. (c) Data interpretation: No single instrument can analyze the full spectrum of the x-axis. An instrument is optimized for detecting certain material characteristics and must therefore affix itself to some point on the x-axis. Any conclusions drawn about data to the left or right of the instrument's position on this axis must necessarily be derived by inference. Hence, it seems logical to include on a mission, instruments that are not closely spaced in their x-axis-position, and if only two analytical methods are used, as shown, they should start at opposite ends of the axis and work towards the center. As examples, we depict a high-resolution camera to evaluate rock type ("aggregate" state) and mineralogy, and an x-ray diffractometer-fluorescence spectrometer (XRD-XRF) to determine elements, minerals, and the degree of order of materials. (d) Mission architecture: No instrument or suite of instruments can be relied upon to always give truly unequivocal analyses. The suite of instruments should therefore permit conclusions of one instrument to be checked against those of another through closed analytical loops. These "loops" can be structured by a combination of orbital imagery, descent imagery, broad-band site viewing/analysis, and data that cover both x and y axes. For example, the detection of a basaltic-looking rock with a microscope should be checked against the elements detected, the appearance of the rock as a lava flow from descent imagery, and so forth. (e) Instrumentation: To satisfy the above criteria, it is necessary to: (i) See the rock or soil with high resolution + magnification, (ii) Examine many samples, (iii) Consume little power per analysis, (iv) Determine elemental species, (v) Determine mineralogy directly (not inferentially) and the degree of ordering of compounds, (vi) Start analyzing from both ends of the x-axis. Every geologist wants to see the hand sample first, and apply a hand lens to its surface. This has not been the starting point for missions to Mars. Thus, our technology satisfies all these criteria . This XRD-XRF-Optical instrument currently being developed, analyses rock or soil surfaces without the need for sample acquisition or preparation; this satisfies the power criterion, and enables many analyses. The device acquires direct mineralogy and determines elemental species. The embedded endoscopic camera satisfies the critical criterion of close inspection of samples; the fiber optic cable can also be used for IR, LTV, or laser sample analysis. Additional information is contained in the original (Figures).

Marshall, J. R.↗

Astronomical data analysis software and systems I; Proceedings of the 1st Annual Conference, Tucson, AZ, Nov. 6-8, 1991

Consideration is given to a definition of a distribution format for X-ray data, the Einstein on-line system, the NASA/IPAC extragalactic database, COBE astronomical databases, Cosmic Background Explorer astronomical databases, the ADAM software environment, the Groningen Image Processing System, search for a common data model for astronomical data analysis systems, deconvolution for real and synthetic apertures, pitfalls in image reconstruction, a direct method for spectral and image restoration, and a discription of a Poisson imagery super resolution algorithm. Also discussed are multivariate statistics on HI and IRAS images, a faint object classification using neural networks, a matched filter for improving SNR of radio maps, automated aperture photometry of CCD images, interactive graphics interpreter, the ROSAT extreme ultra-violet sky survey, a quantitative study of optimal extraction, an automated analysis of spectra, applications of synthetic photometry, an algorithm for extra-solar planet system detection and data reduction facilities for the William Herschel telescope.

Worrall, Diana M.↗

Using computers to analyze continuous data.

Dynamic field measurements often involve large quantities of continuous data, which must be analyzed and interpreted to obtain meaningful information. The processing can often be accomplished by tape-recording the data in analog form, performing off-line digitalization, and using the result as an input to statistical programs on a digital computer. A time series analysis was used to obtain power spectral density (PSD) curves to identify dominant frequencies. Representative PSD plots were obtained for STOL aircraft during cruise. Vibrational energy was clearly concentrated below 0.1 Hz, and was much higher in the vertical than in the lateral direction.

Catherines, J. J.↗

Filamentary galaxy clustering - A mapping algorithm

A simple and objective algorithm is presented which not only accurately identifies the filamentary structures in the Shane-Wirtanen galaxy count catalog, but also finds a set of visually less impressive filaments in a static hierarchical model of the clustering conducted by Soneira and Peebles (1978). The statistical properties of the elements in the model, while very similar to those in the data, show a significant excess of long and bright filaments in the data relative to the model. Two possible interpretations of these results are presented and discussed.

Gott, J. R., III↗

Remote sensing of Earth terrain

Remote sensing of earth terrain is examined. The layered random medium model is used to investigate the fully polarimetric scattering of electromagnetic waves from vegetation. The model is used to interpret the measured data for vegetation fields such as rice, wheat, or soybean over water or soil. Accurate calibration of polarimetric radar systems is essential for the polarimetric remote sensing of earth terrain. A polarimetric calibration algorithm using three arbitrary in-scene reflectors is developed. In the interpretation of active and passive microwave remote sensing data from the earth terrain, the random medium model was shown to be quite successful. A multivariate K-distribution is proposed to model the statistics of fully polarimetric radar returns from earth terrain. In the terrain cover classification using the synthetic aperture radar (SAR) images, the applications of the K-distribution model will provide better performance than the conventional Gaussian classifiers. The layered random medium model is used to study the polarimetric response of sea ice. Supervised and unsupervised classification procedures are also developed and applied to synthetic aperture radar polarimetric images in order to identify their various earth terrain components for more than two classes. These classification procedures were applied to San Francisco Bay and Traverse City SAR images.

Kong, Jin AU↗

PARAGON: A Systematic, Integrated Approach to Aerosol Observation and Modeling

Aerosols are generated and transformed by myriad processes operating across many spatial and temporal scales. Evaluation of climate models and their sensitivity to changes, such as in greenhouse gas abundances, requires quantifying natural and anthropogenic aerosol forcings and accounting for other critical factors, such as cloud feedbacks. High accuracy is required to provide sufficient sensitivity to perturbations, separate anthropogenic from natural influences, and develop confidence in inputs used to support policy decisions. Although many relevant data sources exist, the aerosol research community does not currently have the means to combine these diverse inputs into an integrated data set for maximum scientific benefit. Bridging observational gaps, adapting to evolving measurements, and establishing rigorous protocols for evaluating models are necessary, while simultaneously maintaining consistent, well understood accuracies. The Progressive Aerosol Retrieval and Assimilation Global Observing Network (PARAGON) concept represents a systematic, integrated approach to global aerosol Characterization, bringing together modern measurement and modeling techniques, geospatial statistics methodologies, and high-performance information technologies to provide the machinery necessary for achieving a comprehensive understanding of how aerosol physical, chemical, and radiative processes impact the Earth system. We outline a framework for integrating and interpreting observations and models and establishing an accurate, consistent and cohesive long-term data record.

PARAGON↗