Search NASASearch

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Bayesian Approach to the Joint Inversion of Gravity and Magnetic Data, with Application to the Ismenius Area of Mars

This viewgraph presentation reviews a Bayesian approach to the inversion of gravity and magnetic data with specific application to the Ismenius Area of Mars. Many inverse problems encountered in geophysics and planetary science are well known to be non-unique (i.e. inversion of gravity the density structure of a body). In hopes of reducing the non-uniqueness of solutions, there has been interest in the joint analysis of data. An example is the joint inversion of gravity and magnetic data, with the assumption that the same physical anomalies generate both the observed magnetic and gravitational anomalies. In this talk, we formulate the joint analysis of different types of data in a Bayesian framework and apply the formalism to the inference of the density and remanent magnetization structure for a local region in the Ismenius area of Mars. The Bayesian approach allows prior information or constraints in the solutions to be incorporated in the inversion, with the "best" solutions those whose forward predictions most closely match the data while remaining consistent with assumed constraints. The application of this framework to the inversion of gravity and magnetic data on Mars reveals two typical challenges - the forward predictions of the data have a linear dependence on some of the quantities of interest, and non-linear dependence on others (termed the "linear" and "non-linear" variables, respectively). For observations with Gaussian noise, a Bayesian approach to inversion for "linear" variables reduces to a linear filtering problem, with an explicitly computable "error" matrix. However, for models whose forward predictions have non-linear dependencies, inference is no longer given by such a simple linear problem, and moreover, the uncertainty in the solution is no longer completely specified by a computable "error matrix". It is therefore important to develop methods for sampling from the full Bayesian posterior to provide a complete and statistically consistent picture of model uncertainty, and what has been learned from observations. We will discuss advanced numerical techniques, including Monte Carlo Markov

data analysis

Evaluation of Classifier Complexity for Delay Tolerant Network Routing

The growing popularity of small cost effective satellites (SmallSats, CubeSats, etc.) creates the potential for a variety of new science applications involving multiple nodes functioning together or independently to achieve a task, such as swarms and constellations. As this technology develops and is deployed for missions in Low Earth Orbit and beyond, the use of delay tolerant networking (DTN) techniques may improve communication capabilities within the network. In this paper, a network hierarchy is developed from heterogeneous networks of SmallSats, surface vehicles, relay satellites and ground stations which form an integrated network. There is a tradeoff between complexity, flexibility, and scalability of user defined schedules versus autonomous routing as the number of nodes in the network increases. To address these issues, this work proposes a machine learning classifier based on DTN routing metrics. A framework is developed which will allow for the use of several categories of machine learning algorithms (decision tree, random forest and deep learning) to be applied to a dataset of historical network statistics, which allows for the evaluation of algorithm complexity versus performance to be explored. We develop the emulation of a hierarchical network, consisting of tens of nodes which form a cognitive network architecture. CORE (Common Open Research Emulator) is used to emulate the network using bundle protocol and DTN IP neighbor discovery.

Dudukovich, Rachel

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology

NASA Tech Briefs, June 2014

Topics include: Real-Time Minimization of Tracking Error for Aircraft Systems; Detecting an Extreme Minority Class in Hyperspectral Data Using Machine Learning; KSC Spaceport Weather Data Archive; Visualizing Acquisition, Processing, and Network Statistics Through Database Queries; Simulating Data Flow via Multiple Secure Connections; Systems and Services for Near-Real-Time Web Access to NPP Data; CCSDS Telemetry Decoder VHDL Core; Thermal Response of a High-Power Switch to Short Pulses; Solar Panel and System Design to Reduce Heating and Optimize Corridors for Lower-Risk Planetary Aerobraking; Low-Cost, Very Large Diamond-Turned Metal Mirror; Very-High-Load-Capacity Air Bearing Spindle for Large Diamond Turning Machines; Elevated-Temperature, Highly Emissive Coating for Energy Dissipation of Large Surfaces; Catalyst for Treatment and Control of Post-Combustion Emissions; Thermally Activated Crack Healing Mechanism for Metallic Materials; Subsurface Imaging of Nanocomposites; Self-Healing Glass Sealants for Solid Oxide Fuel Cells and Electrolyzer Cells; Micromachined Thermopile Arrays with Novel Thermo - electric Materials; Low-Cost, High-Performance MMOD Shielding; Head-Mounted Display Latency Measurement Rig; Workspace-Safe Operation of a Force- or Impedance-Controlled Robot; Cryogenic Mixing Pump with No Moving Parts; Seal Design Feature for Redundancy Verification; Dexterous Humanoid Robot; Tethered Vehicle Control and Tracking System; Lunar Organic Waste Reformer; Digital Laser Frequency Stabilization via Cavity Locking Employing Low-Frequency Direct Modulation; Deep UV Discharge Lamps in Capillary Quartz Tubes with Light Output Coupled to an Optical Fiber; Speech Acquisition and Automatic Speech Recognition for Integrated Spacesuit Audio Systems, Version II; Advanced Sensor Technology for Algal Biotechnology; High-Speed Spectral Mapper; "Ascent - Commemorating Shuttle" - A NASA Film and Multimedia Project DVD; High-Pressure, Reduced-Kinetics Mechanism for N-Hexadecane Oxidation; Method of Error Floor Mitigation in Low-Density Parity-Check Codes; X-Ray Flaw Size Parameter for POD Studies; Large Eddy Simulation Composition Equations for Two-Phase Fully Multicomponent Turbulent Flows; Scheduling Targeted and Mapping Observations with State, Resource, and Timing Constraints;

Source record

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term. The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission (TRMM) Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission TRMM Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters

Neural network representation and learning of mappings and their derivatives

Discussed here are recent theorems proving that artificial neural networks are capable of approximating an arbitrary mapping and its derivatives as accurately as desired. This fact forms the basis for further results establishing the learnability of the desired approximations, using results from non-parametric statistics. These results have potential applications in robotics, chaotic dynamics, control, and sensitivity analysis. An example involving learning the transfer function and its derivatives for a chaotic map is discussed.

White, Halbert

Avalanches and the distribution of solar flares

The solar coronal magnetic field is proposed to be in a self-organized critical state, thus explaining the observed power-law dependence of solar-flare-occurrence rate on flare size which extends over more than five orders of magnitude in peak flux. The physical picture that arises is that solar flares are avalanches of many small reconnection events, analogous to avalanches of sand in the models published by Bak and colleagues in 1987 and 1988. Flares of all sizes are manifestations of the same physical processes, where the size of a given flare is determined by the number of elementary reconnection events. The relation between small-scale processes and the statistics of global-flare properties which follows from the self-organized magnetic-field configuration provides a way to learn about the physics of the unobservable small-scale reconnection processes. A simple lattice-reconnection model is presented which is consistent with the observed flare statistics. The implications for coronal heating are discussed and some observational tests of this picture are given.

Lu, Edward T.

Active Learning with Rationales for Identifying Operationally Significant Anomalies in Aviation

A major focus of the commercial aviation community is discovery of unknown safety events in flight operations data. Data-driven unsupervised anomaly detection methods are better at capturing unknown safety events compared to rule-based methods which only look for known violations. However, not all statistical anomalies that are discovered by these unsupervised anomaly detection methods are operationally significant (e.g., represent a safety concern). Subject Matter Experts (SMEs) have to spend significant time reviewing these statistical anomalies individually to identify a few operationally significant ones. In this paper we propose an active learning algorithm that incorporates SME feedback in the form of rationales to build a classifier that can distinguish between uninteresting and operationally significant anomalies. Experimental evaluation on real aviation data shows that our approach improves detection of operationally significant events by as much as 75% compared to the state-of-the-art. The learnt classifier also generalizes well to additional validation data sets.

anomaly detection

Bringing Climate Scientist's Tools into Classrooms to Improve Conceptual Understandings

Efforts to address anthropogenic global climate change (AGCC) require public understanding of Earth and climate science. To meet this need, educational reforms and prominent scientists have called for instructional approaches that teach students how climate scientists examine AGCC. Yet, only a few educational studies have reported clear empirical results on what instructional approaches and climate education technologies best accomplish this goal. This manuscript presents detailed analysis and statistically significant results on the educational impact pre to post of students learning to use a National Aeronautics and Space Administration (NASA) global climate model (GCM). This series of case studies demonstrates that differing instructional approaches and climate education technologies result in differing levels of understanding of AGCC and ability to engage with policies addressing it. Students who learned the scientific process of climate modeling scored significantly higher pre to post on exams (quantitatively) and gained more complete conceptual understandings of the issue (qualitatively). Yet, teaching students to conduct research with complex technology can be difficult. This study also found lecture-based learning better improved recall of facts about GCMs tested by multiple-choice questions. Our findings indicate what educational systems and related technologies might provide the public with the conceptual understandings necessary to engage in the political debate over AGCC.

Public understanding

Metamodels for Computer-Based Engineering Design: Survey and Recommendations

The use of statistical techniques to build approximations of expensive computer analysis codes pervades much of todays engineering design. These statistical approximations, or metamodels, are used to replace the actual expensive computer analyses, facilitating multidisciplinary, multiobjective optimization and concept exploration. In this paper we review several of these techniques including design of experiments, response surface methodology, Taguchi methods, neural networks, inductive learning, and kriging. We survey their existing application in engineering design and then address the dangers of applying traditional statistical techniques to approximate deterministic computer analysis codes. We conclude with recommendations for the appropriate use of statistical approximation techniques in given situations and how common pitfalls can be avoided.

Simpson, Timothy W.

Putting Priors in Mixture Density Mercer Kernels

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. We describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using predefined kernels. These data adaptive kernels can en- code prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS). The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains template for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic- algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code. The results show that the Mixture Density Mercer-Kernel described here outperforms tree-based classification in distinguishing high-redshift galaxies from low- redshift galaxies by approximately 16% on test data, bagged trees by approximately 7%, and bagged trees built on a much larger sample of data by approximately 2%.

Srivastava, Ashok N.

Statistical Practice and Research at NASA

The discipline of statistics has gained recognition within NASA by spurring innovation and efficiency, and it has demonstrated significant impact and value. In aerospace research and development, it accelerates learning, maximizes knowledge, ensures strategic resource investment, and informs rigorous data-driven decisions. In practice, it requires immersive multidisciplinary collaboration to develop solution strategies that integrate statistical methods with subject-matter expertise to address challenging research objectives. This presentation provides an overview of statistical case studies in aeronautics, space exploration, and atmospheric science, and it highlights statistical research motivated by NASA’s challenging applications.

Peter A. Parker

An Ensemble Approach to Building Mercer Kernels with Prior Information

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly dimensional feature space. we describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using pre-defined kernels. These data adaptive kernels can encode prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. Specifically, we demonstrate the use of the algorithm in situations with extremely small samples of data. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS) and demonstrate the method's superior performance against standard methods. The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains templates for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic-algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code.

Srivastava, Ashok N.

Experiences in the Practice of Design of Experiments at NASA

Statistical design of experiments (DOE) has gained recognition within NASA by spurring innovation and efficiency, and it has demonstrated significant impact. Aerospace research and development benefits DOE techniques by accelerating learning, maximizing knowledge, ensuring strategic resource investment, and informing data-driven decisions. In practice, DOE relies on multidisciplinary collaboration to develop solution strategies that integrate statistical methods with subject-matter expertise to meet challenging research objectives. This presentation shares experiences in the practice of design of experiments at NASA in aeronautics, space exploration, and atmospheric science.

Peter A. Parker

A petabyte size electronic library using the N-Gram memory engine

A model library containing petabytes of data is proposed by Triada, Ltd., Ann Arbor, Michigan. The library uses the newly patented N-Gram Memory Engine (Neurex), for storage, compression, and retrieval. Neurex splits data into two parts: a hierarchical network of associative memories that store 'information' from data and a permutation operator that preserves sequence. Neurex is expected to offer four advantages in mass storage systems. Neurex representations are dense, fully reversible, hence less expensive to store. Neurex becomes exponentially more stable with increasing data flow; thus its contents and the inverting algorithm may be mass produced for low cost distribution. Only a small permutation operator would be recalled from the library to recover data. Neurex may be enhanced to recall patterns using a partial pattern. Neurex nodes are measures of their pattern. Researchers might use nodes in statistical models to avoid costly sorting and counting procedures. Neurex subsumes a theory of learning and memory that the author believes extends information theory. Its first axiom is a symmetry principle: learning creates memory and memory evidences learning. The theory treats an information store that evolves from a null state to stationarity. A Neurex extracts information data without a priori knowledge; i.e., unlike neural networks, neither feedback nor training is required. The model consists of an energetically conservative field of uniformly distributed events with variable spatial and temporal scale, and an observer walking randomly through this field. A bank of band limited transducers (an 'eye'), each transducer in a bank being tuned to a sub-band, outputs signals upon registering events. Output signals are 'observed' by another transducer bank (a mid-brain), except the band limit of the second bank is narrower than the band limit of the first bank. The banks are arrayed as n 'levels' or 'time domains, td.' The banks are the hierarchical network (a cortex) and transducers are (associative) memories. A model Neurex was built and studied. Data were 50 MB to 10 GB samples of text, data base, and images: black/white, grey scale, and high resolution in several spectral bands. Memories at td, S(m(sub td)), were plotted against outputs of memories at td-1. S(m(sub td)) was Boltzman distributed, and memory frequencies exhibited self-organized criticality (SOC); i.e., 'l/f(sup beta)' after long exposures to data. Whereas output signals from level n may be encoded with B(sub output) = O(-log(2)f(sup beta)) bits, and input data encoded with B(sub input) = O((S(td)/S(td-1))(sup n)), B(sup output)/B(sub input) is much less than 1 always, the Neurex determines a canonical code for data and it is a lossless data compressor. Further tests are underway to confirm these results with more data types and larger samples.

Bugajski, Joseph M.

Darwinian Spacecraft: Soft Computing Strategies Breeding Better, Faster Cheaper

Computers can create infinite lists of combinations to try to solve a particular problem, a process called "soft-computing." This process uses statistical comparables, neural networks, genetic algorithms, fuzzy variables in uncertain environments, and flexible machine learning to create a system which will allow spacecraft to increase robustness, and metric evaluation. These concepts will allow for the development of a spacecraft which will allow missions to be performed at lower costs.

Noever, David A.