Search NASASearch

SEARCH · Search NASA

Results for “multiple instance regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Multiple-Instance Regression with Structured Data

We present a multiple-instance regression algorithm that models internal bag structure to identify the items most relevant to the bag labels. Multiple-instance regression (MIR) operates on a set of bags with real-valued labels, each containing a set of unlabeled items, in which the relevance of each item to its bag label is unknown. The goal is to predict the labels of new bags from their contents. Unlike previous MIR methods, MI-ClusterRegress can operate on bags that are structured in that they contain items drawn from a number of distinct (but unknown) distributions. MI-ClusterRegress simultaneously learns a model of the bag's internal structure, the relevance of each item, and a regression model that accurately predicts labels for new bags. We evaluated this approach on the challenging MIR problem of crop yield prediction from remote sensing data. MI-ClusterRegress provided predictions that were more accurate than those obtained with non-multiple-instance approaches or MIR methods that do not model the bag structure.

learning

Multiple Instance Regression with Structured Data

This slide presentation reviews the use of multiple instance regression with structured data from multiple and related data sets. It applies the concept to a practical problem, that of estimating crop yield using remote sensed country wide weekly observations.

multiple instance regression

Salience Assignment for Multiple-Instance Regression

We present a Multiple-Instance Learning (MIL) algorithm for determining the salience of each item in each bag with respect to the bag's real-valued label. We use an alternating-projections constrained optimization approach to simultaneously learn a regression model and estimate all salience values. We evaluate this algorithm on a significant real-world problem, crop yield modeling, and demonstrate that it provides more extensive, intuitive, and stable salience models than Primary-Instance Regression, which selects a single relevant item from each bag.

regression

Tuning Parameters in Heuristics by Using Design of Experiments Methods

With the growing complexity of today's large scale problems, it has become more difficult to find optimal solutions by using exact mathematical methods. The need to find near-optimal solutions in an acceptable time frame requires heuristic approaches. In many cases, however, most heuristics have several parameters that need to be "tuned" before they can reach good results. The problem then turns into "finding best parameter setting" for the heuristics to solve the problems efficiently and timely. One-Factor-At-a-Time (OFAT) approach for parameter tuning neglects the interactions between parameters. Design of Experiments (DOE) tools can be instead employed to tune the parameters more effectively. In this paper, we seek the best parameter setting for a Genetic Algorithm (GA) to solve the single machine total weighted tardiness problem in which n jobs must be scheduled on a single machine without preemption, and the objective is to minimize the total weighted tardiness. Benchmark instances for the problem are available in the literature. To fine tune the GA parameters in the most efficient way, we compare multiple DOE models including 2-level (2k ) full factorial design, orthogonal array design, central composite design, D-optimal design and signal-to-noise (SIN) ratios. In each DOE method, a mathematical model is created using regression analysis, and solved to obtain the best parameter setting. After verification runs using the tuned parameter setting, the preliminary results for optimal solutions of multiple instances were found efficiently.

Arin, Arif

Utilization of Machine Learning Techniques for Managing the Tracking and Data Relay Satellite Constellation

National Aeronautics and Space Administration’s (NASA) Goddard Space Flight Center (GSFC) operates a constellation of ten geosynchronous Tracking and Data Relay Satellites (TDRS). The TDRS constellation consists of multiple geosynchronous communication relay satellites located around the equator so they can provide continual coverage of any mission in low earth orbit. The TDRS are located primarily in three oceanic regions around the earth. NASA’s White Sands Complex provides the ground communication support for TDRS located over the Atlantic and Pacific Oceans. Another TDRS ground station in Guam supports the TDRS over the Indian Ocean. With these satellites the TDRS network can provide continuous coverage of satellites in low-earth orbit. The NASA Space Network (SN) project office at GSFC manages the constellation of spacecraft. Major customers of the TDRS constellation include, but are not limited to, the International Space Station and the Hubble Space Telescope. The TDRS constellation has three generations of satellites and has been active for over 30 years providing reliable communication links between customer satellites and corresponding ground stations. However, one of the major concerns for TDRS, and in any space mission, is to ensure the health and safety of the spacecraft. Generally, engineers use telemetry data to monitor and analyze the performance and state of health of the spacecraft. Telemetry data contains hundreds of parameters that monitor each important component in the spacecraft, which can be utilized to recognize and characterize the behavior of the spacecraft. Each parameter contains considerable information to represent time-dependent properties of each spacecraft subsystem and component. During the entire life of a TDRS spacecraft, thousands of gigabytes of telemetry data are transmitted in real-time from the spacecraft to the ground station at the White Sands Complex in Las Cruces, New Mexico, and recorded as historical data sets for engineers to process and analyze the events that occurred on-orbit. These parameters contain the function of multiple spacecraft subsystems, such as the attitude control system (ACS), Thermal, Electrical Power Subsystem (EPS), etc. . The first and second generations have exceeded their required lifetime and NASA is keen to manage these spacecrafts carefully in order to maximize the remaining life using the spacecraft telemetry. The challenge is to know when the risk of losing a spacecraft in geosynchronous orbit exceeds the benefit of continued operations for customer support. In the TDRS fleet, the EPS is the most critical subsystem related to spacecraft operations. Failure of the EPS would strand a spacecraft in geosynchronous orbit. Since EPS provides power to the spacecraft, component failures ultimately lead to the inability to support the spacecraft loads and the communications payload. For instance, TDRS-8 has several anomalies in EPS including the Bus Voltage Limiter (BVL) shunt current, solar array loss of circuits, and failed battery cells. Any of these anomalies can cause critical issues to the spacecraft. Therefore, developing a system to analyze and perform early detection of a potential anomaly is an important issue in telemetry data analysis. In recent years, Telemetry Mining (TM) has been proposed to process telemetry data by using Data Mining (DM) techniques such as classification, clustering, regression and anomaly detection. Anomaly detection, also known as outlier detection, has been widely used in many data mining areas such as remote sensing, medical data processing and digital image processing. The goal of anomaly detection is to detect abnormal data, which contains a relatively low probability of occurrence among the entire data set. Early detection of anomalies is one of the most significant issues in managing the spacecraft configuration. If anomalies can be detected early enough, then the redundant resources can be used to extend the life of the operational spacecraft. We present an unsupervised anomaly detection method to process the EPS data extracted from TDRS-8. This is different from traditional analytical methods, which use telemetry data to illustrate behavior and physical meaning of each spacecraft component. TM connects multiple parameters as a vector and then conducts data analysis on this high dimension telemetry vector. This method is looking at the properties of a high dimensional vector that is able to consider the relationship between different parameters in the anomaly detection problem. This kind of method performs much better than the traditional limit checking method. In addition, we propose a new approach of real-time anomaly detection to process telemetry data in real-time, which can then be applied to spacecraft monitoring with high reliability, low cost and high accuracy.

Machine Learning (ML)

Comparison of Likelihood Methods for Generalized Linear Mixed Models with Application to Quiet Supersonic Flights 2018 Data

Repeated measurement will be a feature of the survey data collected during the Quesst missionX-59 community response tests (CRT). Since each participant will report his or her categorical level of annoyance in response to multiple events, the responses from any single individual may be correlated with one another. Several models within the class of generalized linear mixed models (GLMM) are pertinent to the analysis of correlated categorical outcomes; the random intercept logistic regression model is one example. Both Bayesian and frequentist methods for fitting these models are available, with frequentist methods relying on some form of approximation (of either an integral or the integrand) that appears in the marginal likelihood function. Given several anticipated similarities of the X-59 CRT data to data collected during a past risk reduction, Quiet Supersonic Flights 2018 (QSF18), this short note is intended to create awareness. It documents an instance in which a reported population average dose-response relationship derived from QSF18 single event data was distorted by the integral approximation applied in likelihood-based methods. We review some of the available literature on the topic, compare the outputs of several different computational approaches implemented in available statistical software, and present simple corrective actions that may be useful during the Quesst mission.

dose-response model

Salience Assignment for Multiple-Instance Data and Its Application to Crop Yield Prediction

An algorithm was developed to generate crop yield predictions from orbital remote sensing observations, by analyzing thousands of pixels per county and the associated historical crop yield data for those counties. The algorithm determines which pixels contain which crop. Since each known yield value is associated with thousands of individual pixels, this is a multiple instance learning problem. Because individual crop growth is related to the resulting yield, this relationship has been leveraged to identify pixels that are individually related to corn, wheat, cotton, and soybean yield. Those that have the strongest relationship to a given crop s yield values are most likely to contain fields with that crop. Remote sensing time series data (a new observation every 8 days) was examined for each pixel, which contains information for that pixel s growth curve, peak greenness, and other relevant features. An alternating-projection (AP) technique was used to first estimate the "salience" of each pixel, with respect to the given target (crop yield), and then those estimates were used to build a regression model that relates input data (remote sensing observations) to the target. This is achieved by constructing an exemplar for each crop in each county that is a weighted average of all the pixels within the county; the pixels are weighted according to the salience values. The new regression model estimate then informs the next estimate of the salience values. By iterating between these two steps, the algorithm converges to a stable estimate of both the salience of each pixel and the regression model. The salience values indicate which pixels are most relevant to each crop under consideration.

Wagstaff, Kiri L.

Methane as a Diagnostic Tracer of Changes in the Brewer-Dobson Circulation of the Stratosphere

This study makes use of time series of methane (CH4/ data from the Halogen Occultation Experiment (HALOE) to detect whether there were any statistically significant changes of the Brewer-Dobson circulation (BDC) within the stratosphere during 1992-2005. The HALOE CH4 profiles are in terms of mixing ratio versus pressure altitude and are binned into latitude zones within the Southern Hemisphere and the Northern Hemisphere. Their separate time series are then analyzed using multiple linear regression (MLR) techniques. The CH4 trend terms for the Northern Hemisphere are significant and positive at 10 N from 50 to 7 hPa and larger than the tropospheric CH4 trends of about 3%decade(exp -1) from 20 to 7 hPa. At 60 N the trends are clearly negative from 20 to 7 hPa. Their combined trends indicate an acceleration of the BDC in the middle stratosphere of the Northern Hemisphere during those years, most likely due to changes from the effects of wave activity. No similar significant BDC acceleration is found for the Southern Hemisphere. Trends from HALOE H2O are analyzed for consistency. Their mutual trends with CH4 are anti-correlated qualitatively in the middle and upper stratosphere, where CH4 is chemically oxidized to H2O. Conversely, their mutual trends in the lower stratosphere are dominated by their trends upon entry to the tropical stratosphere. Time series residuals for CH4 in the lower mesosphere also exhibit structures that are anti-correlated in some instances with those of the tracer-like species HCl. Their occasional aperiodic structures indicate the effects of transport following episodic, wintertime wave activity. It is concluded that observed multi-year, zonally averaged distributions of CH4 can be used to diagnose major instances of wave-induced transport in the middle atmosphere and to detect changes in the stratospheric BDC.

Remsberg, E. E.

Organometallic Polymeric Conductors

For aerospace applications, the use of polymers can result in tremendous weight savings over metals. Suitable polymeric materials for some applications like EMI shielding, spacecraft grounding, and charge dissipation must combine high electrical conductivity with long-term environmental stability, good processability, and good mechanical properties. Recently, other investigators have reported hybrid films made from an electrically conductive polymer combined with insulating polymers. In all of these instances, the films were prepared by infiltrating an insulating polymer with a precursor for a conductive polymer (either polypyrrole or polythiophene), and oxidatively polymerizing the precursor in situ. The resulting composite films have good electrical conductivity, while overcoming the brittleness inherent in most conductive polymers. Many aerospace applications require a combination of properties. Thus, hybrid films made from polyimides or other engineering resins are of primary interest, but only if conductivities on the same order as those obtained with a polystyrene base could be obtained. Hence, a series of experiments was performed to optimize the conductivity of polyimide-based composite films. The polyimide base chosen for this study was Kapton. 3-MethylThiophene (3MT) was used for the conductive phase. Three processing variables were identified for producing these composite films, namely time, temperature, and oxidant concentration for the in situ oxidation. Statistically designed experiments were used to examine the effects of these variables and synergistic/interactive effects among variables on the electrical conductivity and mechanical strength of the films. Multiple linear regression analysis of the tensile data revealed that temperature and time have the greatest effect on maximum stress. The response surface of maximum stress vs. temperature and time (for oxidant concentration at 1.2 M) is shown. Conductivity of the composite films was measured for over 150 days in air at ambient temperature. The conductivity of the films dropped only half an order of magnitude in that time. Films aged under vacuum at ambient temperature diminished slightly in conductivity in the first day, but did not change thereafter. An experimental design approach will be applied to maximize the efficiency of the laboratory effort. The material properties (initial and long term) will also be monitored and assessed. The experimental results will add to the existing database for electrically conductive polymer materials. Attachments: 1) Synthesis Crystal Structure, and Polymerization of 1,2:5,6:9,10-Tribenzo-3,7,11,13-tetradehydro(14) annulene. 2) Reinvestigation of the Photocyclization of 1,4-Phenylene Bis(phenylmaleic anhydride): Preparation and Structure of (5)Helicene 5,6:9,10-Dianhydride. 3) Preparation and Structure Charecterization of a Platinum Catecholate Complex Containing Two 3-Ethynyltheophone Groups. and 4) Rigid-Rod Polymers Based on Noncoplanar 4,4'-Biphenyldiamines: A Review of Polymer Properties vs Configuration of Diamines.

Youngs, Wiley J.

Modeling Insights into Deuterium Excess as an Indicator of Water Vapor Source Conditions

Deuterium excess (d) is interpreted in conventional paleoclimate reconstructions as a tracer of oceanic source region conditions, such as temperature, where precipitation originates. Previous studies have adopted co-isotopic approaches to estimate past changes in both site and oceanic source temperatures for ice core sites using empirical relationships derived from conceptual distillation models, particularly Mixed Cloud Isotopic Models (MCIMs). However, the relationship between d and oceanic surface conditions remains unclear in past contexts. We investigate this climate-isotope relationship for sites in Greenland and Antarctica using multiple simulations of the water isotope-enabled Goddard Institute for Space Studies (GISS) ModelE-R general circulation model and apply a novel suite of model vapor source distribution (VSD) tracers to assess d as a proxy for source temperature variability under a range of climatic conditions. Simulated average source temperatures determined by the VSDs are compared to synthetic source temperature estimates calculated using MCIM equations linking d to source region conditions. We show that although deuterium excess is generally a faithful tracer of source temperatures as estimated by the MCIM approach, large discrepancies in the isotope-climate relationship occur around Greenland during the Last Glacial Maximum simulation, when precipitation seasonality and moisture source regions were notably different from present. This identified sensitivity in d as a source temperature proxy suggests that quantitative climate reconstructions from deuterium excess should be treated with caution for some sites when boundary conditions are significantly different from the present day. Also, the exclusion of the influence of humidity and other evaporative source changes in MCIM regressions may be a limitation of quantifying source temperature fluctuations from deuterium excess in some instances.

simulation