Search NASA⌕ Search

SEARCH · Search NASA

Results for “false negative”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Abort Trigger False Positive and False Negative Analysis Methodology for Threshold-Based Abort Detection

This paper describes a quantitative methodology for bounding the false positive (FP) and false negative (FN) probabilities associated with a human-rated launch vehicle abort trigger (AT) that includes sensor data qualification (SDQ). In this context, an AT is a hardware and software mechanism designed to detect the existence of a specific abort condition. Also, SDQ is an algorithmic approach used to identify sensor data suspected of being corrupt so that suspect data does not adversely affect an AT's detection capability. The FP and FN methodologies presented here were developed to support estimation of the probabilities of loss of crew and loss of mission for the Space Launch System (SLS) which is being developed by the National Aeronautics and Space Administration (NASA). The paper provides a brief overview of system health management as being an extension of control theory; and describes how ATs and the calculation of FP and FN probabilities relate to this theory. The discussion leads to a detailed presentation of the FP and FN methodology and an example showing how the FP and FN calculations are performed. This detailed presentation includes a methodology for calculating the change in FP and FN probabilities that result from including SDQ in the AT architecture. To avoid proprietary and sensitive data issues, the example incorporates a mixture of open literature and fictitious reliability data. Results presented in the paper demonstrate the effectiveness of the approach in providing quantitative estimates that bound the probability of a FP or FN abort determination.

risk assessment↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Validation of Atmosphere/Ionosphere Signals Associated with Major Earthquakes by Multi-Instrument Space-Borne and Ground Observations

The latest catastrophic earthquake in Japan (March 2011) has renewed interest in the important question of the existence of pre-earthquake anomalous signals related to strong earthquakes. Recent studies have shown that there were precursory atmospheric/ionospheric signals observed in space associated with major earthquakes. The critical question, still widely debated in the scientific community, is whether such ionospheric/atmospheric signals systematically precede large earthquakes. To address this problem we have started to investigate anomalous ionospheric / atmospheric signals occurring prior to large earthquakes. We are studying the Earth's atmospheric electromagnetic environment by developing a multisensor model for monitoring the signals related to active tectonic faulting and earthquake processes. The integrated satellite and terrestrial framework (ISTF) is our method for validation and is based on a joint analysis of several physical and environmental parameters (thermal infrared radiation, electron concentration in the ionosphere, lineament analysis, radon/ion activities, air temperature and seismicity) that were found to be associated with earthquakes. A physical link between these parameters and earthquake processes has been provided by the recent version of Lithosphere-Atmosphere-Ionosphere Coupling (LAIC) model. Our experimental measurements have supported the new theoretical estimates of LAIC hypothesis for an increase in the surface latent heat flux, integrated variability of outgoing long wave radiation (OLR) and anomalous variations of the total electron content (TEC) registered over the epicenters. Some of the major earthquakes are accompanied by an intensification of gas migration to the surface, thermodynamic and hydrodynamic processes of transformation of latent heat into thermal energy and with vertical transport of charged aerosols in the lower atmosphere. These processes lead to the generation of external electric currents in specific regions of the atmosphere and the modifications, by dc electric fields, in the ionosphere-atmosphere electric circuit. We retrospectively analyzed temporal and spatial variations of four different physical parameters (gas/radon counting rate, lineaments change, long-wave radiation transitions and ionospheric electron density/plasma variations) characterizing the state of the lithosphere/atmosphere coupling several days before the onset of the earthquakes. Validation processes consist in two phases: A. Case studies for seven recent major earthquakes: Japan (M9.0, 2011), China (M7.9, 2008), Italy (M6.3, 2009), Samoa (M7, 2009), Haiti (M7.0, 2010) and, Chile (M8.8, 2010) and B. A continuous retrospective analysis was preformed over two different regions with high seismicity- Taiwan and Japan for 2003-2009. Satellite, ground surface, and troposphere data were obtained from Terra/ASTER, Aqua/AIRS, POES and ionospheric variations from DEMETER and COSMIC-I data. Radon and GPS/TEC were obtaining from monitoring sites in Taiwan, Japan and Italy and from global ionosphere maps (GIM) respectively. Our analysis of ground and satellite data during the occurrence of 7 global earthquakes has shown the presence of anomalies in the atmosphere. Our results for Tohoku M9.0 earthquake show that on March 7th, 2011 (4 days before the main shock and 1 day before the M7.2 foreshock of March 8, 2011) a rapid increase of emitted infrared radiation was observed by the satellite data and an anomaly was developed near the epicenter. The GPS/TEC data indicate an increase and variation in electron density reaching a maximum value on March 8. From March 3 to 11 a large increase in electron concentration was recorded at all four Japanese ground-based ionosondes, which returned to normal after the main earthquake. Similar approach for analyzing atmospheric and ionospheric parameters has been applied for China (M7.9, 2008), Italy (M6.3, 2009), Samoa (M7, 2009), Haiti (M7.0, 2010) and Chile (M8.8, 2010) eahquakes. Results have revealed the presence of related variations of these parameters implying their connection with the earthquake process. The second phase (B) of this validation included 102 major earthquakes (M>5.9) in Taiwan and Japan. We have found anomalous behavior before all of these events with no false negatives. False alarm ratio for false positives is less then 10% and has been calculated for the same month of the earthquake occurrence for the entire period of analysis (2003-2009). The commonalities for detecting atmospheric/ionospheric anomalies are: i.) Regularly appearance over regions of maximum stress (i.e., along plate boundaries); ii.) Anomaly existence over land and sea; and iii) association with M>5.9 earthquakes not deeper than 100km. Due to their long duration over the same region these anomalies are not consistent with a meteorological origin. Our initial results from the ISTF validation of multi-instrument space-borne and ground observations show a systematic appearance of atmospheric anomalies near the epicentral area, one to seven (average) days prior to the largest earthquakes, and suggest that it could be explained by a coupling process between the observed physical parameters and the pre-earthquake preparation processes.

Ouzounov, Dimitar↗

Multi-Sensor Observations of Earthquake Related Atmospheric Signals over Major Geohazard Validation Sites

We are conducting a scientific validation study involving multi-sensor observations in our investigation of phenomena preceding major earthquakes. Our approach is based on a systematic analysis of several atmospheric and environmental parameters, which we found, are associated with the earthquakes, namely: thermal infrared radiation, outgoing long-wavelength radiation, ionospheric electron density, and atmospheric temperature and humidity. For first time we applied this approach to selected GEOSS sites prone to earthquakes or volcanoes. This provides a new opportunity to cross validate our results with the dense networks of in-situ and space measurements. We investigated two different seismic aspects, first the sites with recent large earthquakes, viz.- Tohoku-oki (M9, 2011, Japan) and Emilia region (M5.9, 2012,N. Italy). Our retrospective analysis of satellite data has shown the presence of anomalies in the atmosphere. Second, we did a retrospective analysis to check the re-occurrence of similar anomalous behavior in atmosphere/ionosphere over three regions with distinct geological settings and high seismicity: Taiwan, Japan and Kamchatka, which include 40 major earthquakes (M>5.9) for the period of 2005-2009. We found anomalous behavior before all of these events with no false negatives; false positives were less then 10%. Our initial results suggest that multi-instrument space-borne and ground observations show a systematic appearance of atmospheric anomalies near the epicentral area that could be explained by a coupling between the observed physical parameters and earthquake preparation processes.

Ouzounov, D.↗

TPSAS-NF1676L-32111-DND

Automated in-situ NDE monitoring of component state during flight or at set increments on ground often requires integrated instrumentation. Challenges: false negatives, false positives; system validation/certification (confidence in detectability, assessment of limitations, and assessment of lifetime).

Elizabeth Gregory↗

Bayesian event categorization matrix approach for explosion monitoring

Current efforts to correctly categorize natural events from suspected explosion sources with data that is collected by ground- or space-based sensors presents historical challenges that remain unaddressed by the Event Categorization Matrix (ECM) model. Smaller historical events (lower yield explosions) may have data available from fewer measurement techniques than are available today, and therefore, a historical event record can lack a complete set of discriminants. The covariance structures can also differ between such observations of event (source-type) categories. Both obstacles are problematic for the classic ECM model. Our work addresses this gap and presents a Bayesian update to the previous ECM model, termed the Bayesian Event Categorization Matrix model, which can be trained on partial observations and does not rely on a pooled covariance structure. We further augment the ECM model with Bayesian Decision Theory so that false negative or false positive rates of an event categorization can be reduced in an intuitive manner. To demonstrate improved categorization rates for the Bayesian Event Categorization Matrix model, we compare an array of Bayesian and classic models with multiple performance metrics using Monte Carlo experiments. We use both synthetic and real data. Our Bayesian models show consistent gains in overall accuracy and lower false negative rates relative to the classic ECM model. Here, we propose future avenues to improve Bayesian Event Categorization Matrix models’ decision making and predictive capability.

58 GEOSCIENCES↗

The effect of abnormal cell proportion on specimen classifier performance

An analysis is presented of the results obtained from a cell classifier which is confronted with an abnormal/normal cell ratio which is different from the ratio assumed in the calibration of the classifier. False negative and false positive error rates are determined in advance for classifier operation, along with the necessary sample size in order to validate the predicted distributions. Changes are demonstrated to happen only regarding the false negative rate, where reductions in the abnormal cell rate below the expected rates would cause totally unreliable data. Substantial overproduction of abnormal cells would be quickly noticeable, while production rates beyond, but close to, the expected rates would only require more extensive sampling. Classifier systems for 10% proportions of abnormal cells are concluded to be possible, but difficulties are present with much lower rates

Castleman, K. R.↗

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R↗

Human versus automation in responding to failures: an expected-value analysis

A simple analytical criterion is provided for deciding whether a human or automation is best for a failure detection task. The method is based on expected-value decision theory in much the same way as is signal detection. It requires specification of the probabilities of misses (false negatives) and false alarms (false positives) for both human and automation being considered, as well as factors independent of the choice--namely, costs and benefits of incorrect and correct decisions as well as the prior probability of failure. The method can also serve as a basis for comparing different modes of automation. Some limiting cases of application are discussed, as are some decision criteria other than expected value. Actual or potential applications include the design and evaluation of any system in which either humans or automation are being considered.

NASA Discipline Space Human Factors↗

Multi-Parameter Observation and Detection of Pre-Earthquake Signals in Seismically Active Areas

The recent large earthquakes (M9.0 Tohoku, 03/2011; M7.0 Haiti, 01/2010; M6.7 L Aquila, 04/2008; and M7.9 Wenchuan 05/2008) have renewed interest in pre-anomalous seismic signals associated with them. Recent workshops (DEMETER 2006, 2011 and VESTO 2009 ) have shown that there were precursory atmospheric /ionospheric signals observed in space prior to these events. Our initial results indicate that no single pre-earthquake observation (seismic, magnetic field, electric field, thermal infrared [TIR], or GPS/TEC) can provide a consistent and successful global scale early warning. This is most likely due to complexity and chaotic nature of earthquakes and the limitation in existing ground (temporal/spatial) and global satellite observations. In this study we analyze preseismic temporal and spatial variations (gas/radon counting rate, atmospheric temperature and humidity change, long-wave radiation transitions and ionospheric electron density/plasma variations) which we propose occur before the onset of major earthquakes:. We propose an Integrated Space -- Terrestrial Framework (ISTF), as a different approach for revealing pre-earthquake phenomena in seismically active areas. ISTF is a sensor web of a coordinated observation infrastructure employing multiple sensors that are distributed on one or more platforms; data from satellite sensors (Terra, Aqua, POES, DEMETER and others) and ground observations, e.g., Global Positioning System, Total Electron Content (GPS/TEC). As a theoretical guide we use the Lithosphere-Atmosphere-Ionosphere Coupling (LAIC) model to explain the generation of multiple earthquake precursors. Using our methodology, we evaluated retrospectively the signals preceding the most devastated earthquakes during 2005-2011. We observed a correlation between both atmospheric and ionospheric anomalies preceding most of these earthquakes. The second phase of our validation include systematic retrospective analysis for more than 100 major earthquakes (M>5.9) in Taiwan and Japan. We have found anomalous behavior before all of these events with no false negatives. Calculated false alarm ratio for the for the same month over the entire period of analysis (2003-2009) is less than 10% and was d as the earthquakes. The commonalities in detecting atmospheric/ionospheric anomalies show that they may exist over both land and sea in regions of maximum stress (i.e., along plate boundaries) Our results indicate that the ISTF model could provide a capability to observe pre-earthquake atmospheric/ionospheric signals by combining this information into a common framework.

Ouzounov, D.↗

In-Situ Wire Damage Detection System

An In-Situ Wire Damage Detection System (ISWDDS) has been developed that is capable of detecting damage to a wire insulation, or a wire conductor, or to both. The system will allow for realtime, continuous monitoring of wiring health/integrity and reduce the number of false negatives and false positives while being smaller, lighter in weight, and more robust than current systems. The technology allows for improved safety and significant reduction in maintenance hours for aircraft, space vehicles, satellites, and other critical high-performance wiring systems for industries such as energy production and mining. The integrated ISWDDS is comprised of two main components: (1) a wire with an innermost core conductor, an inner insulation film, a conductive layer or inherently conductive polymer (ICP) covering the inner insulation film, an outermost insulation jacket; and (2) smart connectors and electronics capable of producing and detecting electronic signals, and a central processing unit (CPU) for data collection and analysis. The wire is constructed by applying the inner insulation films to the conductor, followed by the outer insulation jacket. The conductive layer or ICP is on the outer surface of the inner insulation film. One or more wires are connected to the CPU using the smart connectors, and up to 64 wires can be monitored in real-time. The ISWDDS uses time domain reflectometry for damage detection. A fast-risetime pulse is injected into either the core conductor or conductive layer and referenced against the other conductor, producing transmission line behavior. If either conductor is damaged, then the signal is reflected. By knowing the speed of propagation of the pulse, and the time it takes to reflect, one can calculate the distance to and location of the damage.

Williams, Martha↗

Data-driven landslide nowcasting at the global scale

Landslides affect nearly every country in the world each year. To better understand this global hazard, the Landslide Hazard Assessment for Situational Awareness (LHASA) model was developed previously. LHASA version 1 combines satellite precipitation estimates with a global landslide susceptibility map to produce a gridded map of potentially hazardous areas from 60° North-South every 3 h. LHASA version 1 categorizes the world’s land surface into three ratings: high, moderate, and low hazard with a single decision tree that first determines if the last seven days of rainfall were intense, then evaluates landslide susceptibility. LHASA version 2 has been developed with a data-driven approach. The global susceptibility map was replaced with a collection of explanatory variables, and two new dynamically varying quantities were added: snow and soil moisture. Along with antecedent rainfall, these variables modulated the response to current daily rainfall. In addition, the Global Landslide Catalog (GLC) was supplemented with several inventories of rainfall-triggered landslide events. These factors were incorporated into the machine-learning framework XGBoost, which was trained to predict the presence or absence of landslides over the period 2015–2018, with the years 2019–2020 reserved for model evaluation. As a result of these improvements, the new global landslide nowcast was twice as likely to predict the occurrence of historical landslides as LHASA version 1, given the same global false positive rate. Furthermore, the shift to probabilistic outputs allows users to directly manage the trade-off between false negatives and false positives, which should make the nowcast useful for a greater variety of geographic settings and applications. In a retrospective analysis, the trained model ran over a global domain for 5 years, and results for LHASA version 1 and version 2 were compared. Due to the importance of rainfall and faults in LHASA version 2, nowcasts would be issued more frequently in some tropical countries, such as Colombia and Papua New Guinea; at the same time, the new version placed less emphasis on arid regions and areas far from the Pacific Rim. LHASA version 2 provides a nearly real-time view of global landslide hazard for a variety of stakeholders.

XGBoos↗

Exoplanet Biosignatures: Understanding Oxygen as a Biosignature in the Context of Its Environment

Here we review how environmental context can be used to interpret whether O 2 is a biosignature in extrasolar planetary observations. This paper builds on the overview of current biosignature research discussed in Schwieterman et al. (2017), and provides an in-depth, interdisciplinary example of biosignature identification and observation that serves as a basis for the development of the general framework for biosignature assessment described in Catling et al., (2017). O 2 is a potentially strong biosignature that was originally thought to be an unambiguous indicator for life at high-abundance. In exploring O 2 as a biosignature, we describe the coevolution of life with the early Earth's environment, and how the interplay of sources and sinks in the planetary environment may have resulted in suppression of O 2 release into the atmosphere for several billion years, a false negative for biologically generated O 2 . False positives may also be possible, with recent research showing potential mechanisms in exoplanet environments that may generate relatively high abundances of atmospheric O 2 without a biosphere being present. These studies suggest that planetary characteristics that may enhance false negatives should be considered when selecting targets for biosignature searches. Similarly our ability to interpret O 2 observed in an exoplanetary atmosphere is also crucially dependent on environmental context to rule out false positive mechanisms. We describe future photometric, spectroscopic and time-dependent observations of O 2 and the planetary environment that could increase our confidence that any observed O 2 is a biosignature, and help discriminate it from potential false positives. The rich, interdisciplinary study of O 2 illustrates how a synthesis of our understanding of life's evolution and the early Earth, scientific computer modeling of star-planet interactions and predictive observations can enhance our understanding of biosignatures and guide and inform the development of next-generation planet detection and characterization missions. By observing and understanding O 2 in its planetary context we can increase our confidence in the remote detection of life, and provide a model for biosignature development for other proposed biosignatures.

Victoria S Meadows↗

Exploring Applications of Machine Learning for Wildfire Monitoring and Detection using Unmanned Aerial Vehicles

Wildfires are increasing in frequency and severity around the world, including the United States. The losses caused by wildfires could be mitigated if high-risk areas, hotspots, and flare-ups could be monitored continuously, such as through the use of Unmanned Aerial Vehicles (UAVs). This paper documents exploratory efforts using machine learning to determine efficient flight paths for UAVs and to detect wildfires using image classification. On path planning, three machine learning techniques—Genetic Algorithm, Simulated Annealing, and Dynamic Programming—were explored. Genetic Algorithm was found to be an effective approach for path planning for wildfire monitoring and surveillance by UAVs. For a scenario of 25 locations in a circular arrangement, the algorithm was able to return the optimal path. The accuracy and execution time was found to be sensitive to the algorithm hyperparameters selected, which was especially evident in scenarios with hundreds or thousands of locations. Simulated Annealing was also found to be an effective approach for UAV path planning, with a major benefit of avoiding getting trapped in local minima and being straightforward to implement. Like Genetic Algorithm, the performance of Simulated Annealing was also found to be sensitive to the algorithm hyperparameters selected. By comparison, Dynamic Programming guarantees optimality for any number of locations, but it was found to be less practical in terms of execution time for scenarios with more than about a couple dozen locations. On wildfire detection, image classification using deep learning with a convolutional neural network was explored. Transfer learning was found to be a useful technique to efficiently train deep learning models. Also, it was determined that GPU processing can increase training speed by an order of magnitude, which enables significantly faster development. For a validation test set of 500 images, there were only two false negatives and zero false positives. These results demonstrate that detecting wildfires in static cameras using machine learning is feasible and establish a baseline for using images captured by UAVs in flight for wildfire detection.

Wildfire management↗

Detectability of Varied Hybridization Scenarios Using Genome-Scale Hybrid Detection Methods

Hybridization events complicate the accurate reconstruction of phylogenies, as they lead to patterns of genetic heritability that are unexpected under traditional, bifurcating models of species trees. This phenomenon has led to the development of methods to infer these varied hybridization events, both methods that reconstruct networks directly, as well as summary methods that predict individual hybridization events from a subset of taxa. However, a lack of empirical comparisons between methods – especially those pertaining to large networks with varied hybridization scenarios – hinders their practical use. Here, we provide a comprehensive review of popular summary methods: TICR, MSCquartets, HyDe, Patterson’s D-Statistic (ABBA-BABA), D3, and Dp. TICR and MSCquartets are based on quartet concordance factors gathered from gene tree topologies and HyDe, Patterson’s D-Statistic, D3, and Dp use site pattern frequencies to identify hybridization events between sets of three taxa. We then use simulated data to address questions of method accuracy and ideal use scenarios by testing methods against complex networks which depict gene flow events that differ in depth (timing), quantity (single vs. multiple, overlapping hybridizations), and rate of gene flow (γ). We find that deeper or multiple hybridization events may introduce noise and weaken the signal of hybridization, leading to higher relative false negative rates across all methods. Despite some forms of hybridization eluding quartet-based detection methods, MSCquartets displays high precision in most scenarios. While HyDe results in high false negative rates when tested on hybridizations involving extinct or unsampled ghost lineages, HyDe is the only method able to identify the direction of hybridization, distinguishing the source parental lineages from recipient hybrid lineages. Lastly, we test the methods on a dataset of ultraconserved elements from the bee subfamily Nomiinae, finding possible hybridization events between clades which correspond to regions of poor support in the species tree estimated in a previous study.

Bjorner, Marianne B.↗

Using electronic health record metadata to predict housing instability amongst veterans

Housing instability is considered a significant life stressor and preemptive screening should be applied to identify those at risk for homelessness as early as possible so that they can be targeted for specialized care. We developed models to classify patient outcomes for an established VA Homelessness Screening Clinical Reminder (HSCR), which identifies housing instability, in the two months prior to its administration. Logistic Regression and Random Forest models were fit to classify responses using the last 18 months of document activity. We measure concentration of risk across stratifications of predicted probability and observe an enriched likelihood of finding confirmed false negative responses from veterans with diagnosed housing instability. Positive responses were 34 times more likely to be detected within the top 1 % of patients predicted at risk than from those randomly selected. There is a 1 in 4 chance of detecting false negatives within the top 1 % of predicted risk. Machine learning methods can classify between episodes of housing instability using a data-driven approach that does not rely on variables curated from domain experts. This method has the potential to improve clinicians’ ability to identify veterans who are experiencing housing instability but are not captured by HSCR.

60 APPLIED LIFE SCIENCES↗

Planetary Protection Technologies for Planetary Science Instruments, Spacecraft, and Missions: Report of the NASA Planetary Protection Technology Definition Team (PPTDT)

Planetary bodies like Mars, Europa, and Enceladus pose the question, "How to study them without contaminating them and destroying future prospects to detect life, if it is there?" The natural trade-off, of course, is that the cleaner your spacecraft, the more you can explore such a body without risk of contaminating it. As chartered by NASA Headquarters, the Planetary Protection Technology Definition Team (PPTDT) was asked to provide a report covering six different areas related to the engineering and technology challenges of implementing planetary protection requirements on solar system exploration missions, including: Assessment of technical and engineering challenges to applying available microbial-reduction methods, including recontamination prevention, to spacecraft hardware and instruments, to meet current NASA requirements on preventing the forward contamination of potentially habitable worlds by future spacecraft missions (orbiters, atmospheric missions, landers, penetrators, and drills); Identification of spacecraft and instrument materials known to be compatible with existing planetary protection protocols; Planetary protection protocols/processes available or which appear promising, and areas ripe for technological development; The technical and engineering challenges in ensuring that spacecraft hardware and instruments can meet organic cleanliness requirements needed to ensure high confidence in differentiating Earth contamination from extraterrestrial signals to avoid false negative as well as false positive results; Approaches for mitigating the identified challenges that would allow instruments to be flown successfully at the required levels of cleanliness and microbial reduction, beginning with identification of commonly used materials and spacecraft hardware that are compatible (or particularly vulnerable) to planetary protection protocols; Engineering, technology, and scientific research and development that could be funded by NASA to provide future capabilities to field scientific instruments and spacecraft on missions that require either subsystem or system-level microbial reduction and recontamination prevention.

John D. Rummel↗