Search NASA⌕ Search

SEARCH · Search NASA

Results for “false negative analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Open Specy 1.0: Automated (Hyper)spectroscopy for Microplastics

Microplastic spectral analysis is one of the most time-consuming processes in studying microplastic pollution, often requiring days per sample. Researchers are transitioning to automated batch and hyperspectral image analysis techniques to enhance efficiency. Open Specy, initially aimed at manual single-spectrum analysis, has now integrated automated methods. This updated version, Open Specy 1.0, introduces several new features, including two algorithms for automated processing (smoothing and particle compression), an extensive library containing over 40,000 open-source Raman and FTIR spectra, and two machine learning classifiers (logistic regression and k medoids) developed from this library. Furthermore, it includes a revamped user interface, an R package, and a benchmark data set for testing future advancements in automated techniques. Researchers evaluated various configurations for hyperspectral smoothing, particle identification, compression, and splitting, to achieve combined recovery rates between 50 and 150% particle counts, identities, and sizes with a coefficient of variation (CV) of less than 40% (the accredited standard). Mean absorbance times the standard deviation provided a consistent particle identification. Hyperspectral smoothing led to a 96% combined recovery rate and reduced variability (CV = 38%) compared to the 86% recovery (CV = 83%) of nonsmoothed controls. Additionally, compressing spectra for particles was significantly faster (>3x) and showed similar accuracy but with reduced variability than processing each pixel individually. Key challenges persist in automating spectral analysis, particularly in refining particle splitting algorithms, and improving identification routines to minimize false positives and negatives. In conclusion, new methods in sample preparation for better stabilization and dispersion of particles could overcome some of these issues.

13 HYDRO ENERGY↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Chemistry↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Arasteh, Kourosh [Lawrence Livermore National Labo↗

Data-driven modeling of dynamic occupant thermostat override behavior for demand response applications

Buildings consume nearly 40% of global energy and produce similar emissions. Whiletechnological advances address efficiency, occupant behavior causes energy use variations up to 300% between identical buildings. This gap between predicted and actual building performance impacts building design, operations, and grid demand management programs. Through analyses of smart thermostat data from 1,400 single-occupant homes, the researchdemonstrates that occupants respond to 8°F thermostat setpoint changes within a median of 15 minutes, while 2°F changes trigger responses within a median of 30 minutes. This highlights an understudied temporal relationship between thermostat setbacks and response time of occupant behaviors. Models of such behavior dynamics are required to incorporate occupant impacts into building performance simulation. A key contribution of this dissertation is the Thermal Frustration Theory (TFT), which positsthat thermal discomfort driven behaviors are caused by the time-accumulation of discomfort, not simply a temperature deviation threshold or a delay from an initiating event. Using a dataset of 634 thermostats, each with 25+ manual setpoint changes, a comparative analysis of TFT and comfort zone and a delayed response theories demonstrated that personalized TFT models better predict when manual setpoint change occur. This was measured by the area under the curve statistical measure (AUC); all three models perform similarly by a Matthews Correlation Coefficient measure. Higher AUC performance is especially important for modeling occupant behavior in demand response programs where false negatives of rare occupant interactions could adversely affect grid stability. EnergyPlus based simulations were conducted with TFT-derived occupant models, demonstrating the ability to identify parameters of known TFT models from only data observable with smart thermostats, even under the presence of noise from routine overrides. Overall, the dissertation highlights that thermostat interactions are neither static,instantaneous, nor driven solely by the environment. Instead, temporal accumulation of discomfort and routine-based behavior play important roles. The methodology and results offer a pathway towards more accurate modeling of human-building interactions for policy assessment, building design, and demand response programs.

Sharma, Kunind [Northeastern University] (ORCID:00↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R↗

Inferring precocial Chinook Salmon production through single‐parentage assignments

Abstract Objective Parentage analysis is a routine methodology in fisheries research, but study systems exist where it is impractical to sample both parents. The ability to reliably assign offspring to a single parent is beneficial in these situations. We applied single‐parentage assignments to a naturally spawning population of Chinook Salmon Oncorhynchus tshawytscha to quantify production of anadromous returns by unsampled precocial males. Methods We used an approach that focused on two important aspects of parentage analyses: (1) addressing the presence of family structure within the set of sampled parents and (2) controlling for false‐positive and false‐negative assignments. Result Results indicated that 30% of reproductively successful males were precocial males, which produced 20% of the returning anadromous offspring. Conclusion This study provides a framework for applying single‐parent assignments in a salmonid study system while explicitly addressing sources of assignment errors.

Steele, Craig A.↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Red Noise–based False Alarm Thresholds for Astrophysical Periodograms via Whittle’s Approximation to the Likelihood

Astronomers who search for periodic signals using Lomb–Scargle periodograms rely on false alarm level (FAL) estimates to identify statistically significant peaks. Although FALs are often calculated from white noise models, many astronomical time series suffer from red noise. Prewhitening is a statistical technique in which a continuum model is subtracted from the log power spectrum estimate, after which the observer can proceed with a white-noise treatment. Here we present a prewhitening-based method of calculating frequency-dependent FALs. We fit power laws and autoregressive models of order 1 to each Lomb–Scargle periodogram by minimizing the Whittle approximation to the negative log-likelihood (NLL), then calculate FALs based on the best-fit model power spectrum. Our technique is a novel extension of the Whittle NLL to datasets with uneven time sampling. We demonstrate FAL calculations using observations of α Cen B, GJ 581, HD 192310, synthetic data from the radial velocity (RV) fitting challenge, and Kepler observations of a differential rotator. The Kepler data analysis shows that only true rotation signals are detected by red noise FALs, while white noise FALs suggest all spurious peaks in the low-frequency range are significant. A high-frequency sinusoid injected into α Cen B logR$'$ HK observations exceeds the 1% red noise FAL despite having only 8.9% of the power of the dominant rotation signal. In a periodogram of HD 192310 RVs, peaks associated with differential rotation and planets are detected against the 5% red noise FAL without iterative model fitting or subtraction. The software for calculating red noise–based FALs is available on GitHub.

Astrostatistics (1882)↗

Real-Time, Adaptive Radiological Anomaly Detection and Isotope Identification Using Non-Negative Matrix Factorization

Spectroscopic anomaly detection and isotope identification algorithms are integral components in nuclear nonproliferation applications such as search operations. The task is especially challenging in the case of mobile detector systems because the observed gamma-ray background changes more than for a static detector system, and a pretrained background model can easily find itself out of domain. The result is that algorithms may exceed their intended false alarm rate or sacrifice detection sensitivity to maintain the desired false alarm rate. Non-negative matrix factorization (NMF) is a powerful tool for spectral anomaly detection and identification, but, like many similar algorithms that rely on data-driven background models, in its conventional implementation, it is unable to update in real time to account for environmental changes that affect the background spectroscopic signature. Here, we have developed a novel NMF-based algorithm that periodically updates its background model to accommodate changing environmental conditions. The adaptive NMF algorithm involves fewer assumptions about its environment, making it more generalizable than existing NMF-based methods while maintaining or exceeding detection performance on simulated and real-world datasets.

Anomaly detection↗