Search NASASearch

SEARCH · Search NASA

Results for “Model Counting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Poisson Log-Normal Process for Count Data Prediction

Modeling count data is important in physics and other scientific disciplines, where measurements often involve discrete, non-negative quantities such as photon or neutrino detection events. Traditional parametric approaches can be trained to generate integer-count predictions but may struggle with capturing complex, non-linear dependencies often observed in the data. Gaussian process (GP) regression provides a robust non-parametric alternative to modeling continuous data; however, it cannot generate integer outputs. We propose the Poisson Log-Normal (PoLoN) process, a framework that employs GP to model Poisson log-rates. As in GP regression, our approach relies on the correlations between data points captured via GP kernel structure rather than explicit functional parameterizations. We demonstrate that the PoLoN predictive distribution is Poisson-LogNormal and provide an algorithm for optimizing kernel hyperparameters. Furthermore, we adapt the PoLoN approach to the problem of detecting weak localized signals superimposed on a smoothly varying background - a task of considerable interest in many areas of science and engineering. Our framework allows us to predict the strength, location and width of the detected signals. We evaluate PoLoN's performance using both synthetic and real-world datasets, including the open dataset from CERN which was used to detect the Higgs boson at the Large Hadron Collider. Our results indicate that the PoLoN process can be used as a non-parametric alternative for analyzing, predicting, and extracting signals from integer-valued data.

Saha, Anushka [Rutgers U., Piscataway]

Misclassification in Workers’ Telecommuting Frequency Choices Using a Generalized Extreme Value Model

Telecommuting frequency is a response variable collected in travel surveys and is, therefore, prone to errors leading to mismeasurements or misclassification. Misclassification of explanatory variables is a common risk when using statistical modeling techniques. We define “misclassification” as a response reported or recorded in the wrong category; for example, a variable is recorded as a 1 when it should be 0. Here, in this context, this study aims to develop a statistical model to analyze telecommuting data which accounts for potential misclassification errors by building on existing literature in econometrics. The empirical analysis was undertaken using the 2017 National Household Travel Survey (NHTS) and the general extreme value (GEV) models available in the literature. Specifically, the frequency of telecommuting days was analyzed using the negative binomial (NB) model recast as the multinomial logit (MNL) model. By nature—and consistent with other studies—NHTS data are prone to errors that can be classified as intentional or unintentional misinformation provided by the person being interviewed. Ignoring these errors while modeling telecommuting frequencies using standard discrete count models can result in biased parameter estimates. The misclassification parameter was calculated for both over-reporting and under-reporting scenarios. The misclassification errors can be as high as 14% over-reported and 10% under-reported, particularly for the neighboring values. Statistical fit comparison between the models shows that models that ignore misclassification have worse data fit and biased parameter estimates with significant policy implications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Practical Probabilistic Programming

Recent advances in probabilistic programming languages (PPLs) have provided the capability for exact inference: computing a closed-form probability distribution for a given probabilistic program. In particular, the new language Roulette uses a language oriented programming (LOP) approach, wherein analysts build new programming languages on top of a set of primitives provided by Roulette, which then translates these structures into a weighted model counting problem which can be solved by automated reasoning tools. However, because Roulette provides few convenience features, developing these new languages is challenging even for expert users. We developed a standard library of common probability functions for Roulette with the goal of improved usability. This included approximation of continuous probability density functions using discrete probability mass functions. We demonstrated this approach by modeling a cosmic ray striking a RAM controller. We found that Roulette provides a powerful interface for highly expressive probabilistic programs to be generated. In collaboration with the NNSA Advanced Simulation and Computing program, which resulted in development of a tool called Circulette, we were able to model complex circuits expressed in Verilog using probabilistic programs with an expressivity not previously possible. Our research question that motivated the development of a Roulette standard library was to determine whether non-experts could use a PPL to model relevant problems regarding radiation effects on microelectronics. This standard library improved the expressivity of Roulette by implementing common probability density functions, mathematical operators on distributions, and support for empirical distributions. While Roulette is a powerful modeling language, the untyped, LOP approach makes error messages difficult to understand and requires expert aid. We recommend further research on Roulette, especially with its error messages, to enable improved usability. At the same time, this project demonstrated that for users familiar with Roulette and the LOP approach, Roulette provides powerful new capabilities that can be integrated with other Sandia modeling capabilities.

97 MATHEMATICS AND COMPUTING

Radon emanation rate measurements using liquid scintillation counting

This article describes a radon emanation measurement technique using liquid scintillator counting. A model for radon loading and transport is described, along with its calibration. Detector background and blank have been studied and quantified. The Minimal detectable activity has been determined for the counting setup using a toy Monte Carlo simulation. Here, the measurement technique is validated using a butyl rubber sample previously used for cross-calibration between different radon counting facilities.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Quantifying health benefits of sustainable aviation fuels: Modeling decreased ultrafine particle emissions and associated impacts on communities near the Seattle-Tacoma International Airport

Exposure to ultrafine particles (UFP, ≤100 nm) is an emerging health concern linked to premature mortality, with jet fuel combustion identified as a significant source of UFPs near airports. Sustainable aviation fuel (SAF) adoption has the potential to reduce aviation-related UFPs and may particularly benefit populations who reside nearby. However, assessing aviation-specific impacts on health remains challenging due to the lack of tools capable of addressing: fine-scale exposure evaluation, novel ambient pollutants, and groups with increased exposure or susceptibility. We develop and apply a method to estimate reductions in mortality associated with aviation-related UFP reductions at the Seattle-Tacoma (SEA-TAC) International Airport under SAF adoption scenarios, with a focus on near-airport communities. Using UFP exposure surfaces generated from AERMOD modeling, flight count data, and UFP measurements, we evaluated UFP reductions under various control scenarios. We estimated mortality reductions by combining this with population data, baseline mortality, and a hazard ratio of 1.012 (95 % confidence interval: 1.010, 1.015) per interquartile range increment of 2723 particles/cm 3 . Our analysis included 412 census tracts representing almost 1.5 million adults. Baseline aviation-related UFP exposures averaged 1145 (SD: 277) particles/cm 3 . The highest baseline concentrations and subsequent reductions under SAF scenarios were near SEA-TAC. Mortality case reductions averaged between 3.1 (95 % range: 2.5–3.7) for a 5 % UFP reduction to 31.0 (24.6–37.4) for a 50 % reduction, with corresponding mortality rate reductions of 0.2 (0.2–0.3) to 2.1 (1.7–2.5) cases per 100,000 people per year. Mortality rate reductions were larger among populations residing closer to SEA-TAC, including those that were Hispanic or Latino, below-poverty, and did not identify as White. Reducing aviation-related UFPs through SAF adoption could lead to lower mortality, particularly in near-airport communities. This reproducible approach can be adapted to other settings to evaluate health benefits from aviation-related UFP reductions.

Aviation-related air pollution

Strategies for the in-orbit gain tracking using the modulated X-ray sources for the Resolve microcalorimeter spectrometer on the X-ray Imaging and Spectroscopy Mission

Accurate and precise correction of the gain drift is the key to achieve the required energy resolution of the Resolve microcalorimeter spectrometer on the X-ray Imaging and Spectroscopy Mission (XRISM). Therefore, Resolve is equipped with highly configurable X-ray sources called the modulated X-ray source (MXS). The pulsed nature allows us to separate calibration and astrophysical X-rays by time interval selections. However, undesirable characteristics of the MXS, such as the afterglow X-rays, restrict the allowed configuration range. Moreover, the nonlinear and discontinuous behaviors of the calibration line count rate make the determination of the optimal setting highly complex. The MXS count rate model has been established using measurements in the spacecraft thermal vacuum test with the flight detector and MXS. A trade-off study using the model enables us to choose a few settings for the continuous use of the MXS optimized for different ranges of target gain tracking intervals. An alternative approach, where the MXS is used only intermittently, has also been developed and implemented. This new mode enables us to reconstruct the drift without having most of the undesirable effects in science data. This also forms the basis of the gain tracking under the current Resolve configuration with the closed gate valve.

Astronomy and AstroPhysics

Near-Efficient and Non-Asymptotic Multiway Inference

We establish non-asymptotic efficiency guarantees for tensor decomposition–based inference in count data models. Under a Poisson framework, we consider two related goals: (i) parametric inference , the estimation of the full distributional parameter tensor, and (ii) multiway analysis , the recovery of its canonical polyadic (CP) decomposition factors. Our main result shows that in the rank-one setting, a rank-constrained maximum-likelihood estimator achieves multiway analysis with variance matching the Cramér–Rao Lower Bound (CRLB) up to absolute constants and logarithmic factors. This provides a general framework for studying “near-efficient” multiway estimators in finite-sample settings. For higher ranks, we illustrate that our multiway estimator may not attain the CRLB; nevertheless, CP-based parametric inference remains nearly minimax optimal, with error bounds that improve on prior work by offering more favorable dependence on the CP rank. Numerical experiments corroborate near-efficiency in the rank-one case and highlight the efficiency gap in higher-rank scenarios.

97 MATHEMATICS AND COMPUTING

How are Heterogeneous Nucleation Rate Observations Influenced by Instrument Resolution?

Experimental measurements of the heterogeneous nucleation rate rely on counting the number of nuclei with time. However, the size of a thermodynamically stable nucleus is often a few nanometers in diameter and is below the resolution of most (in situ) measurement techniques that provide a statistically valid sample. Due to the finite resolution of the instruments and analysis methods, it is challenging to capture the incipient nuclei and the subsequent evolution of nuclei density over time. In this work, we demonstrate the impact of instrument resolution on observed nuclei densities by comparing numerical modeling with experimental results. Further, to achieve this, we implemented heterogeneous nucleation within the pore-scale reactive transport modeling framework using classical nucleation theory (CNT). We compared the modeling results with nucleation rates measured using X-ray nanotomography (XnT) and evaluated how these impact the apparent values of the prefactor and interfacial energy based on CNT and the crystal growth rate. Specifically, we applied a resolution threshold (artificial resolution limit) in the model during nuclei counting to resemble an experimental resolution, ranging from 15 to 500 nm. The findings reveal that the instrument resolution significantly impacts the apparent prefactor and interfacial energy. Both apparent prefactor and interfacial energy decrease with a decrease in the instrument resolution. While deviation in the prefactor due to resolution is anticipated, those in the interfacial energy are unexpected. The approach described here allows one to correct apparent nucleation rates that depend on the instrument’s resolution to derive “intrinsic” CNT parameters for the prefactor and interfacial energy.

47 OTHER INSTRUMENTATION

Expression of a mammalian RNA demethylase increases flower number and floral stem branching in Arabidopsis thaliana

Abstract RNA methylation plays a central regulatory role in plant biology and is a relatively new target for plant improvement efforts. In nearly all cases, perturbation of the RNA methylation machinery results in deleterious phenotypes. However, a recent landmark paper reported that transcriptome‐wide use of the human RNA demethylase FTO substantially increased the yield of rice and potatoes. Here, we have performed the first independent replication of those results and demonstrated broader transferability of the trait, finding increased flower and fruit count in the model species Arabidopsis thaliana . We also performed RNA‐seq of our FTO‐transgenic plants, which we analyzed in conjunction with previously published datasets to detect several previously unrecognized patterns in the functional and structural classification of the upregulated and downregulated genes. From these, we present mechanistic hypotheses to explain these surprising results with the goal of spurring more widespread interest in this promising new approach to plant engineering.

59 BASIC BIOLOGICAL SCIENCES

Leveraging 13C-Labeling to Assign Molecular Formulas to Unknown Yeast Metabolites

Mass spectrometry analyses have identified tens of thousands of unknown small molecule-associated peaks in different biological specimens. Notably, even the simplest and best studied organisms like Escherichia coli and Saccharomyces cerevisiae yield thousands of unknown peaks. A key question is how many of these reflect actual novel endogenous metabolites. To explore this, Mahieu and Patti used complete 13 C -labeling in E. coli to credential peaks as biological. This reduced the number of unknowns by more than 90%. Here, we carry out similar uniform 13 C-labeling in the Baker’s yeast S. cerevisiae and two less-studied bioenergy-relevant yeasts Rhodotorula toruloides (lipid producer) and Issatchenkia orientalis (organic acid producer). Identification of unknown metabolite peaks and their molecular formulas is facilitated through software tailored for 13 C labeling data and resulting knowledge of carbon atom count. A classification model evaluates the plausibility of each candidate formula, with peaks lacking plausible candidate formulas unlikely to reflect metabolite molecular ions. This approach prioritizes about one hundred candidate abundant unknown metabolites with logical molecular formulas. Most of these are species-specific rather than conserved across yeasts, and more are found in the nonmodel yeasts than S. cerevisiae. Thus, 13 C-labeling data on unknown metabolites highlights the potential for discovering new metabolites and pathways in nonmodel yeasts.

Carbon

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics

Kinetics of photogenerated carbon dangling bonds in organic photovoltaic thin Films: An EPR study

Here, we report an investigation of the early kinetics of photogenerated carbon dangling bond (CDB) formation and annealing in organic photovoltaic bulk heterojunction (BHJ) thin film blends under oxygen- and moisture-free conditions, using X-band electron paramagnetic resonance (EPR) spectroscopy. The study focuses on donor:acceptor BHJ blends of PCE12:PCBM and PCE12:ITIC films, where PCE12 is PBDB-T. The time evolution of CDBs in such drop-cast BHJ films irradiated at 300 nm is monitored. The early kinetics of CDB formation, critical for understanding OPV degradation mechanisms, is studied. Theoretical analysis of the defect growth mechanism suggests a monomolecular defect creation model where the defect count follows a power-law t β with irradiation time t, where β ∼ 0.55–0.58, in excellent agreement with the theoretically expected value of β = 1/2. This model is compatible with CDB formation by the holes in donor sites adjacent to acceptors, likely assisted by energy released from quenching of nearby excitons by the holes, elucidating the physical mechanism underlying CDB formation. This is significant for designing improved materials, which mitigate defect creation, and consequently advancing the development of stable OPV systems.

42 ENGINEERING

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence

High-count rate effects in event processing for XRISM/ Resolve X-ray microcalorimeter: I. Ground test

The spectroscopic performance of an X-ray microcalorimeter is compromised at high count rates. We utilize the Resolve X-ray microcalorimeter onboard the XRISM satellite to examine the effects observed during high-count rate measurements and propose modeling approaches to mitigate them. We specifically address the following instrumental effects that impact performance: CPU limit, pile-up, and untriggered electrical cross-talk. Experimental data at high count rates were acquired during ground testing using the flight model instrument and a calibration X-ray source. In the experiment, data processing not limited by the performance of the onboard CPU was run in parallel, which cannot be done in orbit. This makes it possible to access the data degradation caused by limited CPU performance. We use these data to develop models that allow for a more accurate estimation of the aforementioned effects. To illustrate the application of these models in observation planning, we present a simulated observation of GX 13+1. Understanding and addressing these issues is crucial to enhancing the reliability and precision of X-ray spectroscopy in situations characterized by elevated count rates.

47 OTHER INSTRUMENTATION

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE

Prime VI

SAND2025-03757O Prime VI is a distribution-of-disease outbreak model calibration code based on variational inference. It accompanies a publication for submission to Statistics in Medicine journal, and the code will be maintained for open-source use on Sandia's GitLab. The software provides methods for calibrating an epidemiological model to measured case-count data for a multitude of correlated spatial regions. The code solves a Bayesian inverse problem for model calibration where the posterior over-model parameters are approximated through a custom implementation of variational inference. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin

ThermoPore: Predicting part porosity based on thermal images using deep learning

Part qualification is often a critical and labor-intensive process in additive manufacturing, particularly in the detection of defects such as porosity, which stands to benefit significantly from advancements in machine learning. We present a deep learning approach for quantifying and localizing ex-situ porosity within Laser Powder Bed Fusion fabricated samples utilizing in-situ thermal image monitoring data. Our goal is to build the real time porosity map of parts based on thermal images acquired during the build. The quantification task builds upon the established Convolutional Neural Network model architecture to predict pore count and the localization task leverages the spatial and temporal attention mechanisms of the novel Video Vision Transformer model to indicate areas of expected porosity. Our model for porosity quantification achieved a R 2 score of 0.57 and our model for porosity localization produced an average Intersection over Union (IoU) score of 0.32 and a maximum of 1.0. This work is setting the foundations of part porosity “Digital Twins” based on additive manufacturing monitoring data and can be applied downstream to reduce time-intensive post-inspection and testing activities during part qualification and certification. In addition, we seek to accelerate the acquisition of crucial insights normally only available through ex-situ part evaluation by means of machine learning analysis of in-situ process monitoring data.

Deep learning