Search NASASearch

SEARCH · Search NASA

Results for “Under-reporting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Misclassification in Workers’ Telecommuting Frequency Choices Using a Generalized Extreme Value Model

Telecommuting frequency is a response variable collected in travel surveys and is, therefore, prone to errors leading to mismeasurements or misclassification. Misclassification of explanatory variables is a common risk when using statistical modeling techniques. We define “misclassification” as a response reported or recorded in the wrong category; for example, a variable is recorded as a 1 when it should be 0. Here, in this context, this study aims to develop a statistical model to analyze telecommuting data which accounts for potential misclassification errors by building on existing literature in econometrics. The empirical analysis was undertaken using the 2017 National Household Travel Survey (NHTS) and the general extreme value (GEV) models available in the literature. Specifically, the frequency of telecommuting days was analyzed using the negative binomial (NB) model recast as the multinomial logit (MNL) model. By nature—and consistent with other studies—NHTS data are prone to errors that can be classified as intentional or unintentional misinformation provided by the person being interviewed. Ignoring these errors while modeling telecommuting frequencies using standard discrete count models can result in biased parameter estimates. The misclassification parameter was calculated for both over-reporting and under-reporting scenarios. The misclassification errors can be as high as 14% over-reported and 10% under-reported, particularly for the neighboring values. Statistical fit comparison between the models shows that models that ignore misclassification have worse data fit and biased parameter estimates with significant policy implications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Weak baselines and reporting biases lead to overoptimism in machine learning for fluid-related partial differential equations

One of the most promising applications of machine learning in computational physics is to accelerate the solution of partial differential equations (PDEs). The key objective of machine-learning-based PDE solvers is to output a sufficiently accurate solution faster than standard numerical methods, which are used as a baseline comparison. Here, we first perform a systematic review of the ML-for-PDE-solving literature. Out of all of the articles that report using ML to solve a fluid-related PDE and claim to outperform a standard numerical method, we determine that 79% (60/76) make a comparison with a weak baseline. Second, we find evidence that reporting biases are widespread, especially outcome reporting and publication biases. We conclude that ML-for-PDE-solving research is overoptimistic: weak baselines lead to overly positive results, while reporting biases lead to under-reporting of negative results. To a large extent, these issues seem to be caused by factors similar to those of past reproducibility crises: researcher degrees of freedom and a bias towards positive results. We call for bottom-up cultural changes to minimize biased reporting as well as top-down structural reforms to reduce perverse incentives for doing so.

97 MATHEMATICS AND COMPUTING

Experimental uncertainty quantification using templates of expected measurement uncertainties for fast neutron-induced total, capture, and scattering cross sections

Careful experimental uncertainty quantification (UQ) is key for developing trustworthy evaluated nuclear data. Templates to account for missing or under-reported experimental uncertainties were recently developed by the covariance committee of Cross Section Evaluation Working Group (CSEWG). In this work, we illustrate the practical application and limitations of these templates for selected neutron-induced reactions, including (n, tot), (n, γ), and (n, xn) in the fast energy range, to illustrate their use in data analyses for nuclear data evaluations. We show that while the templates provide consistent framework, proper implementation still requires detailed knowledge of experimental conditions and careful treatment of nonlinear effects in cross section derivation. Case studies highlight how template-assisted UQ improves consistency with previous evaluations such as ENDF/B and reveals open challenges in propagating uncertainties across different energy regimes. The main contribution of this paper is to connect formal template recommendations with their use in practical evaluation workflows, clarifying both their benefits and current limitations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Structure-aware Initialization via Numerical Continuation and Informed Priors

Scientific machine learning (SciML) often operates in ill-conditioned, weakly identifiable regimes due to limited data or indirect observations. In such settings, optimization and inference are highly sensitive to the starting point, making initialization--often under-reported--a consequential degree of freedom. Random initialization is not a neutral default as it induces an implicit prior over candidate solutions and can systematically bias the result, producing large run-to-run variability. Here, we formalize this view by treating initialization as a hidden confounder in SciML and develop a unifying theory for structure-aware initialization via numerical continuation, constructing warm starts from related problem instances. Across representative tasks, including physics-informed neural networks, maximum likelihood estimation, and variational inference, warm starts have been shown to consistently reduce optimization effort and improve reliability.

Data integrity

Machine learning mathematical models for incidence estimation during pandemics

Accurate estimates of the incidence of infectious diseases are key for the control of epidemics. However, healthcare systems are often unable to test the population exhaustively, especially when asymptomatic and paucisymptomatic cases are widespread; this leads to significant and systematic under-reporting of the real incidence. Here, we propose a machine learning approach to estimate the incidence of a pandemic in real-time, using reported cases and the overall test rate. In particular, we use Bayesian symbolic regression to automatically learn the closed-form mathematical models that most parsimoniously describe incidence. We develop and validate our models using COVID-19 incidence values for nine different countries, confirming their ability to accurately predict daily incidence. Remarkably, despite the differences in epidemic trajectories and dynamics across countries, we find that a single model for all countries offers a more parsimonious description and is more predictive of actual incidence compared to separate models for each country. Our results show the potential to accurately model incidence in real-time using closed-form mathematical models, providing a valuable tool for public health decision-makers.

Fajardo-Fontiveros, Oscar (ORCID:0000000207058972)

2013 New Mexico Mid-Region Travel Survey

The 2013 New Mexico Mid-Region Travel Survey assessed travel behavior patterns to update a travel demand model for the Albuquerque Metropolitan Planning Area, which consists of Bernalillo County, Valencia County, and southern Sandoval County. It includes the cities of Albuquerque, Rio Rancho, Los Lunas, and Belen as well as some tribal lands. The Mid-Region Council of Governments contracted with Westat to conduct the survey, which included the collection of socio-demographic data and a one-day (24-hour) period of household travel behavior collected during weekdays (Monday through Friday). The survey also included a random selection of a 20% subsample of households (1,023 participants) to take part in a wearable global positioning system, technology-based component of the study, which was used to assess the level of trip under-reporting from the self-reported component of the survey.

1Hz data

2014 Southern Nevada Household Travel Survey

The 2014 Southern Nevada Household Travel Survey collected information from residents in the Las Vegas area to update the regional travel demand model and assess travel behavior. The Regional Transportation Commission of Southern Nevada contracted with Westat to conduct the survey. The survey was conducted in two phases—from March to May 2014 and from August to October 2014. Participants provided demographic and travel data (via a travel log). A 10% subsample (1,694 participants) was randomly selected to take part in a wearable global positioning system (GPS) technology-based component of the study, the purpose of which was to assess the level of trip under-reporting in the self-reported travel logs. Participants in the GPS portion of the study were instructed to record trips in their travel logs on the first day only, while passively recording their travel for three full days.

1Hz data