Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Towards retrieving cloud top entrainment velocities from MISR cloud motion vectors

Although important, direct retrievals of entrainment rates in cloud-topped planetary boundary layer (PBL) remain elusive. Here we present a novel technique for retrieving cloud-top entrainment velocities using only Multi-angle Imaging Spectro-Radiometer (MISR) stereoscopic retrievals of cloud-motion vectors (CMVs) and cloud-top heights (CTHs). Mesoscale vertical air velocity at CTH is diagnosed from the continuity equation and then used to derive entrainment velocities from the PBL mass-budget equation. The algorithm is demonstrated through a case of marine stratocumulus deck off the California coast, with comparisons made against data from the European Centre for Medium-range Weather Forecasts (ECMWF) reanalysis (ERA5) and the data from other satellites. MISR low-cloud CTH for this case were lower than the ERA5 reported PBL depth by 189 ± 87 m. These differences in cloud top heights partly modulate the differences in the ERA5 and MISR horizontal winds, with larger differences in meridional over zonal wind components. Average difference between ERA5 and MISR derived mesoscale vertical air motion at cloud top was 0.14 ± 0.73 cm s −1 , while the same for entrainment rate was −0.09 ± 0.46 cm s −1 . The uncertainties in the utilized CTHs and CMVs are propagated to derive systematic and random retrieval uncertainties. Fractional uncertainty is lower than 25 % when the retrieved mesoscale vertical air motion is stronger than ±0.04 cm s −1 and entrainment velocities are stronger than ±0.03 cm s −1 . These results showcase the ability to derive mesoscale vertical air motion and entrainment rates from MISR observations and motivate its extension to generate a global climatology leveraging its full 23-year record (2000–2022). Nonetheless comprehensive validation of the retrievals is warranted through comparisons with estimates from an independent dataset across diverse weather conditions.

Mitra, Arka [Argonne National Laboratory (ANL), Ar↗

Over three decades, and counting, of near-surface turbulent flux measurements from the Atmospheric Radiation Measurement (ARM) user facility

Processes mediating the coupling of terrestrial, aquatic, biospheric, and atmospheric systems influence weather, climate, and ecosystem dynamics via transfer of energy, momentum, water, and carbon (or other species). These exchange processes are quantified by measurements of near-surface turbulent fluxes. Understanding processes at these interfaces provides insight toward understanding and predicting current and future states within the Earth system. The Atmospheric Radiation Measurement (ARM) user facility has been conducting measurements of near-surface turbulent fluxes since the early 1990s at long-term fixed locations and shorter-term mobile deployments across the Earth. ARM has utilized two established methods for conducting these measurements: energy balance Bowen ratio (EBBR) and eddy covariance (EC). Primary measurements from the former include sensible and latent heat flux, while the latter also measures fluxes of momentum and carbon (primarily carbon dioxide, with methane fluxes measured at two locations to date). The EBBR systems have been deployed at 22 locations, and, to date, the EC systems have been deployed at over 50 sites, with plans for additional novel site locations in the future. Herein, the history, evolution, and key aspects of these instrument systems are documented, along with information on data quality assurance and post-processing, as well as best use practices. Additionally, three data validation experiments were recently conducted, and their key findings are summarized. Finally, ancillary datasets acquired by ARM, which can contextualize and aid interpretation of the near-surface turbulent flux measurements, are discussed. The datasets described herein include the eddy correlation flux measurement system: 30ECOR (https://doi.org/10.5439/1879993, Sullivan et al., 1997), 30QCECOR (https://doi.org/10.5439/1097546, Gaustad, 2003), ECORSF (https://doi.org/10.5439/1494128, Sullivan et al., 2019a), and associated AmeriFlux and Methane Value-Added Product, AMCMETHANE (https://doi.org/10.5439/1508268, Billesbach, 2011); the energy balance Bowen ratio system: 30EBBR (https://doi.org/10.5439/1023895, Sullivan et al., 1993) and 30BAEBBR (https://doi.org/10.5439/1027268, Gaustad and Xie, 1993); and the carbon dioxide flux measurement system: CO2FLX (https://doi.org/10.5439/1287574, https://doi.org/10.5439/1287575, https://doi.org/10.5439/1287576, Koontz et al., 2015a, b, c; https://doi.org/10.5439/1989774, https://doi.org/10.5439/1989776, https://doi.org/10.5439/1992202, Biraud and Chan, 2002a, b, c). These data can be found by searching the above data stream names at https://adc.arm.gov/discovery/#/results/ (last access: 8 September 2025).

Sullivan, Ryan C. [Argonne National Laboratory (AN↗

Physics-coupled data-driven design of high-temperature alloys

We present a materials design loop, which streamlines physics-coupled machine learning (ML) surrogate models to discover new alloy chemistries with improved properties. The efficacy is demonstrated by discovering a high-temperature alumina-forming austenitic (AFA) stainless steel with enhanced creep, followed by experimental validation. The ML models have been trained using a well-curated, highly consistent experimental dataset augmented with synthetic microstructural features from a computational thermodynamic approach. We have populated a large number of hypothetical AFA alloys to explore the high-dimensional composition space and have predicted their creep properties by providing the same synthetic input features obtained from the trained ML models. Uncertainties from the ML training were taken as thresholds for truncating predicted results to identify alloys with improved or deteriorated creep. Individual elemental compositions have been determined via probability density distribution analysis from the group of alloys at the top and bottom of the predicted creep values for further virtual and experimental validations. In conclusion, we anticipate that this workflow can be applied to screen desired conditions, such as chemistry and processing parameters, in high-dimensional space through physics-guided data analytics.

Alloy design↗

Microstructural Topology as a Prescriptor for Quantum Coherence: Towards A Unified Framework for Decoherence in Superconducting Qubits

In superconducting quantum circuits, decoherence improvements are frequently obtained through process interventions that simultaneously modify surface chemistry, microstructural topology, and device geometry, leaving mechanistic attribution structurally underdetermined. Predictive materials engineering requires measurable structural statistics to be separated from geometry-dependent coupling coefficients into independently testable factors. We introduce the concept of classical and quantum microstructure. In that context, we formulate a channel-wise separable framework for decoherence in superconducting transmon qubits in which each loss channel is described by a reduced prescriptor. Here, a channel-specific microstructural state variable is determined independently of device geometry, and a geometry-dependent coupling functional is computable from field solutions without reference to surface chemistry. We derive this product form from a spatially resolved kernel representation and establish a perturbative separability criterion that defines the regime where independent variation of the variables is valid. The framework specifies five prescriptor classes for dominant loss pathways in transmon-class devices. Falsifiability is operationalized through a pre-committed 2x2 experimental protocol in which the variables must satisfy independent ratio checks within propagated uncertainty. A Minimum-Dataset Specification standardizes reporting for cross-laboratory inference. Part I establishes the conceptual and mathematical architecture; coordinated experimental validation is reserved for Part II.

Dravid, Vinayak P. [Northwestern U.]↗

Development and transferability of neural-network models for plasma-surface interactions

Plasma-surface interactions are increasingly critical to modern technologies; yet, accurate molecular dynamics simulations remain limited by the capabilities of interatomic potentials. Deep Potentials (DPs) promise to revolutionize the field by providing a systematic method for producing accurate interatomic potentials. The primary challenge of DP development is selecting a dataset, which efficiently spans the set of atomic environments one expects to encounter in the subsequent molecular dynamics simulations. The computational cost of density functional theory calculations, which are the typical basis for DP development, makes it impossible to directly verify the quality of a given DP. To address this challenge, we explore the development of a deep-learned interatomic potential, “DeepREBO,” trained to reproduce the behavior of the REBO2 empirical potential, enabling direct validation of training methodology and transferability. Using an active learning framework, we begin with a minimal dataset and iteratively expand it to train a Deep Potential-Smooth Edition model that faithfully reproduces REBO2 results for 25 eV hydrogen bombardment of diamond (001), a particularly challenging case. We show that small, carefully curated datasets can outperform large, unguided ones, with effective models requiring fewer than 15 000 snapshots. Subsequent transferability tests demonstrate that while DeepREBO generalizes well to diamond (111) surfaces, performance degrades for amorphous carbon or higher-energy impacts, highlighting the need for use-case-specific training data. We also evaluate methods to improve short-range repulsion. This study outlines best practices for training robust deep potentials and underscores the importance of dataset design for predictive plasma simulations.

Ab-initio molecular dynamics↗

Foundational Dataset for Developing Large-Sample Stream Temperature Models in the Conterminous United States

This dataset provides inputs, evaluation results, and trained weights from a large-sample Long Short-Term Memory (LSTM) model designed to predict daily stream temperatures across unregulated river reaches in the conterminous United States (CONUS). It includes dynamic meteorological and hydrologic forcings, static physiographic attributes, and model outputs from cross-validation experiments spanning 300 basins. It supports reproducible modeling, direct application for new basins, and provides data suitable for integration with reservoir and river simulations under current and future climates. It contains two .zip files described below · RQ-AI_runs.zip: Model outputs from 10-fold cross-validation experiments, including observed and predicted daily stream temperatures, along with test performance metrics for water years 2017–2019. Two versions are included: 1. Model trained and validated using subbasin-area weighted dynamic features. 2. Model trained and validated using whole-basin area weighted dynamic features. · RQ-AI_inputs.zip: Collection of all formatted dynamic and static predictor datasets (meteorological, hydrologic, and physiographic features) used in model training and analysis. Detailed instructions and data structure is held at the following GitLab repository: https://code.ornl.gov/tempwise/training.

Gomez-Velez, Jesus [Oak Ridge National Laboratory ↗

Geometry-complete diffusion for 3D molecule generation and optimization

Abstract Generative deep learning methods have recently been proposed for generating 3D molecules using equivariant graph neural networks (GNNs) within a denoising diffusion framework. However, such methods are unable to learn important geometric properties of 3D molecules, as they adopt molecule-agnostic and non-geometric GNNs as their 3D graph denoising networks, which notably hinders their ability to generate valid large 3D molecules. In this work, we address these gaps by introducing the Geometry-Complete Diffusion Model (GCDM) for 3D molecule generation, which outperforms existing 3D molecular diffusion models by significant margins across conditional and unconditional settings for the QM9 dataset and the larger GEOM-Drugs dataset, respectively. Importantly, we demonstrate that GCDM’s generative denoising process enables the model to generate a significant proportion of valid and energetically-stable large molecules at the scale of GEOM-Drugs, whereas previous methods fail to do so with the features they learn. Additionally, we show that extensions of GCDM can not only effectively design 3D molecules for specific protein pockets but can be repurposed to consistently optimize the geometry and chemical composition of existing 3D molecules for molecular stability and property specificity, demonstrating new versatility of molecular diffusion models. Code and data are freely available on GitHub .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The DECam MAGIC Survey $-$ Mapping the Ancient Galaxy in CaHK: Overview and Summary of Early Science

We present the DECam Mapping the Ancient Galaxy in CaHK (MAGIC) survey, a 54-night NOIRLab Survey Program to image $\gtrsim$5,000$\,$deg$^2$ of the southern hemisphere using a metallicity-sensitive narrow-band filter covering the Ca$\,$ii$\,$H&K lines centered at 3955$\,$A. This filter is installed on the Dark Energy Camera (DECam), mounted on the 4-m NSF Víctor M. Blanco Telescope. The survey reaches typical $10σ$ depths of $\text{mag}_{\text{CaHK}} \approx 22.5$, 3$-$4$\,$mag deeper than comparable surveys in the southern hemisphere. By combining photometry from this Ca$\,$ii$\,$H&K filter with existing DECam $g,r,i$ broadband photometry from the DECam Local Volume Exploration (DELVE) survey, MAGIC is deriving photometric metallicities for red giant branch stars down to the magnitude limit of usable proper motions from Gaia data release 3 (DR3). MAGIC has already imaged $\sim$3,000$\,$deg$^2$, supplemented by other affiliated observing programs that have used this filter to image star clusters, dwarf galaxies, and stellar streams. We overview MAGIC's survey strategy, describe data processing through the derivation of metallicities and photometric distances, and summarize early science results that have been published with this dataset. In addition, we present several new results, including the confirmation of a distant ($>5\,r_h$) member of the Reticulum II ultra-faint dwarf galaxy, on-sky density maps of low-metallicity stars into the distant Milky Way halo ($\sim150\,$kpc) recovering 13/14 ultra-faint dwarf galaxies in the current footprint, and a validation of our initial targeting of extremely metal-poor stars. Collectively, these results demonstrate that the MAGIC dataset enables cutting-edge studies of the faint, low-metallicity regime of the Milky Way and its substructures.

Chiti, A. [KIPAC, Menlo Park] (ORCID:0000000271556↗

Estimates of Southern Hemispheric Gravity Wave Momentum Fluxes across Observations, Reanalyses, and Kilometer-Scale Numerical Weather Prediction Model

Abstract Gravity waves (GWs) are among the key drivers of the meridional overturning circulation in the mesosphere and upper stratosphere. Their representation in climate models suffers from insufficient resolution and limited observational constraints on their parameterizations. This obscures assessments of middle atmospheric circulation changes in a changing climate. This study presents a comprehensive analysis of stratospheric GW activity above and downstream of the Andes from 1 to 15 August 2019, with special focus on GW representation ranging from an unprecedented kilometer-scale global forecast model (1.4 km ECMWF IFS), ground-based Rayleigh lidar (CORAL) observations, modern reanalysis (ERA5), to a coarse-resolution climate model (EMAC). Resolved vertical flux of zonal GW momentum (GWMF) is found to be stronger by a factor of at least 2–2.5 in IFS compared to ERA5. Compared to resolved GWMF in IFS, parameterizations in ERA5 and EMAC continue to inaccurately generate excessive GWMF poleward of 60°S, yielding prominent differences between resolved and parameterized GWMFs. A like-to-like validation of GW profiles in IFS and ERA5 reveals similar wave structures. Still, even at ∼1 km resolution, the resolved waves in IFS are weaker than those observed by lidar. Further, GWMF estimates across datasets reveal that temperature-based proxies, based on midfrequency approximations for linear GWs, overestimate GWMF due to simplifications and uncertainties in GW wavelength estimation from data. Overall, the analysis provides GWMF benchmarks for parameterization validation and calls for three-dimensional GW parameterizations, better upper-boundary treatment, and vertical resolution increases commensurate with increases in horizontal resolution in models, for a more realistic GW analysis. Significance Statement Gravity wave–induced momentum forcing forms a key component of the middle atmospheric circulation. However, complete knowledge of gravity waves, their atmospheric effects, and their long-term trends are obscured due to limited global observations, and the inability of current climate models to fully resolve them. This study combines a kilometer-scale forecast model, modern reanalysis, and a coarse-resolution climate model to first compare the resolved and parameterized momentum fluxes by gravity waves generated over the Andes, and then evaluate the fluxes using a state-of-the-art ground-based Rayleigh lidar. Our analysis reveals shortcomings in current model parameterizations of gravity waves in the middle atmosphere and highlights the sensitivity of the estimated flux to the formulation used.

Meteorology & Atmospheric Sciences↗

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David↗

From Latent Dynamics to Meaningful Representations

While representation learning has been central to the rise of machine learning and artificial intelligence, a key problem remains in making the learnt representations meaningful. For this the typical approach is to regularize the learned representation through prior probability distributions. However such priors are usually unavailable or are ad hoc. To deal with this, recent efforts have shifted towards leveraging the insights from physical principles to guide the learning process. In this spirit, we propose a purely dynamics-constrained representation learning framework. Instead of relying on predefined probabilities, we restrict the latent representation to follow overdamped Langevin dynamics with a learnable transition density — a prior driven by statistical mechanics. We show this is a more natural constraint for representation learning in stochastic dynamical systems, with the crucial ability to uniquely identify the ground truth representation. We validate our framework for different systems including a real-world fluorescent DNA movie dataset. Here, we show that our algorithm can uniquely identify orthogonal, isometric and meaningful latent representations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient Probabilistic Visualization of Local Divergence of 2D Vector Fields with Independent Gaussian Uncertainty

This work focuses on visualizing uncertainty of local divergence of two-dimensional vector fields. Divergence is one of the fundamental attributes of fluid flows, as it can help domain scientists analyze potential positions of sources (positive divergence) and sinks (negative divergence) in the flow. However, uncertainty inherent in vector field data can lead to erroneous divergence computations, adversely impacting downstream analysis. While Monte Carlo (MC) sampling is a classical approach for estimating divergence uncertainty, it suffers from slow convergence and poor scalability with increasing data size and sample counts. Thus, we present a two-fold contribution that tackles the challenges of slow convergence and limited scalability of the MC approach. (1) We derive a closed-form approach for highly efficient and accurate uncertainty visualization of local divergence, assuming independently Gaussian-distributed vector uncertainties. (2) We further integrate our approach into Viskores, a platform-portable parallel library, to accelerate uncertainty visualization. In our results, we demonstrate significantly enhanced efficiency and accuracy of our serial analytical (speed-up up to 1946×) and parallel Viskores (speed-up up to 19698×) algorithms over the classical serial MC approach. We also demonstrate qualitative improvements of our probabilistic divergence visualizations over traditional mean-field visualization, which disregards uncertainty. We validate the accuracy and efficiency of our methods on wind forecast and ocean simulation datasets.

Ouermi, Timbwaoga [University of Utah]↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

EXCLUSIVE NEUTRAL PION ELECTROPRODUCTION CROSS SECTION MEASUREMENTSWITHANEUTRALPARTICLE SPECTROMETER

Deep Virtual Compton Scattering (DVCS), the exclusive electron-proton scattering process ep ¿e'p'¿, provides access to generalized parton distributions (GPDs), which correlate information about the longitudinal momentum and transverse spatial structure of quarks inside the nucleon. Experiment E12-13-010 in Hall C at Jefferson Lab was designed to take high-precision measurements of the DVCS cross section over an extended kinematic range using the newly commissioned Neutral Particle Spectrometer (NPS). The NPS features a high-resolution electromagnetic calorimeter and a streaming data acquisition system optimized for operation at high luminosities. This thesis presents the detector and analysis work carried out to support the NPS DVCS program. In particular, it focuses on the hardware design, calibration, and performance of the calorimeter. A development of a waveform reconstruction analysis of the calorimeter signals enabled improved extraction of pulse amplitudes and times. The waveform analysis was also extended to operate in a multithreaded environment, substantially reducing processing time for large datasets. Analysis of exclusive neutral pion electroproduction events in the calorimeter gives a strong validation of the calorimeter’s performance and resolution. Together these developments establish a foundation for future analyses and extraction of the DVCS cross section and its use in constraining the GPDs.

Kerver, Mitchell [Old Dominion Univ., Norfolk, VA ↗

Focused Ion Beam Tomography of Alloy 617 Corroded in Molten Chloride Salt

Materials qualification of reactor structural materials is a critical step in rapid implementation of advanced nuclear reactor technologies, particularly to assess the corrosion performance in these designs. Accelerated qualification of reactor structural materials requires incorporating powerful computational toolsets, such as phase field modelling in the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, to predict the evolution of structural materials due to corrosion. Accordingly, computational toolsets will require experimental data generated at appropriate length scales to validate accuracy. Focused ion beam (FIB) provides a high degree of control over manipulation of materials for analytical purposes, including capturing data on the evolution in the microstructure and elemental composition of materials at the mesoscale, an appropriate length scale for phase field modelling of intergranular diffusion phenomena using the MOOSE framework. For instance, the FEI Helios G4 UX dual beam plasma FIB microscope at the Irradiated Materials Characterization Laboratory (IMCL) is capable of backscatter diffraction (EBSD) and energy-dispersive x-ray spectroscopy (EDS) documenting the evolution in the microstructure and elemental composition, respectively. The Helios can perform EDS and EBSD three-dimensionally (3D) using tomography, which is then combined using different software packages to visualize 3D volumes correlating elemental composition to microstructural data. The purpose of this investigation was to develop a streamlined characterization and data processing workflow for 3D tomography studies on the FEI Helios G4 plasma FIB. The investigation is segmented into three parts: 1) Optimizing the data collection workflow, 2) identifying appropriate data processing and visualization software (i.e. DREAM.3D, MIPAR, and VGStudioMax), and 3) establishing an infrastructure for public release. The optimization of the data collection workflow is in collaboration with members of the U220 department to setup formal training on the tomography operation of the G4, through ThermoFisher Scientific, and exploring DREAM.3D, MIPAR, and VGStudioMax data processing/visualization software packages. VGStudioMax currently demonstrates the most promise for future use. Optimization of the data collection and processing workflow is still ongoing. A collaboration with INL High Performance Computing (HPC) established an open-source license for expediting the public release of FIB tomography datasets through HPC. FIB tomography data generated by the G4 will provide comprehensive data for validating 3D phase field mesoscale modelling tools within the MOOSE framework for accelerated qualification of reactor structural materials.

Copeland-Johnson, Trishelle↗

An open retail boundary dataset for South Korea using open data and computer vision technique

Although delineating retail boundaries is important to explore and comprehend the dynamics of the retail sector, it is hard to find studies specifically addressing it in the South Korean context. This study fills this gap by proposing new retail boundaries across South Korea. To achieve this goal, we employed a variety of retailers and building datasets and proposed a unique computer vision-based framework with a deep ensemble voting technique. As a result, we delineated 6,636 distinct retail boundaries that were validated against existing reference retail boundaries. These newly delineated retail boundaries provide valuable insights for researchers, governments, and other relevant stakeholders by enhancing their understanding of retail geography. This dataset can be used as a foundational resource for analyses on topics such as pandemic recovery, retail gentrification, and the resilience of retail spaces in response to e-commerce growth, ultimately contributing to more robust retail sector research in South Korea.

97 MATHEMATICS AND COMPUTING↗

Using Temporal Information from Human Mobility Data to Detect Anchor Points

Spatiotemporal mobility data are available in massive quantities, but large quantities of data typically include fewer variables or data fields. Often, the only available fields are User ID, Longitude, Latitude, Timestamp (ULLT). This raises an important question: how much can we infer about human mobility patterns using only these four fields? With ULLT data, we do not know individuals' socioeconomic status information or when they are visiting their anchor points (AP) or locations (such as homes, places of employment, or schools), and it is a modern challenge to use this data to infer these characteristics. When detecting anchor locations with limited input information, verification and validation (VV) are significant challenges. This paper addresses the problem of identifying individuals' anchor locations using only temporal information from spatiotemporal datasets with limited attributes. Our approach does not explicitly use latitude and longitude during analysis. Locationbased information is only employed in the preprocessing stage to identify periods of movement (trips) and stops (dwelling). Beyond this step, all analysis is based on temporal patterns. In theory, if stops and dwell times could be detected through alternative means, our method could function entirely without location-based input. We demonstrate this methodology on the 2017 National Household Travel Survey (NHTS) data, because it includes a carefully designed and collected time use survey with representative sampling and labeled ground truth. The high-quality survey data allows us to test the accuracy of our methods because NHTS contains intended place labels and agent/user characteristics. We have also applied our validated AP identification algorithm on very large-scale GPS based trajectory data for Patterns-of-Life (PoL) assessment and other applications, but due to space limit that could not be presented here.

McBride, Liz [ORNL] (ORCID:0000000286925869)↗

Predicting cutoff L-shells of solar protons using the GPPSn particle dataset

Solar energetic protons (SEPs) arriving at the Earth trigger severe radiation storms in the near-Earth space, directly impacting space missions operating at various altitudes. Therefore, monitoring SEP events and predicting the penetration depths of solar protons are critical for aerospace sectors. Building on previous efforts, here we demonstrate the feasibility of using proton measurements from the Global Prompt Proton Sensor network (GPPSn), enabled by Los Alamos National Laboratory developed combined X-ray dosimeters aboard GPS satellites, to characterize and predict the penetration of solar protons into the geomagnetic field. The inclined medium-Earth-orbits (MEOs) of the global GPS constellation offer a unique advantage of allowing simultaneous measurements of penetrating solar protons inside both open- and closed-field line regions. Therefore, the L-profiles of ∼10s–100 MeV solar protons and their associated cutoff L-shells can be determined from the GPPSn dataset, using predefined threshold proton flux values rather than traditional flux ratios. After examining a list of SEP event intervals across solar cycles 23, 24 and 25—including the 2024 Mother’s Day superstorm, we showcase how the latest GPPSn proton dataset (release v1.10), reprocessed and calibrated, can not only be used to monitor solar proton distributions inside the dynamic geomagnetic field for individual events, but also to derive a new empirical model linking cutoff L-shells with several key space weather parameters. This newly developed SEPCL-MEO model demonstrates high predictive performance; for example, predictions for > 30 MeV solar protons yield a correlation coefficient of 0.85 and performance efficiency of 0.67 when validated against GPPSn observations. Results from this pilot study underscores the scientific and operational value of the GPPSn dataset, and this dataset—when paired with machine-learning techniques—can play a critical role in observing and predicting the effects of future incoming SEP events, including extreme ones.

58 GEOSCIENCES↗