Search NASA⌕ Search

SEARCH · Search NASA

Results for “data volume”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

O'Hare Airport roadway traffic prediction via data fusion and Gaussian process regression

This study proposes an approach of leveraging information gathered from multiple traffic data sources at different resolutions to obtain approximate inference on the traffic distribution of Chicago's O'Hare Airport area. Specifically, it proposes the ingestion of traffic datasets at different resolutions to build spatiotemporal models for predicting the distribution of traffic volume on the road network. Due to its good adaptability and flexibility for spatiotemporal data, the Gaussian process (GP) regression was employed to provide short-term forecasts using data collected by loop detectors (sensors) and supplemented by telematics data. The GP regression is used to make predictions of the distribution of the proportion of sensor data traffic volume represented by the telematics data for each location of the sensors. Consequently, the fitted GP model can be used to determine the approximate traffic distribution for a testing location outside of the training points. Policymakers in the transportation sector can find the results of this work helpful for making informed decisions relating to current and future transportation conditions in the area.

42 ENGINEERING↗

A Scalable Multi-Modal Framework for High-Fidelity Distributed Human Mobility Simulations

The development of data-driven models for human mobility in urban settings requires access to substantial and diverse real-world data. However, existing historical data often presents challenges such as limited volume, variety, and veracity, as well as missing data and privacy preservation concerns. Also, urban mobility modeling is inherently time-variant, complex, and multi-modal, encompassing everything from individual walking and running to private road travel and large-scale public transportation. These challenges call for innovative solutions to overcome data limitations and compute needs to model mobility behaviors accurately. To address these challenges, we propose a distributed, co-simulation-based architecture DURMOSim that integrates real-world data with scalable, high-fidelity simulations, demonstrating distributed co-simulation feasibility with existing mobility models. DURMOSim underpins a modular integration that would enable using any available mobility simulators for greater extensibility and scalability in performing various urban scenarios. In this paper, we present the design, implementation, and performance evaluation of DURMOSim, highlighting its capability to model population-scale mobility patterns. Our initial results show its ability to dynamically synchronize multiple simulation models at runtime with negligible computational overhead. We believe DURMOSim could be a robust tool for advancing urban mobility research and intelligent transportation systems.

Yoginath, Srikanth [ORNL] (ORCID:0000000184236050)↗

ACE-ENA: Fast Liquid Water Content

These data were collected during the Aerosol and Cloud Experiments in the Eastern North Atlantic field campaign as part of ARM Aerial Facility deployment (ACE-ENA, https://www.arm.gov/research/campaigns/aaf2017ace-ena). The ARM Aerial Facility Gulfstream-1 was deployed at Lajes Air Base (IATA: TER, ICAO: LPLA), on Terceira Island in the Azores, Portugal, for the two Intensive Observation Periods from June 20 through July 22, 2017 (IOP#1) and from January 11 through February 22, 2018 (IOP#2). The G-1 aircraft performed 20+19 research flights over the ARM Eastern North Atlantic (ENA) site and Atlantic Ocean to measure atmospheric turbulence, cloud water content and drop size distributions, aerosol precursor gases, aerosol chemical composition and size distributions. The current data set presents re-processed Particle Volume Monitor PVM-100A (aka Gerber probe) data: Liquid Water Content (LWC), Particle Surface Area (PSA), and the effective droplet radius (re) averaged to 50 Hz, 10 Hz, and 1 Hz.

54 ENVIRONMENTAL SCIENCES↗

CACTI: Fast Liquid Water Content

These data were collected during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI; https://www.arm.gov/research/campaigns/amf2018cacti ) field campaign in the Sierras de Córdoba mountain range of north-central Argentina as part of the ARM Aerial Facility (AAF) deployment. The ARM Aerial Facility Gulfstream-1 was operated from Las Higueras Airport (IATA: RCU, ICAO: SAOC), Río Cuarto, Córdoba, Argentina, for the Intensive Observation Period (IOP) from Nov. 1 through Dec. 15, 2018. The G-1 aircraft performed 22 research flights over the first ARM Mobile Facility (AMF1) location in the Sierras de Córdoba mountain range to measure atmospheric state and turbulence, cloud water content and droplet size distributions, aerosol precursor gases, and aerosol chemical composition and size distributions. The current data set presents re-processed Particle Volume Monitor PVM-100A (aka Gerber probe) data: Liquid Water Content (LWC), Particle Surface Area (PSA), and the effective droplet radius (re) averaged to 50Hz, 10Hz, and 1Hz.

54 ENVIRONMENTAL SCIENCES↗

Energy-resolved neutron imaging and diffraction including grain orientation mapping using event camera technology

Time-of-flight neutron diffraction and energy-resolved imaging each provide unique perspectives into material properties. Neutron diffraction is useful for assessing microstructural parameters such as phase composition, texture, and dislocation densities, though it typically provides averaged data over the sampled volume. Energy-resolved imaging, on the other hand, offers both spatial and spectral information by detecting Bragg edges and neutron absorption resonances, which enables detailed mapping of microstructure and isotopic composition. When combined, these techniques have the potential to enrich our understanding of material behavior across different scales, enhancing our understanding of complex materials. Traditionally, these modalities are conducted on separate instruments, which is time-consuming and poses challenges for data integration. Here, we report the integration of the LumaCam, an event-mode energy-resolved neutron imaging camera with the HIPPO time-of-flight diffractometer at LANSCE. This integration enables simultaneous diffraction and imaging across the full spectrum, with analysis optimized for diffraction and Bragg-edge imaging in the thermal range (0.45–10 Å) and resonance imaging in the epithermal range (0.5–3000 eV), facilitating comprehensive multi-modal analysis. We demonstrate its capabilities through case studies, including spatial mapping of grain orientations in a steel sample and accurate thickness estimations for irregular samples including a depleted uranium cylinder and a natural silver-containing mineral specimen. The combined setup enhances real-time sample alignment and provides comprehensive data for crystal structure, texture, and isotopic composition analysis. This approach opens new possibilities for advanced applications in nuclear engineering, archaeology, and materials science.

36 MATERIALS SCIENCE↗

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning↗

DESIVAST: Catalogs of Low-redshift Voids Using Data from the DESI Data Release 1 Bright Galaxy Survey

We present three separate void catalogs created using a volume-limited sample of the DESI Data Release 1 Bright Galaxy Survey. We use the algorithms VoidFinder and V 2 to construct void catalogs out to a redshift of z = 0.24. Excluding voids affected by the boundaries of the survey, we obtain 1489 voids with VoidFinder, 389 with V 2 using REVOLVER pruning, and 297 with V 2 using VIDE pruning. Comparing our catalogs with overlapping Sloan Digital Sky Survey void catalogs, we find generally consistent void properties but significant differences in the void volume overlap, which we attribute to differences in the galaxy selection and survey masks. These catalogs are suitable for studying the variation in galaxy properties with cosmic environment and for cosmological studies.

79 ASTRONOMY AND ASTROPHYSICS↗

The DECADE cosmic shear project I: A new weak lensing shape catalog of 107 million galaxies

We present the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. This catalog was assembled from public DECam data including survey and standard observing programs. These data were consistently processed with the Dark Energy Survey Data Management pipeline as part of the DECADE campaign and serve as the basis of the DECam Local Volume Exploration survey (DELVE) Early Data Release 3 (EDR3). We apply the Metacalibration measurement algorithm to generate and calibrate galaxy shapes. After cuts, the resulting cosmology-ready galaxy shape catalog covers a region of $5,\!412 \,\,{\rm deg}^2$ with an effective number density of $4.59\,\, {\rm arcmin}^{-2}$. The coadd images used to derive this data have a median limiting magnitude of $r = 23.6$, $i = 23.2$, and $z = 22.6$, estimated at ${\rm S/N} = 10$ in a 2 arcsecond aperture. We present a suite of detailed studies to characterize the catalog, measure any residual systematic biases, and verify that the catalog is suitable for cosmology analyses. In parallel, we build an image simulation pipeline to characterize the remaining multiplicative shear bias in this catalog, which we measure to be $m = (-2.454 \pm 0.124) \times10^{-2}$ for the full sample. Despite the significantly inhomogeneous nature of the data set, due to it being an amalgamation of various observing programs, we find the resulting catalog has sufficient quality to yield competitive cosmological constraints.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DyG-DPCD: A Distributed Parallel Community Detection Algorithm for Large-Scale Dynamic Graphs

Dynamic (Temporal) graphs capture the valuable evolution of real-world systems, from the continuously evolving patterns of social interactions and genetic pathways to the dynamic fluctuations of economic forces. Detecting communities for such evolving networks poses unique challenges. Detecting and analyzing the evolution of communities within dynamic graphs unlocks valuable insights into the underlying structural and temporal patterns of real-world systems. However, the sheer volume of modern graph data and the inherent complexity of the temporal dimension pose significant challenges to scalable community detection algorithms. Addressing this gap, our work explores the limited landscape of scalable distributed-memory parallel methods specifically designed for dynamic network community detection. We propose a novel parallel algorithm, DyG-DPCD (Dynamic Graph Distributed Parallel Community Detection), to detect communities in dynamic networks using the Message Passing Interface (MPI) framework. We present a vertex-centric approach, allowing us to detect communities through local optimization. Furthermore, we enhance our baseline algorithm by incorporating three heuristics, which improve the algorithm’s performance significantly while maintaining the quality of the solutions. We demonstrate the efficiency of our algorithm by experimenting on several real-world large-scale networks with hundreds of millions of edges spanning diverse domains. Notably, DyG-DPCD achieves speedups between 25× and 30× for large networks that we experimented on using NERSC compute nodes. In conclusion, our algorithm outperforms the STINGER parallel re-agglomeration algorithm by 30×.

97 MATHEMATICS AND COMPUTING↗

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)↗

A unified large language model–based framework for heterogeneous PV image diagnosis

With advances in imaging technologies, modern photovoltaic (PV) systems generate large volumes of heterogeneous image data, including visible, electroluminescence (EL), and infrared (IR) images. Existing PV image analysis models, particularly deep learning approaches, are typically task-specific and lack cross-modality generalization. To address this limitation, this paper proposes an open-source large language model (LLM)–based unified framework for heterogeneous PV image diagnostics. Through task-aware diagnostic prompting, the framework enables analysis of visible, EL, and IR images within a single pipeline, supporting both zero-shot and few-shot inference and binary and multiclass classification. It is compatible with state-of-the-art multimodal LLMs, including ChatGPT, Gemini, Claude, Qwen, and CLIP. The framework is evaluated on PV module condition classification (clean, soiling, snow, hail, and bird droppings) using visible images, cell crack detection using EL images, and hotspot detection using IR images. GPT-5.1 in few-shot mode achieves the best performance, with classification accuracy exceeding 97.3%. Open-source models such as Qwen and CLIP also deliver competitive results on visible images (around 90% accuracy), though their performance is more limited on EL and IR modalities. On the full ELPV dataset, the framework achieves 83.5% zero-shot accuracy, within 2.8% of the supervised CNN baseline, confirming scalability to larger benchmarks. Practical aspects such as reproducibility, response latency, and confidence estimation are systematically analyzed. The framework operates across PV image modalities without modality- or task-specific training, making it well suited as a rapid pre-screening tool to support downstream detailed diagnostics. A benchmark dataset of diverse labeled PV images is also released.

Li, Baojie↗

Low energy neutron light output characterization of EJ301D and deuterated stilbene with a comparison of light output characterization methods

The neutron-induced light yield of a 2.54 cm diameter by 2.54 cm long right circular cylinder of EJ301D and a (5.08 cm)3 custom made cube of deuterated trans-stilbene-d12 (d-stilbene) were measured over incident neutron energies from 300 keV to 2.2 MeV and 200 keV to 2.4 MeV, respectively. The measurements were performed using a time-of-flight experiment with a Cf-252 source and an approximately 1.5 m flight path. We compare three light output spectrum full energy deposition edge estimation methods: (1) simulating the neutron energy spectrum edge and fitting it to the light output spectrum, (2) using the inflection point of the light output spectrum edge (derivative method, a.k.a. Kornilov’s method), and (3) using an empirical model fit to the edge of the light output spectrum. Both the derivative and equation fit methods do not account for physical processes such as multiple neutron scattering in the detectors. They instead rely on assumptions about the linear shape continuum shape of the light output spectrum and the direct correlation between the location of the spectrum’s inflection point and maximum energy deposition. These assumptions were found to introduce bias into those methods when tested against simulated spectra with known edge locations. When tested against measured spectra the derivative method was found to differ from the simulation fit by greater than 30% at low energies with large discontinuities for adjacent data points above 800 keV incident neutron energy. The empirical equation fitting method was found to also exhibit bias of a similar magnitude, but with significantly more continuous behavior, especially with the lower count data of the smaller volumed EJ301D scintillator. Experimental light output yield for this neutron energy range is reported using the simulated spectrum fitting method because it includes physics neglected by the other methods, and did not exhibit the bias observed in the other methods

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Brochure for the DOE Office of Science Workshop on Envisioning Frontiers in AI and Computing for Biological Research

In February of 2025 a joint ASCR/BER workshop was held to identify key transformational research directions for understanding biology using artificial intelligence (AI), digital twins and high-performance (HPC) computational methods to facilitate scientific discovery and innovation in support of the Department of Energy mission. AI technologies offer exciting new groundbreaking methods to analyze large volumes of complex biological data, thereby greatly accelerating the ability to understand, predict, and design biological processes for beneficial purposes. In the laboratory, the bridging of AI-enabled automated experimental technologies, HPC and digital twins will provide potent tools for researchers to explore the fundamental nature of biology and harness its inherent metabolic potential for a variety of beneficial purposes. The focus of this workshop was on how high-performance computational methods can impact this objective by exploring digital twins, foundational models, and data-driven approaches with applications to advance automated laboratory experiments, modeling of complex living systems and engineering new functions into plants and microbial systems relevant to DOE mission. Workshop attendees with expertise in plant science, microbiology, mathematics, computer science, and AI assessed the current state of the science, trends, and AI challenges at the interface of plant and microbial systems biology and computational science to identify opportunities for high-impact research. This collaborative effort capitalized on ASCR's advancements in applied mathematics, computer science, and Exascale systems, and BER's expertise in basic genomics-enabled research on DOE relevant plant and microbial systems. The workshop culminated in four key priority research directions to guide future research and development within DOE Office of Science programs.

59 BASIC BIOLOGICAL SCIENCES↗

A galactic approach to neutron scattering science

Neutron scattering science is leading to significant advances in our understanding of materials and will be key to solving many of the challenges that society is facing today. Improvements in scientific instruments are actually making it more difficult to analyze and interpret the results of experiments due to the vast increases in the volume and complexity of data being produced and the associated computational requirements for processing that data. New approaches to enable scientists to leverage computational resources are required, and Oak Ridge National Laboratory (ORNL) has been at the forefront of developing these technologies. We recently completed the design and initial implementation of a neutrons data interpretation platform that allows seamless access to the computational resources provided by ORNL. For the first time, we have demonstrated that this platform can be used for advanced data analysis of correlated quantum materials by utilizing the world's most powerful computer system, Frontier. In particular, we have shown the end-to-end execution of the DCA++ code to determine the dynamic magnetic spin susceptibility χ(q, ω) for a single-band Hubbard model with Coulomb repulsion U/t = 8 in units of the nearest-neighbor hopping amplitude t and an electron density of n = 0.65. The following work describes the architecture, design, and implementation of the platform and how we constructed a correlated quantum materials analysis workflow to demonstrate the viability of this system to produce scientific results.

97 MATHEMATICS AND COMPUTING↗

DELVE-ing into the Milky Way’s Globular Clusters: Assessing Extratidal Features in NGC 5897, NGC 7492, and Testing Detectability with Deeper Photometry

Extratidal features around globular clusters (GCs) are tracers of their disruption, stellar stream formation, and their host’s gravitational potential. However, these features remain challenging to detect due to their low surface brightness. We conduct a systematic search for such features around 19 GCs in the DECam Local Volume Exploration (DELVE) survey Data Release 2, discovering a new extra-tidal envelope around NGC 5897 and find tentative evidence for an extended envelope surrounding NGC 7492. Through a combination of dynamical modeling and analyzing synthetic stellar populations, we demonstrate these envelopes may have formed through tidal disruption. We use these models to explore the detectability of these features in the upcoming Legacy Survey of Space and Time (LSST), finding that while LSST’s deeper photometry will enhance detection significance, additional methods for foreground removal like proper motions or metallicities may be important for robust stream detection. Our results both add to the sample of globular clusters with extratidal features and provide insights on interpreting similar features in current and upcoming data.

Chiti, A. [Univ. of Chicago, IL (United States); S↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Combining Observations and Models: A Review of the CARDAMOM Framework for Data‐Constrained Terrestrial Ecosystem Modeling

The rapid increase in the volume and variety of terrestrial biosphere observations (i.e., remote sensing data and in situ measurements) offers a unique opportunity to derive ecological insights, refine process‐based models, and improve forecasting for decision support. However, despite their potential, ecological observations have primarily been used to benchmark process‐based models, as many past and current models lack the capability to directly integrate observations and their associated uncertainties for parameterization. In contrast, data assimilation frameworks such as the CARbon DAta MOdel fraMework (CARDAMOM) and its suite of process‐based models, known as the Data Assimilation Linked Ecosystem Carbon Model (DALEC), are specifically designed for model‐data fusion. This review, motivated by a recent CARDAMOM community workshop, examines the development and applications of CARDAMOM, with an emphasis on its role in advancing ecosystem process understanding. CARDAMOM employs a Bayesian approach, using a Markov Chain Monte Carlo algorithm to enable data‐driven calibration of DALEC parameters and initial states (i.e., carbon pool sizes) through observation operators. CARDAMOM's unique ability to retrieve localized model process parameters from diverse datasets—ranging from in situ measurements to global satellite observations—makes it a highly flexible tool for analyzing spatially variable ecosystem responses to environmental change. However, assimilating these data also presents challenges, including data quality issues that propagate into model skill, as well as trade‐offs between model complexity, parameter equifinality, and predictive performance. We discuss potential solutions to these challenges, such as reducing parameter equifinality by incorporating new observations. This review also offers community recommendations for incorporating emerging datasets, integrating machine learning techniques, strengthening collaboration with remote sensing, field, and modeling communities, and expanding CARDAMOM's relevance for localized ecosystem monitoring and decision‐making. CARDAMOM enables a deep, mechanistic understanding of terrestrial ecosystem dynamics that cannot be achieved through empirical analyses of observational datasets or weakly constrained models alone.

Bayesian inference↗

Peer Review of X441A Fission Gas Release Data

The purpose of this document is to demonstrate completion of the independent peer review for the fission gas release measurement of fuel pins irradiated in the Experimental Breeder Reactor-II (EBR-II) as part of the X441A experiment. The specific data reviewed includes reported plenum volume and pressure from the measurement from the fuel pins. Fraction percent of fission gas release is derived in part from these measurements but is excluded from the data reviewed.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗