Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

New constraint on the Np 237 ( n , γ ) Np 238 integral cross section using the Godiva-IV critical assembly

Accurate knowledge of the 237 Np(n, γ) 238 Np cross section at fast neutron energies is important for applied nuclear science. The presently available experimental data has large disagreements in the fast neutron region. Perform a model-independent measurement of the 237 Np(n, γ) 238 Np integral cross section using a well characterized fast neutron source and compare the result with previous measurements and current nuclear data evaluations. Provide an integral measurement that can be used as a benchmark for current evaluations. Multiple samples of 237 Np were irradiated in the Godiva-IV critical assembly. Following the irradiation, the samples placed in a γ-ray counting setup and the γ-rays emitted from the decay of 238 Np were measured over a time period of approximately 7 days. Multiple γ-ray decay branches of 238 Np were observed. The observed activity of 238 Np was used to calculate the amount of 238 Np produced during the irradiation via the 237 Np(n, γ) 238 Np reaction and an integral cross section of 342(11) mb was measured for the Godiva-IV neutron spectrum. Further, the 238 Np half-life has been measured with a result of 50.31(5) hours. The 237 Np(n, γ) 238 Np integral cross section measured in this work is in agreement with overlapping 1σ error bands to ENDF/B-VIII.0. However, the measured value is 3σ away from the calculated integral cross section using JENDL-5. This measurement offers a reliable benchmark for future 237 Np(n, γ) 238 Np cross section evaluations.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science [Slides]

Machine intelligence has the potential to revolutionize materials science, enabling autonomous synthesis, self-driving characterization, and accelerated modeling. However, despite the promise, successful implementation of these methods in day-to-day research remains a challenge. This talk will delve into the reasons behind this, exploring how truly intelligent experiments are hindered by opaque experiment control, a lack of domain-specific models, and human-centric design. Through a focus on the characterization of next-generation microelectronics and energy storage materials, I will share insights from both successful and failed attempts to implement machine intelligence. We will then explore the next steps necessary to unlock the full potential of machine intelligence in materials science, creating a future where intelligent systems work seamlessly alongside researchers to drive innovation and discovery.

36 MATERIALS SCIENCE↗

CHUWD-H v1.0: a comprehensive historical hourly weather database for U.S. urban energy system modeling

Reliable and continuous meteorological data are crucial for modeling the responses of energy systems and their components to weather and climate conditions, particularly in densely populated urban areas. However, existing long-term datasets often suffer from spatial and temporal gaps and inconsistencies, posing great challenges for detailed urban energy system modeling and cross-city comparison under realistic weather conditions. Here we introduce the Historical Comprehensive Hourly Urban Weather Database (CHUWD-H) v1.0, a 23-year (1998-2020) gap-free and quality-controlled hourly weather dataset covering 550 weather station locations across all urban areas in the contiguous United States. CHUWD-H v1.0 synthesizes hourly weather observations from stations with outputs from a physics-based solar radiation model and a reanalysis dataset through a multi-step gap filling approach. A 10-fold Monte Carlo cross-validation suggests that the accuracy of this gap filling approach surpasses that of conventional gap filling methods. Designed primarily for urban energy system modeling, CHUWD-H v1.0 should also support historical urban meteorological and climate studies, including the validation and evaluation of urban climate modeling.

54 ENVIRONMENTAL SCIENCES↗

Hacking Kilometer-Scale Models: A Participative Model for Climate Information

In May 2025, nearly 700 participants from all around the world coalesced at 10 regional nodes and a few satellite nodes to take part in a global hackathon of kilometer-scale (horizontal grid spacing < 10 km) regional and global Earth system models. Exciting science is emerging from these efforts, ranging across novel model analysis, new ways of integrating with satellite data, and emulation with machine learning. New technologies were trialed that enable the community to work in new and complementary ways to democratize access to global information at a local scale from a set of the world’s highest-resolution climate models. The hackathon demonstrated how exascale data can be organized to be accessible to anyone. Fundamentally, the community could apply these techniques and technologies to move toward more participative models for coproduction and delivery of diverse sources of climate information for climate scientists and citizens alike.

Climate models↗

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Dataset about Warming Effects on Carbon Cycling and Greenhouse Gas Fluxes in Permafrost Ecosystems

Field observations provide direct evidence of how does carbon cycling in permafrost ecosystems respond to climate change. This study provides a comprehensive dataset on the impact of warming on carbon cycling and greenhouse gas (GHG) fluxes in permafrost ecosystems. The dataset is extracted and integrated from 132 peer-reviewed studies with 1430 paired observations across eight major permafrost ecosystems, including Arctic and subarctic tundra and wetland, and alpine meadow, steppe, tundra and wetland. This dataset includes 17 variables from experiments conducted during the growing season, covering the plant and soil carbon pools, soil nitrogen pool, and GHG (i.e., CO 2 , CH 4 , and N 2 O) fluxes, among others. Background information on site climate conditions, vegetation and soil characteristics, and details of the warming experiments, including timing, methods, and warming magnitude, are also contained in the dataset. This dataset facilitates a comprehensive understanding of the impact of warming on carbon cycling and GHG fluxes in permafrost ecosystems, and provides supports for meta-analyses and literature reviews, remote sensing data validation, and land model development and parameterization.

Bao, Tao [Chinese Academy of Sciences (CAS), Beiji↗

Deep learning-assisted modeling for χ (2) nonlinear optics

Modeling second-order (χ(2)) nonlinear optical processes remains computationally expensive due to the need to resolve fast field oscillations and simulate wave propagation using methods such as the split-step Fourier method (SSFM). This can become a bottleneck in real-time applications, such as high-repetition-rate laser systems requiring rapid feedback and control. We present a long short-term memory-based surrogate model trained on SSFM simulations generated from a start-to-end model of the photocathode drive laser at SLAC National Accelerator Laboratory’s Linac Coherent Light Source II. The model achieves over 250× speedup while maintaining high fidelity, enabling future real-time optimization and laying the foundation for data-integrated modeling frameworks and digital twins of laser systems.

Accelerator Physics (physics.acc-ph)↗

MSD CoP Webinar: "Advances in MSD-LIVE to Support the MSD Community of Practice"

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Advances in MSD-LIVE to Support the MSD Community of Practice Presenters: Casey Burleyson and Zoe Guillen (Pacific Northwest National Laboratory) Abstract: The MultiSector Dynamics Living, Intuitive, Value-adding, Environment (MSD-LIVE; msdlive.org) is a cloud-based data management system and advanced computing platform that enables MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and workflows within the MSD Community of Practice. Recently, several high-profile datasets have attracted many new users to MSD-LIVE. This webinar has two goals: 1) To refamiliarize the MSD community and new users with the components of the platform (e.g., the data repository, model training notebooks, and data dashboards) and to highlight examples of how these components are advancing MSD science and 2) To demonstrate new features in v3 of the platform, released in late 2025. The main new feature in v3 is the ability to interactively explore data in MSD-LIVE without downloading it. MSD-LIVE users can now click a button in our data repository and launch a blank Jupyter notebook with access to the underlying data on AWS. Users can use the notebook to write analysis, visualization, or subsetting routines that process the data directly on the AWS cloud. We also added a GitHub integration feature that allows users to share analysis or visualization code they develop with the community of MSD-LIVE users. The webinar will wrap up with a look at what's coming next for MSD-LIVE in 2026. Moderator: Patrick M. Reed (MSD CoP Facilitation Team) This webinar was held on: May 12th, 2026 from 1-2 PM EST.

Open Science↗

Probabilistic Mixture Model-Based Spectral Unmixing

Spectral unmixing attempts to decompose a spectral ensemble into the constituent pure spectral signatures (called endmembers) along with the proportion of each endmember. This is essential for techniques like hyperspectral imaging (HSI) used in environment monitoring, geological exploration, etc. Several spectral unmixing approaches have been proposed, many of which are connected to hyperspectral imaging. However, most extant approaches assume highly diverse collections of mixtures and extremely low-loss spectroscopic measurements. Additionally, current non-Bayesian frameworks do not incorporate the uncertainty inherent in unmixing. We propose a probabilistic inference algorithm that explicitly incorporates noise and uncertainty, enabling us to unmix endmembers in collections of mixtures with limited diversity. We use a Bayesian mixture model to jointly extract endmember spectra and mixing parameters while explicitly modeling observation noise and the resulting inference uncertainties. We obtain approximate distributions over endmember coordinates for each set of observed spectra while remaining robust to inference biases from the lack of pure observations and the presence of non-isotropic Gaussian noise. As a direct impact of our methodology, access to reliable uncertainties on the unmixing solutions would enable robust solutions to noise, as well as informed decision-making for HSI applications and other unmixing problems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Discrete global grid system-based flow routing datasets in the Amazon and Yukon basins

Abstract. Discrete global grid systems (DGGS) are emerging spatial data structures widely used to organize geospatial datasets across scales. While DGGS have found applications in various scientific disciplines, including atmospheric science and ecology, their integration into physically based hydrological models and Earth system models (ESMs) has been hindered by the lack of flow routing datasets based on DGGS. In response to this gap, this study pioneers the development of new flow routing datasets using icosahedral Snyder equal-area (ISEA) DGGS and a novel mesh-independent flow direction model. We present flow routing datasets for two large basins, the tropical Amazon River basin and the Arctic Yukon River basin. These datasets (1) facilitate the adoption of DGGS for hydrological models and (2) provide flow routing inputs for evaluation of DGGS-based flow routing in the Amazon and Yukon river basins. The data are available at https://doi.org/10.5281/zenodo.8377765 (Liao, 2023).

54 ENVIRONMENTAL SCIENCES↗

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Phonon second harmonic generation in NaBr studied by inelastic neutron scattering and computer simulation

The phenomenon of second harmonic generation (SHG) was found for phonons in anharmonic NaBr by inelastic neutron scattering. The temperature dependence of this phonon SHG was measured from 300 K to 650 K. At 300 K the second harmonic (SH) is seen as a high-energy branch around 33 meV, nearly independent of $\overrightarrow{Q}$. The temperature effective potential (TDEP) method and classical molecular dynamics (MD) simulation with machine learning interatomic potential were able to reproduce the SH, and showed that SHG occurs with the flat transverse optical (TO) phonon branch. A classical model of a nonlinear medium explains the intensity and lifetime of the SH, compared to those of the TO modes. Also successful was a quantum model based on the Heisenberg-Langevin equation for interacting phonons coupled to a thermal bath, which also predicts a spectral distribution of the SH. In conclusion, the measured temperature dependence of the intensity of the second harmonic showed that it follows the Planck distribution of a one-phonon quasiparticle, and not two TO phonons.

36 MATERIALS SCIENCE↗

Efficient First-Order Algorithms for Large-Scale, Non-Smooth Maximum Entropy Models with Application to Wildfire Science

Maximum entropy (MaxEnt) models are a class of statistical models that use the maximum entropy principle to estimate probability distributions from data. Due to the size of modern data sets, MaxEnt models need efficient optimization algorithms to scale well for big data applications. State-of-the-art algorithms for MaxEnt models, however, were not originally designed to handle big data sets; these algorithms either rely on technical devices that may yield unreliable numerical results, scale poorly, or require smoothness assumptions that many practical MaxEnt models lack. In this paper, we present novel optimization algorithms that overcome the shortcomings of state-of-the-art algorithms for training large-scale, non-smooth MaxEnt models. Our proposed first-order algorithms leverage the Kullback–Leibler divergence to train large-scale and non-smooth MaxEnt models efficiently. For MaxEnt models with discrete probability distribution of n elements built from samples, each containing m features, the stepsize parameter estimation and iterations in our algorithms scale on the order of O(mn) operations and can be trivially parallelized. Moreover, the strong ℓ1 convexity of the Kullback–Leibler divergence allows for larger stepsize parameters, thereby speeding up the convergence rate of our algorithms. To illustrate the efficiency of our novel algorithms, we consider the problem of estimating probabilities of fire occurrences as a function of ecological features in the Western US MTBS-Interagency wildfire data set. Our numerical results show that our algorithms outperform the state of the art by one order of magnitude and yield results that agree with physical models of wildfire occurrence and previous statistical analyses of wildfire drivers.

Physics↗

Developing Partnership between San Jose State University and DOE Lawrence Livermore National Laboratory to Enhance Climate Research Equity and Inclusion

One of the key objectives of this project was to develop partnership between San Jose State University (SJSU) and the Lawrence Livermore National Laboratory (LLNL), a US Department of Energy (DOE) funded national laboratory. Both institutions are closely located within the San Francisco Bay Area in California and their researchers share overlapping research interests and expertise related to Earth system sciences. By leveraging their connections with LLNL, faculty and students from SJSU gained exposure to state-of-the-art observations and simulations related to Earth system sciences, including but not limited to the usage of the facility data provided by the DOE Atmospheric Radiation Measurement (ARM) program and data analysis techniques for interpreting and analyzing the DOE Energy Exascale Earth System Model (E3SM) simulations.

54 ENVIRONMENTAL SCIENCES↗

Model Data Archive Associated with Manuscript "Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon"

This data package supports the publication “Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon” by Li et al. (2026). The package contains processed model inputs, configuration files, restart files, simulation outputs, scripts, and visualization products used to evaluate post-fire dissolved organic carbon (DOC) dynamics in the Naches River Watershed, Washington, USA, following the 2021 Schneider Springs Fire. The modeling workflow couples ELM-BGC, the biogeochemistry-enabled Energy Exascale Earth System Model Land Model; ATS, the Advanced Terrestrial Simulator for integrated surface-subsurface hydrology; and PFLOTRAN, a reactive transport model for multicomponent aqueous geochemistry. Together, these models simulate how wildfire-induced changes in vegetation, litter, coarse woody debris, and soil organic matter influence DOC production, transport, and reaction from burned hillslopes to stream networks. The archive includes preprocessed meteorological, geospatial, hydrologic, and biogeochemical forcing data; ELM-BGC-derived DOC source terms; ATS mesh files; PFLOTRAN reactive-transport inputs; model configuration files; spin-up and transient restart files; watershed-scale diagnostic outputs; stream concentration time series; and figures or visualization files used to inspect and reproduce key results. File types include Hierarchical Data Format 5 (HDF5) files for gridded forcing and model-coupling data, model input and configuration files for ELM-BGC, ATS, and PFLOTRAN, restart and simulation-output files generated by the modeling workflow, tabular or time-series diagnostic outputs, scripts for post-processing and figure generation, and image or visualization products associated with the manuscript. Use of the package depends on the intended task. Re-running the simulations requires the relevant modeling software, including ELM-BGC, ATS, and PFLOTRAN as ATS's geochemical engine. Inspecting outputs and reproducing figures requires Python with scientific plotting libraries such as Matplotlib, and three-dimensional model outputs may be viewed with ParaView. Geographic information system files or maps may be inspected with ArcGIS Pro or comparable GIS software. The data package is intended to enable traceability, reuse, and partial reproduction of the coupled land-to-watershed hydro-biogeochemical modeling workflow used to test how wildfire disturbance affects terrestrial carbon pools and downstream DOC dynamics.

ATS↗