Search NASA⌕ Search

SEARCH · Search NASA

Results for “data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Event generators for high-energy physics experiments

We provide an overview of the status of Monte-Carlo event generators for high-energy particle physics. Guided by the experimental needs and requirements, we highlight areas of active development, and opportunities for future improvements. Particular emphasis is given to physics models and algorithms that are employed across a variety of experiments. These common themes in event generator development lead to a more comprehensive understanding of physics at the highest energies and intensities, and allow models to be tested against a wealth of data that have been accumulated over the past decades. A cohesive approach to event generator development will allow these models to be further improved and systematic uncertainties to be reduced, directly contributing to future experimental success. Event generators are part of a much larger ecosystem of computational tools. They typically involve a number of unknown model parameters that must be tuned to experimental data, while maintaining the integrity of the underlying physics models. Making both these data, and the analyses with which they have been obtained accessible to future users is an essential aspect of open science and data preservation. It ensures the consistency of physics models across a variety of experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING↗

Wind Turbine Sound Setbacks and Supply Curves: Ordinances and Extrapolated Trends, 110 Hub Height, 130 Rotor Diameter

This dataset provides a comprehensive set of wind turbine sound setbacks from every residential structure in the contiguous United States (CONUS). A sound setback is defined as the minimum required distance between a residential structure and a hypothetical turbine installation site to ensure that modeled sound levels received at the residence do not exceed local sound ordinances, which are commonly expressed in A-weighted decibels (dBA). Therefore, sound setbacks are a local spatial assessment combining multiple factors, including the sound pressure curve as a function of the observer location (distance and direction) relative to the turbine, local sound regulations, and the geographical distribution of residential structures. The dataset is organized into multiple scenario-based products, detailed as follows: 1. Existing and extrapolated sound setbacks. An existing scenario characterizes sound setbacks only in states or counties that have implemented sound regulations as of 2022. The extrapolated scenarios extend a constant sound threshold to counties that lack explicit sound regulations, with thresholds ranging from 35 to 60 dBA, in 5-dBA increments reflecting the variation observed in current sound ordinances. 2. Sound setbacks in directional and worst scenarios. The directional scenario accounts for the distance and orientation of residential structures relative to a hypothetical turbine location, utilizing the turbine's sound emissions in that specific direction. In contrast, the worst scenario takes loudest sound level at each distance step from the turbine, irrespective of directional considerations, which aligns with current industry practice. 3. Supply curves for Open and Reference Access scenarios. This dataset includes supply curves generated by the reV model, which integrates each of the above sound setbacks into both Open and Reference siting scenarios. In addition, two Open and Reference baselines scenarios were included which do not consider sound setbacks for comparative analysis. All sound setback data are stored in TIF files, with partial maps of the data provided in PNG format. The values in the sound setback raster range from 0 to 1, representing the fraction of developable land within a 90 meter by 90 meter pixel due to sound ordinances. A value of 0 indicates areas where wind energy development is prohibited, while a value of 1 signifies areas fully permissible. The wind turbine parameters used in the sound modeling are based on the land-based turbine from International Energy Agency (IEA), featuring a rated electrical power of 3.4 MW, a rotor diameter of 130 meters, and a hub height of 110 meters. The atmospheric conditions, including wind speed/direction, turbulence, air temperature, relative humidity, and air pressure, that drive the sound generation are obtained from the WIND Toolkit dataset.

Array↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Retrospective on decadal progress of the NOAA/NPS ocean noise reference station network

The National Oceanic and Atmospheric Administration (NOAA), in partnership with the U.S. National Park Service (NPS), established the Ocean Noise Reference Station Network (NRS) in 2014 as a foundational component of NOAA’s Ocean Noise Strategy. This long-term effort aims to characterize baseline ocean ambient sound conditions across diverse marine environments and to inform management of noise impacts on protected species and habitats within U.S. waters. The NRS is now composed of 13 autonomous passive acoustic monitoring stations strategically positioned across the U.S. Exclusive Economic Zone (EEZ), extending from Arctic regions to tropical waters in depths ranging from 33 to 4,790 m. These locations include several National Marine Sanctuaries and National Parks, such as the recently designated Chumash Heritage National Marine Sanctuary off the coast of California. Each station is equipped to continuously sample low-frequency underwater sound at five kHz, enabling the detection of anthropogenic, geophysical, and biological acoustic signals. To date the network has sampled over 72 years of calibrated acoustic data. The spatial breadth and consistent methodology of the NRS allow for comparative acoustic assessments across diverse marine ecosystems. In addition to applied research functions, the NRS has served as a platform for education and training, offering opportunities for students to develop skills for marine science and data analysis. Looking forward, the NRS project team is focused on network expansion, improved data delivery, and broader integration with collaborative scientific initiatives. NRS recordings are being archived in partnership with NOAA’s National Centers for Environmental Information to enhance accessibility and long-term utility. Efforts are underway to develop standardized metadata and summary products to accompany raw audio files, making the data more usable for a wide range of stakeholders in the ocean science community. The NRS is evolving into a fully integrated national framework for ocean sound monitoring that supports scientific inquiry, management decision-making, national security interests, and public engagement with ocean acoustic environments.

Long-term monitoring↗

Mind the gap: Bridging the divide between AI aspirations and the reality of autonomous microscopy

What does materials science look like in the “Age of Artificial Intelligence?” Each material’s domain—synthesis, characterization, and modeling—has a different answer to this question, motivated by unique challenges and constraints. This work focuses on the tremendous potential of autonomous characterization within electron microscopy. We present our recent advancements in developing domain-aware, multimodal models for microscopy analysis capable of describing complex atomic systems. We then address the critical gap between the theoretical promise of autonomous microscopy and its current practical limitations, showcasing recent successes while highlighting the necessary developments to achieve robust, real-world autonomy.

2D materials↗

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Recovered supernova Ia rate from simulated LSST images

Aims.TheVera C. RubinObservatory’s Legacy Survey of Space and Time (LSST) will revolutionize time-domain astronomy by detecting millions of different transients. In particular, it is expected to increase the number of known type Ia supernovae (SN Ia) by a factor of 100 compared to existing samples up to redshift ∼1.2. Such a high number of events will dramatically reduce statistical uncertainties in the analysis of the properties and rates of these objects. However, the impact of all other sources of uncertainty on the measurement of the SN Ia rate must still be evaluated. The comprehension and reduction of such uncertainties will be fundamental both for cosmology and stellar evolution studies, as measuring the SN Ia rate can put constraints on the evolutionary scenarios of different SN Ia progenitors. Methods.We used simulated data from the Dark Energy Science Collaboration (DESC) Data Challenge 2 (DC2) and LSST Data Preview 0 to measure the SN Ia rate on a 15 deg 2 region of the “wide-fast-deep” area. We selected a sample of SN candidates detected in difference images, associated them to the host galaxy with a specially developed algorithm, and retrieved their photometric redshifts. We then tested different light-curve classification methods, with and without redshift priors (albeit ignoring contamination from other transients, as DC2 contains only SN Ia). We discuss how the distribution in redshift measured for the SN candidates changes according to the selected host galaxy and redshift estimate. Results.We measured the SN Ia rate, analyzing the impact of uncertainties due to photometric redshift, host-galaxy association and classification on the distribution in redshift of the starting sample. We find that we are missing 17% of the SN Ia, on average, with respect to the simulated sample. As 10% of the mismatch is due to the uncertainty on the photometric redshift alone (which also affects classification when used as a prior), we conclude that this parameter is the major source of uncertainty. We discuss possible reduction of the errors in the measurement of the SN Ia rate, including synergies with other surveys, which may help us to use the rate to discriminate different progenitor models.

Astronomy & Astrophysics↗

Constraining gas motion and non-thermal pressure beyond the core of the Abell 2029 galaxy cluster with XRISM

We report on a detailed spectroscopic study of the gas dynamics and hydrostatic mass bias of the galaxy cluster Abell 2029, utilizing high-resolution observations from XRISM Resolve. Abell 2029, known for its cool core and relaxed X-ray morphology, provides an excellent opportunity to investigate the influence of gas motions beyond the central region. Expanding upon prior studies that revealed low turbulence and bulk motions within the core, our analysis covers regions out to the scale radius $R_{2500}$ (670 kpc) based on three radial pointings extending from the cluster center toward the northern side. We obtain accurate measurements of bulk and turbulent velocities along the line of sight. The results indicate that non-thermal pressure accounts for no more than 2% of the total pressure at all radii, with a gradual decrease outward. The observed radial trend differs from many numerical simulations, which often predict an increase in non-thermal pressure fraction at larger radii. These findings suggest that deviations from hydrostatic equilibrium are small, leading to a hydrostatic mass bias of around 2% across the observed area.

X-rays: galaxies: clusters↗

Navigating the Noise: Bringing Clarity to ML Parameterization Design With O $\boldsymbol{\mathcal{O}}$(100) Ensembles

Abstract Machine‐learning (ML) parameterizations of subgrid processes (here of turbulence, convection, and radiation) may one day replace conventional parameterizations by emulating high‐resolution physics without the cost of explicit simulation. However, uncertainty about the relationship between offline and online performance (i.e., when integrated with a large‐scale general circulation model) hinders their development. Much of this uncertainty stems from limited sampling of the noisy, emergent effects of upstream ML design decisions on downstream online hybrid simulation. Our work rectifies the sampling issue via the construction of a semi‐automated, end‐to‐end pipeline for size ensembles of hybrid simulations, revealing important nuances in how systematic reductions in offline error manifest in changes to online error and online stability. For example, removing dropout and switching from a Mean Squared Error to a Mean Absolute Error loss both reduce offline error, but they have opposite effects on online error and online stability. Other design decisions, like incorporating memory, converting moisture input from specific humidity to relative humidity, using batch normalization, and training on multiple climates do not come with any such compromises. Finally, we show that ensemble sizes of may be necessary to reliably detect causally relevant differences online. By enabling rapid online experimentation at scale, we can empirically settle debates regarding subgrid ML parameterization design that would have otherwise remained unresolved in the noise.

Lin, Jerry [Department of Earth System Sciences Un↗

Predictions for the Detectability of Milky Way Satellite Galaxies and Outer-Halo Star Clusters with the Vera C. Rubin Observatory

We predict the sensitivity of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) to faint, resolved Milky Way satellite galaxies and outer-halo star clusters. We characterize the expected sensitivity using simulated LSST data from the LSST Dark Energy Science Collaboration (DESC) Data Challenge 2 (DC2) accessed and analyzed with the Rubin Science Platform as part of the Rubin Early Science Program. We simulate resolved stellar populations of Milky Way satellite galaxies and outer-halo star clusters over a wide range of sizes, luminosities, and heliocentric distances, which are broadly consistent with expectations for the Milky Way satellite system. We inject simulated stars into the DC2 catalog with realistic photometric uncertainties and star/galaxy separation derived from the DC2 data itself. We assess the probability that each simulated system would be detected by LSST using a conventional isochrone matched-filter technique. We find that assuming perfect star/galaxy separation enables the detection of resolved stellar systems with $M_V$ = 0 mag and $r_{1/2}$ = 10 pc with >50% efficiency out to a heliocentric distance of ~250 kpc. Similar detection efficiency is possible with a simple star/galaxy separation criterion based on measured quantities, although the false positive rate is higher due to leakage of background galaxies into the stellar sample. When assuming perfect star/galaxy classification and a model for the galaxy-halo connection fit to current data, we predict that 89 +/- 20 Milky Way satellite galaxies will be detectable with a simple matched-filter algorithm applied to the LSST wide-fast-deep data set. Different assumptions about the performance of star/galaxy classification efficiency can decrease this estimate by ~7%-25%, which emphasizes the importance of high-quality star/galaxy separation for studies of the Milky Way satellite population with LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

Overionized plasma in the supernova remnant Sagittarius A East anchored by XRISM observations

Sagittarius A East is a supernova remnant with a unique surrounding environment, as it is located in the immediate vicinity of the supermassive black hole at the Galactic center, Sagittarius A$^{*}$. The X-ray emission of the remnant is suspected to show features of overionized plasma, which would require peculiar evolutionary paths. We report on the first observation of Sagittarius A East with the X-Ray Imaging and Spectroscopy Mission (XRISM). Equipped with a combination of a high-resolution microcalorimeter spectrometer and a large field-of-view CCD imager, we for the first time resolved the Fe xxv K-shell lines into fine structure lines and measured the forbidden-to-resonance intensity ratio to be $1.39 \pm 0.12$, which strongly suggests the presence of overionized plasma. We obtained a reliable constraint on the ionization temperature just before the transition into the overionization state, of $\gt\! 4\:$keV. The recombination timescale was constrained to be $\lt\! 8 \times 10^{11} \:$cm$^{-3}\:$s. The small velocity dispersion of $109 \pm 6\:$km$\:$s$^{-1}$ indicates a low Fe ion temperature $\lt\! 8\:$keV and a small expansion velocity $\lt\! 200\:$km$\:$s$^{-1}$. The high initial ionization temperature and small recombination timescale suggest that either rapid cooling of the plasma via adiabatic expansion from dense circumstellar material or intense photoionization by Sagittarius A$^{*}$ in the past may have triggered the overionization.

galaxy center↗

Thermal and kinematic properties of ejecta in SN1987A revealed by XRISM

We present an analysis of high-resolution spectra from the shock-heated plasmas in SN 1987A, based on an observation using the Resolve instrument onboard the X-Ray Imaging and Spectroscopy Mission (XRISM). The 1.7–10 keV Resolve spectra are accurately represented by a single-component, plane-parallel shock plasma model, with a temperature of $2.84_{-0.08}^{+0.09}$ keV and an ionization parameter of $2.64_{-0.45}^{+0.58}$ × $10^{11}\,\,{\rm s\,\, cm}^{-3}$. The Resolve spectra are also well reproduced by the 3D magneto-hydrodynamic simulation presented by Orlando et al. (2020, A&A, 636, A22) suggesting substantial contribution from the ejecta. The metal abundances obtained with Resolve align with the Large Magellanic Cloud value, indicating that the X-rays in 2024 originate from “non-metal-rich” shock-heated ejecta and the reverse shock has not reached the inner metal-rich region of ejecta. Doppler widths of the atomic lines from Si, S, and Fe correspond to velocities of 1500–1700 km s$^{-1}$, where the thermal broadening effects in this non-metal-rich plasma are negligible. Therefore, the line broadening seen in Resolve spectra is determined by the large bulk motion of ejecta. For reference, we determined a $90\%$ upper limit on non-thermal emission from a pulsar wind nebula at $4.3 \times 10^{-13}$ erg cm$^{-2}$ s$^{-1}$ in the 2–10 keV range, aligning with NuSTAR findings by Greco et al. (2022, ApJ, 931, 132). Additionally, we searched for the $^{44}$Sc K line feature and found a $1\sigma$ upper limit of $1.0 \times 10^{-6}$ photons cm$^{-2}$ s$^{-1}$, which translates to an initial $^{44}$Ti mass of approximately $2 \times 10^{-4}\, M_{\odot }$, consistent with previous X-ray to soft gamma-ray observations (Boggs et al. 2015, Science, 348, 670; Grebenev et al. 2012, Nature, 490, 373; Leising 2006, ApJ, 651, 1019).

ISM: supernova remnants↗