Search NASA⌕ Search

SEARCH · Search NASA

Results for “science data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Validation of a Global Geospace Model With a Systems Science Approach Based on Canonical Correlation Analysis

A systems science approach based on canonical correlation analysis (CCA) is applied as a new, behavioral way to validate global geospace models. The biggest novelty of the technique is that it validates models at a system level, whereby a side‐by‐side comparison is performed of CCA applied to a 30‐day observational and the corresponding simulation data sets comprising quiet, moderate and active times. The simulation used the Multiscale Atmosphere‐Geospace Environment (MAGE) model. It is shown that (a) CCA must be combined with sensitivity analysis to be effective, (b) the MAGE model generally reproduces the observed behavior (more so for quieter time intervals), quantified by the intercorrelations between different variables and (c) the technique identifies the SuperMAG SML index as a quantity for which refinements of the model are needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Opportunities for Earth Observation to Inform Risk Management for Ocean Tipping Points

Abstract As climate change continues, the likelihood of passing critical thresholds or tipping points increases. Hence, there is a need to advance the science for detecting such thresholds. In this paper, we assess the needs and opportunities for Earth Observation (EO, here understood to refer to satellite observations) to inform society in responding to the risks associated with ten potential large-scale ocean tipping elements: Atlantic Meridional Overturning Circulation; Atlantic Subpolar Gyre; Beaufort Gyre; Arctic halocline; Kuroshio Large Meander; deoxygenation; phytoplankton; zooplankton; higher level ecosystems (including fisheries); and marine biodiversity. We review current scientific understanding and identify specific EO and related modelling needs for each of these tipping elements. We draw out some generic points that apply across several of the elements. These common points include the importance of maintaining long-term, consistent time series; the need to combine EO data consistently with in situ data types (including subsurface), for example through data assimilation; and the need to reduce or work with current mismatches in resolution (in both directions) between climate models and EO datasets. Our analysis shows that developing EO, modelling and prediction systems together, with understanding of the strengths and limitations of each, provides many promising paths towards monitoring and early warning systems for tipping, and towards the development of the next generation of climate models.

Wood, Richard A. (ORCID:0000000239609513)↗

Analog Computing for Science

Conventional digital computing faces fundamental physical limits: large scale computing systems already con sume tens of Megawatts of power, Dennard scaling has ended, and data movement costs dominate application performance. Next generation experimental facilities generate data at rates that overwhelm conventional pro cessing and demand real-time analysis at the source. Analog computing, which exploits the continuous dynamics of physical systems to perform computation, promises a transformative path toward orders-of-magnitude gains in energy efficiency and time-to-solution for scientific workloads.

97 MATHEMATICS AND COMPUTING↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Galaxy-multiplet clustering from DESI DR2

We present an efficient estimator for higher-order galaxy clustering using small groups of nearby galaxies, or multiplets. Using the Luminous Red Galaxy (LRG) sample from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2, we identify galaxy multiplets as discrete objects and measure their cross-correlations with the general galaxy field. Our results show that the multiplets exhibit stronger clustering bias as they trace more massive dark matter halos than individual galaxies. When comparing the observed clustering statistics with the mock catalogs generated from the N-body simulation AbacusSummit, we find that the mocks underpredict multiplet clustering despite reproducing the galaxy two-point auto-correlation reasonably well. This discrepancy indicates that the standard Halo Occupation Distribution (HOD) model is insufficient to describe the properties of galaxy multiplets, revealing the greater constraining power of this higher-order statistic on galaxy-halo connection and the possibility that multiplets are specific to additional assembly bias. We demonstrate that incorporating secondary biases into the HOD model improves agreement with the observed multiplet statistics, specifically by allowing galaxies to preferentially occupy halos in denser environments. Our results highlight the potential of utilizing multiplet clustering, beyond traditional two-point correlation measurements, to break degeneracies in models describing the galaxy-dark matter connection.

cosmology↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

SpecFIDLER User Manual (Software V.2.6.0)

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Novel Design and Fabrication of a High Frequency Transient Heat Flux Sensor for Use in an RDE

Rotating detonation engine (RDE) combustion systems have been a topic of interest in the pressure gain combustion community for their benefits over traditional gas turbine engine combustors. However, cooling requirements for these engines are significantly higher and less predictable than non-detonating engines. To understand the high-speed heat transfer dynamics inside an RDE, a novel, high-frequency heat flux gage is presented. This study aims to design a robust, single-sided sensor that can withstand the high temperature and harsh environment of an RDE for extended durations. Sensor bench testing is performed using a hot plate as a heat source, and the sensor response is compared to a finite-element analysis (FEA) model. The sensor response is then tested inside a water-cooled RDE and the wall heat flux is compared to calorimetry data.

rotating detonation engines↗

Novel Design and Fabrication of a High Frequency Transient Heat Flux Sensor for Use in an RDE

Rotating detonation engine (RDE) combustion systems have been a topic of interest in the pressure gain combustion community for their benefits over traditional gas turbine engine combustors. However, cooling requirements for these engines are significantly higher and less predictable than non-detonating engines. To understand the high-speed heat transfer dynamics inside an RDE, a novel, high-frequency heat flux gage is presented. This study aims to design a robust, single-sided sensor that can withstand the high temperature and harsh environment of an RDE for extended durations. Sensor bench testing is performed using a hot plate as a heat source, and the sensor response is compared to a finite-element analysis (FEA) model. The sensor response is then tested inside a water-cooled RDE and the wall heat flux is compared to calorimetry data.

rotating detonation engines↗

Precision measurements of EFT parameters and BAO peak shifts for the Lyman- α forest

We present precision measurements of the bias parameters of the one-loop power spectrum model of the Lyman- α (Ly- α ) forest, derived within the effective field theory (EFT) of large-scale structure. We fit our model to the three-dimensional flux power spectrum measured from the ACCEL 2 hydrodynamic simulations. The EFT model fits the data with an accuracy of below 2% up to k = 2 h Mpc − 1 . Further, we analytically derive how nonlinearities in the three-dimensional clustering of the Ly- α forest introduce biases in measurements of the baryon acoustic oscillations (BAOs) scaling parameters in radial and transverse directions. From our EFT parameter measurements, we obtain a theoretical error budget of Δ α ∥ = − 0.2 % ( Δ α ⊥ = − 0.3 % ) for the radial (transverse) parameters at redshift z = 2.0 . This corresponds to a shift of − 0.3 % (0.1%) for the isotropic (anisotropic) distance measurements. We provide an estimate for the shift of the BAO peak for Ly- α -quasar cross-correlation measurements assuming analytical and simulation-based scaling relations for the nonlinear quasar bias parameters resulting in a shift of − 0.2 % ( − 0.1 % ) for the radial (transverse) dilation parameters, respectively. This analysis emphasizes the robustness of Ly- α forest BAO measurements to the theory modeling. We provide informative priors and an error budget for measuring the BAO feature—a key science driver of the currently observing Dark Energy Spectroscopic Instrument (DESI). Our work paves the way for full-shape cosmological analyses of Ly- α forest data from DESI and upcoming surveys such as the Prime Focus Spectrograph, WEAVE-QSO, and 4MOST. Published by the American Physical Society 2025

de Belsunce, Roger (ORCID:0000000336604028)↗

Opportunities in Hydropower and Pumped Storage Hydropower

Hydropower and pumped storage hydropower (PSH) are established technologies with a long-standing presence in the United States, but the industry remains dynamic with an active relicensing pipeline and growing interest in PSH to provide large-scale energy storage and grid services. NLR has developed a range of tools, data, and analysis that supports the evaluation of hydropower and PSH opportunities, complementing other DOE-supported efforts in this space. This presentation initiates a conversation around the future of hydropower and PSH and how it can inform the National Academy of Sciences investigation of possible DOE-supported regional energy-water technology pilots.

13 HYDRO ENERGY↗

Multi-omics data resource: Data package 25 (Pck025)

This data package comprises omics datasets from human pancreatic islets treated with IL-1β + IFNγ or with estrogen (E2) for 18 h. Two RNA-seq datasets are available: the first is a discovery dataset involving human islets treated with or without IL-1β + IFNγ for 18 hours; the second is a validation dataset, where human islets are treated with or without IL-1β + IFNγ or E2 for 18 hours. DIA proteomic analysis was performed on the same validation dataset samples. Data contributors: Kiersten L. Webster, Sarah Tersey & Raghavendra G. Mirmir: Kovler Diabetes Center and Department of Medicine, The University of Chicago, Chicago, IL, 60637, USA. Soumyadeep Sarkar, Raghavendra Mirmira, Ernesto S. Nakayasu: Biological Sciences Division, Pacific Northwest National Laboratory, Richland, WA, 99354, USA. Data repository: RNA-seq: GSE310965 Proteomics: MSV000101892 Publication: PMID 41279069

Sarkar, Soumyadeep [Pacific Northwest National Lab↗

Introducing Molecular Hypernetworks for Discovery in Multidimensional Metabolomics Data

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and address challenges of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used in the Global Natural Products Social Molecular Networking platform, are a prominent strategy for exploring and visualizing molecular relationships and improving annotation. MNs are mathematical graphs showing the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated with a single molecular identity. Here, this paper introduces “molecular hypernetworks” (MHNs) as more complex MN models able to natively represent multiway relationships among observations. Compared to MNs, MHNs can more parsimoniously represent the inherent complexity present among groups of observations, initially supporting improved exploratory data analysis and visualization. MHNs also promise to increase confidence in annotation propagation, for both human and analytical processing. We first illustrate MHNs with simple examples, and build them from liquid chromatography- and ion mobility spectrometry-separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Assessment of Extinction‐, Satellite‐, and Model‐Based Vertical Cloud Condensation Nuclei (CCN) Retrieval Methods Using Airborne CCN Measurements Over the Southern Great Plains

Abstract Accurate estimates of the vertical profile of cloud condensation nuclei (CCN) concentration are crucial to better quantify aerosol‐cloud interactions. We assessed the correlation between the vertical CCN concentrations obtained from extinction‐, satellite‐, and model‐based retrieval methods and airborne CCN concentrations collected at 0.24% supersaturation within the 3, 9, 27, and 81 km regions centered over the U.S. Department of Energy's Atmospheric Radiation Measurement User Facility Southern Great Plains (SGP) site during the spring and summer of 2016. The extinction profiles at a wavelength 355 nm were provided by the ground‐based Raman lidar. Our analysis showed moderate correlation between dry‐corrected extinction and airborne CCN data. We found the retrieved number concentration of CCN (RNCCN) method showed regression best‐fit slopes close to unity and consistent prediction errors for the majority of the data. The Lenhardt et al. (2023, https://doi.org/10.5194/amt‐16‐2037‐2023 ) method showed similar conclusions but only during spring, whereas the Mamouri and Ansmann (2016, https://doi.org/10.5194/acp‐16‐5905‐2016 ) method showed poor correlation. The Shinozuka et al. (2015, https://doi.org/10.5194/acp‐15‐7585‐2015 ) satellite‐based method exhibited reasonable agreement during summer but poor correlation during periods where both high (∼1,400 #/cm 3 ) and low (∼50 #/cm 3 ) airborne CCN concentrations were observed. The Copernicus Atmosphere Monitoring Service reanalysis modeled 3‐D CCN data set showed a moderate to weak positive correlation but performed poorly at high airborne CCN concentrations. Our analysis suggests the extinction‐based RNCCN method performed better than other methods across most observation periods under the diverse meteorological conditions observed at the SGP site.

54 ENVIRONMENTAL SCIENCES↗

Time series methods for the analysis of soundscapes and other cyclical ecological data

Biodiversity monitoring has entered an era of ‘big data’, exemplified by a near-continuous collection of sounds, images, chemical and other signals from organisms in diverse ecosystems. Such data streams have the potential to help identify new threats, assess the effectiveness of conservation interventions, as well as generate new ecological insights. However, appropriate analytical methods are often still missing, particularly with respect to characterizing cyclical temporal patterns. Here, we present a framework for characterizing and analysing ecological responses that represent nonstationary, complex temporal patterns and demonstrate the value of using Fourier transforms to decorrelate continuous data points. In our example, we use a framework based on three approaches (spectral analysis, magnitude squared coherence, and principal component analysis) to characterize differences in tropical forest soundscapes within and across sites and seasons in Gabon. By reconstructing the underlying, cyclic behaviour of the soundscape for each site, we show how one can identify circadian patterns in acoustic activity. Soundscapes in the dry season had a complex diel cycle, requiring multiple harmonics to represent daily variation, while in the wet season there was less variance attributable to the daily cyclic patterns. Our framework can be applied to most continuous, or near-continuous ecological data collected at a fine temporal resolution, allowing ecologists to explore patterns of temporal autocorrelation at multiple levels for biologically meaningful trends. Such methods will become indispensable as biological big data are used to understand the impact of anthropogenic pressures on biodiversity and to inform efforts to mitigate them.

54 ENVIRONMENTAL SCIENCES↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

2023 Project Peer Review Report

The Bioenergy Technologies Office (BETO) within the U.S. Department of Energy’s Office of Energy Efficiency and Renewable Energy supports the research, development, and demonstration (RD&D) of technologies aimed at mobilizing domestic renewable carbon resources for the reduction of greenhouse gas emissions across the U.S. economy. BETO systematically prioritizes RD&D into technology opportunities across a range of emerging scientific breakthroughs and technology readiness levels in the subprogram areas illustrated in Figure 1. This approach supports a diverse portfolio while developing the most promising and widely applicable technologies, testing technologies as integrated processes, and demonstrating integrated processes to support scale-up. These technologies will use a broad variety of renewable carbon resources to produce increasing volumes of biofuels and bioproducts. More information on BETO’s mission, goals, and strategic approaches can be found in the Bioenergy Technologies Office Multi-Year Program Plan. The biennial Peer Review process enables external stakeholders to provide feedback on the responsible use of taxpayer funding and develop recommendations for the most efficient and effective ways to accelerate the development of a bioenergy industry. This report includes the results of the Project Peer Review meeting held on April 3–7, 2023, in Denver, Colorado.

09 BIOMASS FUELS↗

The U.S. Agrivoltaic Shading Tool: A National-Scale Interface for Modeling Light and Shade Patterns in Ten Common Agrivoltaic Configurations

Agrivoltaic systems are dual-use configurations that co-locate agriculture and photovoltaic (PV) infrastructure and require careful design to balance crop performance and energy generation. A critical element of agrivoltaic design is the spatial and temporal distribution of irradiance and shade within and around PV arrays. To support research, planning, and stakeholder decision-making, we introduce the U.S. Agrivoltaic Shading Tool, a novel web-based application that delivers high-resolution irradiance and photosynthetically active radiation (PAR) modeling for ten standardized PV configurations across the conterminous United States. The tool leverages the National Laboratory of the Rockies (NLR) System Advisor Model (SAM) to perform detailed irradiance simulations, using meteorological data from the National Solar Radiation Database (NSRDB). Outputs include seasonal, monthly, weekly, and diurnal patterns of available sunlight, amount of shade, irradiance, and PAR at ground level within agrivoltaic system footprints. For a user's selected location, these results are visualized through interactive visualizations, heatmaps, and time-series plots, designed to be accessible to both technical and non-technical users. In addition to facilitating rapid spatial exploration of agrivoltaic light environments, the tool will offer seamless integration with the InSPIRE Agrivoltaics Design and Analysis Model (ADAM). This optional workflow will allow users to port selected site and configuration parameters into a more advanced modeling environment for further customization of structural layouts, crop-system compatibility, power generation, and technoeconomic performance. Finally, to promote open science, the entire dataset will be hosted and available for open access through the OpenEI platform. By standardizing and disseminating high-quality irradiance data and design tools, the U.S. Agrivoltaic Shading Tool supports a wide range of users, including researchers, landowners, energy developers, and policymakers, in evaluating the agronomic and energetic feasibility of agrivoltaic systems across the United States.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗