Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data structures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Staying Competitive in Clean Manufacturing: Insights on Barriers from Industry Interviews

While industrial emissions research has historically focused on energy-intensive sectors like steel, cement, and chemicals, this study addresses a critical gap by examining barriers across all the manufacturing industry in the U.S. Sectors like food processing, retail, plastics, and transportation face unique challenges distinct from heavy industry, operating on thin margins with limited bargaining power while experiencing heightened consumer and stakeholder pressure for improved environmental responsibility. Through structured interview data collection process and using quantitative ratings and qualitative analysis, this research identifies and categorizes emission reduction barriers across four key themes: financial, technical, organizational, and regulatory. Unlike energy-intensive industries that may pursue hydrogen or carbon capture technologies, discrete manufacturing industry like automotive, electrical and electronics, and machine manufacturers typically focus on energy efficiency, electrification of thermal processes, and alternate fuel switching, solutions better aligned with their lower-temperature processes and distributed facility profiles. The study’s primary contribution lies in documenting specific barrier manifestations within organizations and identifying proven mitigation strategies that companies have successfully implemented or observed among peers.

business competitiveness↗

Data-driven particle dynamics: Structure-preserving coarse-graining for emergent behavior in non-equilibrium systems

Multiscale systems are ubiquitous in science and technology, but are notoriously challenging to simulate as short spatiotemporal scales must be appropriately linked to emergent bulk physics. When expensive high-dimensional dynamical systems are coarse-grained into low-dimensional models, the entropic loss of information leads to emergent physics which are dissipative, history-dependent, and stochastic. To machine learn coarse-grained dynamics from time-series observations of particle trajectories, we propose a framework using the metriplectic bracket formalism that preserves these properties by construction; most notably, the framework guarantees discrete notions of the first and second laws of thermodynamics, conservation of momentum, and a discrete fluctuation-dissipation balance crucial for capturing non-equilibrium statistics. We introduce the mathematical framework abstractly before specializing to a particle discretization. As labels are generally unavailable for entropic state variables, we introduce a novel self-supervised learning strategy to identify emergent structural variables. We validate the method on benchmark systems and demonstrate its utility on two challenging examples: (1) coarse-graining star polymers at challenging levels of coarse-graining while preserving non-equilibrium statistics, and (2) learning models from high-speed video of colloidal suspensions that capture coupling between local rearrangement events and emergent stochastic dynamics. We provide open-source implementations in both PyTorch and LAMMPS, enabling large-scale inference and extensibility to diverse particle-based systems.

Computational Engineering, Finance, and Science (c↗

Replication Data for: Measurement of the mean number of muons with energies above 500 GeV in air showers detected with the IceCube Neutrino Observatory

<b>Measurement of the mean number of muons with energies above 500 GeV in air showers detected with the IceCube Neutrino Observatory</b> <br><br> This data release accompanies results submitted to Physical Review D describing the measurement of the average multiplicity of TeV muons with IceCube. It contains the data necessary to reproduce the main plots from the paper (Figs. 7 and 9), i.e. the numerical results for the average number of muons with energies above 500 GeV as a function of primary cosmic ray energy. <br><br> For any questions about this data release, please write to analysis@icecube.wisc.edu. <br><br> Files included in this release: <ul> <li>A README file <li>Files including data to reproduce the results plots from the paper (see below for details) <li>An example python script showing how to read and plot the data </ul> <br> <u>What is in the files icecube_Nmu500_X_Y.txt:</u> <br> Y indicates wether the file contains values obtained from experimental data (Y="data") or air-shower simulations (Y="MC"). <br> X indicates the hadronic interaction model for which the plot is made. If Y="data", this means that the experimental data was interpreted using this model. If Y="MC", it means that the simulations were performed with this model. The three models included are Sibyll 2.1, QGSJet-II.04, and EPOS-LHC (see paper for references). The file with X="modelaverage" gives the average over the three individual results with the deviations from the average included in the systematic uncertainties. <br><br> Please see the README file for details on how the data is structured in the files.

Astroparticle Physics↗

Structuring and storing signals with their metadata: practical considerations

This chapter aims to provide a comprehensive overview on structuring signal data and their metadata, highlighting key considerations for optimal management and storage. Specifically, it focuses on: (a) the relevance of data organization; (b) what data to store and what to keep; and (c) data management methods. Thus, this chapter answers where and how to store metadata efficiently. Chapter 3 explains what is considered metadata. Chapters 5 and 6 present how to collect certain metadata through dedicated sensor validation tests (Chapter 5) or algorithmic analysis (Chapter 6).

Nicolaï, Niels↗

xCDAT: A Python Package for Simple and Robust Analysis of Climate Data

xCDAT (Xarray Climate Data Analysis Tools) is an open-source Python package that extends Xarray (Hoyer & Hamman, 2017) for climate data analysis on structured grids. xCDAT streamlines analysis of climate data by exposing common climate analysis operations through a set of straightforward APIs. Some of xCDAT’s key features include spatial averaging, temporal averaging, and regridding. These features are inspired by the Community Data Analysis Tools (CDAT) library (Dean N. Williams et al., 2009) (D. N. Williams, 2014) (Doutriaux et al., 2019) and leverage powerful packages in the Xarray ecosystem including xESMF (Zhuang et al., 2023), xgcm (Abernathey et al., 2022), and CF xarray (Cherian et al., 2023). To ensure general compatibility across various climate models, xCDAT operates on datasets that are compliant with the Climate and Forecast (CF) metadata conventions (Hassell et al., 2017).

54 ENVIRONMENTAL SCIENCES↗

A new chapter for RCSB Protein Data Bank Molecule of the Month in 2025

The online Molecule of the Month series authored by David S. Goodsell and published by the Research Collaboratory for Structural Biology Protein Data Bank at PDB101.RCSB.org has highlighted stories about the biomolecular structures driving fundamental biology, biomedicine, bioenergy, and biotechnology since January 2000. A new chapter begins in 2025: Janet Iwasa has taken over as the series creator of stories about critically important biological macromolecules in a rapidly changing world.

Bioenergy↗

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (↗

Data about data – when, why and how metadata can support the digital plant

A structured approach for recording data quality and contextual information about how and why a signal exists – i.e. metadata – is central to interpret and use sensor data correctly. This is becoming increasingly important with the global trend with data-driven applications such as digital twins and AI-models. But a structured metadata collection and organization of sensor data is not routine in most plants, which can result in lost information and missed opportunities to make use of the investments made in the data collection. Therefore, the IWA task group on Metadata Collection and Organization in wastewater resource recovery systems (MetaCO) was initiated in 2020 and recently delivered the IWA scientific and technical report number 31. The report gives and in-depth description about metadata in water resources recovery facilities (WRRFs) and is available as open access at IWA publishing. The report is the outcome of the collaboration between more than 80 water professionals with the intention to serve WRRF data users with a guide on how to structure and make use of metadata throughout the data pipeline in order to maximize the value of sensor data.

Alferes, Janelcy [VITO, Belgium]↗

Emergent Nanostructure and Ion Transport in Polyzwitterion/Polyanion Blends

We investigated blends of poly(1-(3-sulfonatopropyl)-2-vinylpyridinium) (P2VPPS) and poly(lithium (trifluoromethane)sulfonimide methacrylate) (poly(MTFSI)Li) at varying molar ratios to gain a mechanistic understanding of ionic conductivity in a miscible polyzwitterion/polyanion system. This dataset contains the raw numerical data corresponding to the figures in the manuscript. The data files include the following information: (1) Experimental Data – includes X-ray and neutron scattering measurements, broadband dielectric spectroscopy (BDS) data, extracted DC conductivity values, differential scanning calorimetry (DSC) and thermogravimetric analysis (TGA) results, extracted glass transition temperatures, etc. (2) CGMD Data – includes molecular dynamics (MD) trajectory files and computed structural correlations. All data files are organized/named according to the figure numbers in the manuscript.

36 MATERIALS SCIENCE↗

Data for Magnetic Anisotropy due to Localized Structural Defects in Strained Yttrium Iron Garnet Thin Films

This dataset contains electron microscopy, magnetic, and X-ray diffraction characterization data for YIG/YSGG thin films associated with the manuscript "Magnetic anisotropy due to localized structural defects in strained yttrium iron garnet thin films". Characterizations include: scanning electron nano diffraction (SEND, or 4D-STEM), dark field transmission electron microscopy (DF-TEM), energy dispersive X-ray spectroscopy (EDS), ferrimagnetic resonance (FMR), superconducting quantum interference device (SQUID), scanning transmission electron microscopy (STEM), X-ray diffraction (XRD).

Electron microscopy↗

Population structure limits the use of genomic data for predicting phenotypes and managing genetic resources in forest trees

There is overwhelming evidence that forest trees are locally adapted to climate. Thus, genecological models based on population phenotypes have been used to measure local adaptation, infer genetic maladaptation to climate, and guide assisted migration. However, instead of phenotypes, there is increasing interest in using genomic data for gene resource management. We used whole-genome resequencing and common-garden experiments to understand the genetic architecture of adaptive traits in black cottonwood. We studied the potential of using genome-wide association studies (GWAS) and genomic prediction to detect causal loci, identify climate-adapted phenotypes, and inform gene resource management. We analyzed population structure by partitioning phenotypic and genomic (single-nucleotide polymorphism) variation among 840 genotypes collected from 91 stands along 16 rivers. Most phenotypic variation (60 to 81%) occurred among populations and was strongly associated with climate. Population phenotypes were predicted well using genomic data (e.g., predictive abilityr> 0.9) but almost as well using climate or geography (r> 0.8). In contrast, genomic prediction within populations was poor (r< 0.2). We identified many GWAS associations among populations, but most appeared to be spurious based on pooled within-population analyses. Hierarchical partitioning of linkage disequilibrium and haplotype sharing suggested that within-population genomic prediction and GWAS were poor because allele frequencies of causal loci and linked markers differed among populations. Given the urgent need to conserve natural populations and ecosystems, our results suggest that climate variables alone can be used to predict population phenotypes, delineate seed zones and deployment zones, and guide assisted migration.

Science & Technology - Other Topics↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]↗

Degradation and deforestation increase the sensitivity of the Amazon Forest to climate extremes

About 40% of the Brazilian Amazon has been deforested or suffered changes in forest structure through degradation (selective logging, fires, and fragmentation). The impact of forest degradation on the forest’s sensitivity to climate extremes has not been fully explored because of a lack of data and the complex interplay of forest structure and climate drivers. Here, we combined forest structure data from 545 airborne lidar transects (375 ha each) across the Brazilian Amazon with the Ecosystem Demography Model (ED2). We explore the forest’s functional response to near-present (1981–2019) climate extremes under observed forest structure from lidar ( Control ) and two forest structure change scenarios: (1) forest recovery by excluding all future deforestation and degradation ( Recovery ) and (2) expansion of selective logging and deforestation ( Degradation ). Using the Control simulation, we found a close and positive association between local forest aboveground biomass and the predicted gross primary productivity (GPP) and evapotranspiration (ET). Moreover, both GPP and ET respond negatively to extremes in vapor pressure deficit and downwelling shortwave irradiance in degraded forests in Eastern and Southern Amazon, indicating high sensitivity to droughts. Locally high-biomass forest patches showed little or no negative response of GPP and ET to extreme drought conditions whereas low-biomass forest patches in the same locations—typically degraded forest canopies—responded negatively to higher moisture stress. The results from the Recovery scenario showed similar results to simulations with observed structure; however, under the Degradation scenario, low-biomass forest patches became more abundant, resulting in more regions where GPP and ET are negatively impacted by hot drought conditions according to the ED2 model. Our results suggest that local forest structure is a critical determinant of an ecosystem’s response to climate variability, and that the loss of canopy trees in the Amazon through forest degradation could increase and expand forest vulnerability to droughts.

54 ENVIRONMENTAL SCIENCES↗

Extraction of the neutron F 2 structure function from inclusive proton and deuteron deep-inelastic scattering data

The available world deep-inelastic scattering (DIS) data on proton and deuteron structure functions F 2 p , F 2 d , and their ratios are leveraged to extract the free neutron F 2 n structure function, the F 2 n / F 2 p ratio, and associated uncertainties using the latest nuclear effect calculations in the deuteron. Special attention is devoted to the normalization of the proton and deuteron experimental datasets and to the treatment of correlated systematic errors, as well as the quantification of procedural and theoretical uncertainties. The extracted F 2 n dataset is utilized to evaluate the Q 2 dependence of the Gottfried sum rule and the nonsinglet F 2 p − F 2 n moments. To facilitate replication of our study, as well as for general applications, we provide a comprehensive DIS database including all recent Jefferson Lab 6 GeV measurements, the extracted F n 2 , a modified CTEQ-JLab global parton distribution function fit named CJ15nlo_mod, and grids with calculated proton, neutron, and deuteron DIS structure functions. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)↗

VA EDH Data Curation Documentation FY25-Q1

This data source documentation report provides researchers with valuable insights into the structure, contents, and data sources used to compile the datasets. It specifically covers the Fiscal Year 2024, Fourth Quarter (FY25-Q1) dataset curation documentation for the Environmental Determinants of Health (EDH) project.

97 MATHEMATICS AND COMPUTING↗