Search NASASearch

SEARCH · Search NASA

Results for “Data structures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles

buhito

buhito is a Python library for graph analysis and machine learning. Graphs can represent networks with objects as nodes and their relationships as edges. buhito focuses on graphlet methods that study graphs through enumerating their component subgraphs to enable interpretable and fast models of complex systems. The package provides tools for different algorithmic designs for computing, analyzing, and applying graphlets to research problems such as machine learning, data compression, and anomaly detection in graph-structured data. A central feature is performing decomposition data analysis on graphs for machine learning models. Implemented in Python and built upon open-source scientific libraries such as NetworkX, NumPy, and SciPy, buhito provides high-performance methods for researchers exploring the mathematical and computational foundations of graphlet analysis applicable to systems of different sizes.

Pimonova, Yulia

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence

plexosdb: A Modular Library for Programmatic PLEXOS Model Construction

plexosdb is a lightweight Python library for constructing PLEXOS models using a SQLite-backed data structure. It provides a clear, modular interface that maps relational data directly to model components. By leveraging SQLite and idiomatic Python, it enables fast iteration and reproducible workflows. The result is a performant, composable foundation for scalable PLEXOS model development.

24 POWER TRANSMISSION AND DISTRIBUTION

Recommended Nuclear Structure and Decay Data for A=206 Isobars

Here, evaluated nuclear structure and decay data for all nuclei with mass number A=206 ( 206 Pt, 206 Au, 206 Hg, 206 Tl, 206 Pb, 206 Bi, 206 Po, 206 At, 206 Rn, 206 Fr, 206 Ra and 206 Ac), are presented. All available experimental data are compiled and evaluated, and best values for level and γ-ray energies, quantum numbers, lifetimes, γ-ray intensities and transition probabilities, as well as other nuclear properties, are recommended. Inconsistencies and discrepancies that exist in the literature are discussed. A number of computer codes (https://www-nds.iaea.org/public/ensdf_pgm/) developed by members of the NSDD network were used during the evaluation process. This work supersedes the earlier evaluation by F.G. Kondev (2008Ko21), published in Nuclear Data Sheets 109, 1527 (2008).

Kondev, F. G. [Argonne National Laboratory (ANL),

Enabling event-by-event precision in γ-ray cascades for neutron-induced reactions

Neutron-induced γ-ray spectra provide key inputs for modern active interrogation applications. A precise modeling of the nuclear reaction and subsequent emission of γ rays is challenging and often impossible due to limitations on evaluated data file formats and nuclear transport simulation codes. We present a framework that addresses these challenges by combining experimental data and reaction-model calculation outputs into an extended candidate version of the Generalized Nuclear Data Structure (GNDS) file, the successor format for the legacy Evaluated Nuclear Data File (ENDF-6). This proposed GNDS hierarchical format contains all the necessary ingredients for inline γ-ray cascade reproduction with event-by-event precision, including continuum–continuum and continuum–discrete transitions following neutron-capture and inelastic neutron scattering reactions. Cascade-event generation based on our approach demonstrates improved energy conservation on an event-by-event basis and permits the use of γ-γ coincidences in applications. This work offers, for the first time, a method to generate neutron-capture and inelastic neutron-scattering γ-ray cascades where energy conservation, correlations, and experimental primaries are fully accounted for.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Getting Started with Evaluations with Means and Uncertainties (EMU 3.0)

One of the most fundamental quantities in nuclear physics is the reaction cross section. A cross section represents the probability that a nuclear reaction will resolve through a given channel given a target nucleus and a projectile with a certain energy. A nuclear evaluation is a set of discrete data and interpolation rules to convert those discrete nuclear reaction data—such as the cross section—into a continuous function at arbitrary energies. Evaluated nuclear data files that can appear in Evaluated Nuclear Data File (ENDF) and Generalized Nuclear Data Structure (GNDS) formats, storing a “most-complete” discretized representation of nuclear data, based on both experimental measurements and theory models.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure

First-principles fusion plasma simulations are both compute and memory intensive, and CGYRO is no exception. The use of many HPC nodes to fit the problem in the available memory thus results in significant communication overhead, which is hard to avoid for any single simulation. That said, most fusion studies are composed of ensembles of simulations, so we developed a new tool, named XGYRO, that executes a whole ensemble of CGYRO simulations as a single HPC job. By treating the ensemble as a unit, XGYRO can alter the global buffer distribution logic and apply optimizations that are not feasible on any single simulation, but only on the ensemble as a whole. The main saving comes from the sharing of the collisional constant tensor structure, since its values are typically identical between parameter-sweep simulations. This data structure dominates the memory consumption of CGYRO simulations, so distributing it among the whole ensemble results in drastic memory savings for each simulation, which in turn results in overall lower communication overhead.

CGYRO

Fox Trails

1. This software utilizes python pandas to pull data from P6 databases or XER files. The software transforms the datasets into multiple main tables by joining, filtering, iteratively flattening hierarchical structured data, and pivoting datasets to give simple flat output tables. The activity table includes all of the information related to an activity including activity codes, global, EPS, and project codes, UDFs, and WBS information as separate columns. This includes the code id, code value and sequence number for all levels in hierarchical codes. The resource table is similar to the activity table and includes all of the information related to resources on activities including UPFs and resource codes. The resource time phased table takes the resource information and time phases it for the budget, forecast, late, and actual dates/units/costs that closely matches P6's user interface's values as it implements the resource curve and calendars. The wbs table contains the WBS structure broken out by levels and includes UDFs, codes, and notebook topics. The final P6 data table is the relationships table which simply contains the relationships. 2. When a user updates the tool with data (via giving it P6 project names with database username/password information or XER files) the system creates the data in #1, then creates a networkx graph with the activity data imbedded in the node data and the relationships added as edges. Each edge also has it's float calculated (working time distance between the predecessor and successor) and attached to the edge. Activities are also tagged as a potential start of a path based on their constraints, constraint dates, remaining start date, and activity status. When a user enters an activity ID into the UI, it runs a shortest path calculation on the network graph between each node tagged as potential start to the entered activity id based on the float tagged on the edge. Each path returned by the algorithm contains all of the nodes on the path in order, as well as the total float of the edges that make the path. This data is then collected and returned to the user in the form of a gantt chart with groupings for each path that includes the total float for each group. 3. Similar to 2, if the user passes through a reference dataset each activity set in the path is checked to see if it had a path in the reference dataset, if that path was the primary path between the start and end activities, and what has changed regarding logic and durations. These changes are color coded and summarized before sent to the user to be displayed by the UI for simple discovery. 4. Utilizing the data from #1, the user can submit desired grouping code(s) and filters to the system. The system will then pull the activities, resources, and relationships and create a gantt chart based on the groupings sent and filtered based on the filters sent. 5. The system will produce a gantt chart in a similar method to #4, but allows interactivity with the data. As the user interacts with the gantt chart, the software captures the changes and stores it with the user making the change so that project controls and implement those changes in P6.

Fox, Ben

Insights into the Complexation of Actinides by Diethylenetriaminepentaacetic Acid from Characterization of the Americium(III) Complex

Diethylenetriaminepentaacetic acid (DTPA) is a frequently used chelator in the nuclear and medical industries, especially for the complexation of trivalent actinides. However, structural data on these complexes in the solid-state have long remained elusive. Herein, a detailed structural analysis of the presented crystal structures of [C(NH 2 ) 3 ] 4 [Nd(DTPA)] 2 · n H 2 O and [C(NH 2 ) 3 ] 4 [Am(DTPA)] 2 · n H 2 O, where [C(NH 2 ) 3 ] + is guanidinium, details the subtle differences in the Lewis acidity between a lanthanide/actinide pair of similar ionic sizes. Contractions in nitrogen–metal bond lengths between neodymium(III) and americium(III) were observed, while the metal–oxygen bonds remained relatively consistent, highlighting the marginal favorability for actinide complexation over the lanthanides with moderately soft N-donors. Spectroscopic analysis shows significant splitting of many transitions and relatively strong electronic interactions with traditionally low-intensity transitions in the americium complex, as is demonstrated in the 7 F 0 → 7 F 5 transitions. Pressure-induced spectroscopic analysis showed surprisingly little effect on the americium complex, with 5 f →5 f transitions either not shifting or marginally shifting from 2 to 3 nm at 11.93 ± 0.06 GPa─atypical of a soft, N-donor americium complex under pressure. Finally, large voids occupied by water molecules in between the complexes within the crystal structure may be responsible for the lack of pressure response in the 5 f →5 f transitions.

absorption spectroscopy

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.

kilic, Ozgur Ozan [Brookhaven National Laboratory

PyTrac

SAND2025-00635O PyTrac is a software tool that analyzes and visualizes PTRAC event files generated by MCNP 6.3. It converts these files into a graph network that makes it easier to interpret individual histories. The software includes command line tools for viewing the graph data structure in both 2D and 3D formats. PyTrac also integrates with MCNP to run simulations and manage data files. This provides a streamlined approach to analyzing and understanding the complex data generated by MCNP simulations. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Nowack, Aaron [Sandia National Lab. (SNL-CA), Live

Cataloging Legacy Data from the Tritium Systems Test Assembly Program

The Tritium Systems Test Assembly (TSTA) at Los Alamos National Laboratory, operational from 1984 to 2001, was critical in advancing fusion fuel cycle technologies, including tritium storage, gas separation, and pumping. TSTA’s contributions, particularly in safe tritium operations, have influenced subsequent fusion projects. This paper discusses the ongoing effort to digitize and catalog TSTA’s historical data to create a searchable resource for the fusion research community. While the long-term objective is to develop a relational database for structured data management, the project remains in the early phase, with current efforts focused on scanning and indexing physical documents. Initial plans for database implementations are also presented, outlining key considerations for structure, query indexing, and standardization. As digitization progresses, future discussions will refine these implantation details to ensure an efficient and comprehensive system. This initiative aims to preserve critical legacy data, enhance the design of tritium system facilities, and support the next generation of fusion energy research.

42 ENGINEERING

Structural Evidence of Interanionic Hydrogen Bonding in Phosphoric Acid Solutions

Interanionic hydrogen bonding (IAHB) is a noncovalent interaction between like-charged ions that challenges conventional electrostatic understanding. This study provides direct structural evidence of IAHB in concentrated aqueous phosphoric acid (PA) solutions, which exhibit >60% dissociation under these conditions. Oxygen K-edge X-ray absorption fine structure spectroscopy, combined with electron affinity time-dependent density functional theory calculations, reveals the formation of stable, cyclic phosphate-phosphate IAHB dimers at PA concentrations ≥7 M. Extended X-ray absorption fine structure data show distinct long-range ordering consistent with these dimers, and near-edge X-ray absorption fine structure spectra confirm a concentration-dependent transition from monomeric to dimeric species. Energy decomposition analysis through density functional theory shows that the formation of solution-phase IAHB is energetically favored and is attributed to polarization of and the charge transfer between the two fragments driven by the surrounding solvent molecules, in addition to the permanent electrostatics. These findings offers crucial structural insights into the H-bonded networks in concentrated PA, highlighting the critical role of solvent in facilitating anion–anion association.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Evidence for a Single Holocene Paleoseismic Event on the Pajarito Fault, Northern New Mexico

Low-slip rate fault systems tend to be less studied than their high-slip rate counterparts, and paleoseismic techniques used to study them may pose challenges in interpretation that differ from high-slip rate systems. A good example of this is the Pajarito fault system (PFS), a normal fault complex within the Rio Grande rift. Despite numerous previous paleoseismic trenching studies conducted on the PFS between 1990 and 2003, considerable uncertainty remains regarding its Holocene paleoseismic history, particularly for the primary Pajarito fault (PF). To further clarify the PF paleoseismic history, we present data from paleoseismic investigations of 6 trenches at 3 distinct locations along the PF. Though the totality of the age and structural data obtained in this study is complex and not entirely consistent with any one interpretation, a single Holocene paleoearthquake occurring younger than ∼1,600 to 2,300 kcal yr BP is the simplest interpretation. It is possible that the PF records two Holocene events, with a penultimate event 6.9–2.4 kcal yr BP event and the aforementioned most recent event (MRE) between 2.3 and 1.6 kcal yr BP. However, only a single wall of one trench, out of a total of 12 walls in our 6 trenches, provides evidence supporting that interpretation. This study finds evidence of a single late Holocene paleoseismic event on the PF and sparse evidence for 2 Holocene paleoseismic events on the PF and highlights the benefits of logging multiple trench walls to better understand the complexity that results from this low-slip rate, low-deposition-rate fault system.

58 GEOSCIENCES