Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data structures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Molecular model of TFIIH recruitment to the transcription-coupled repair machinery

Transcription-coupled repair (TCR) is a vital nucleotide excision repair sub-pathway that removes DNA lesions from actively transcribed DNA strands. Binding of CSB to lesion-stalled RNA Polymerase II (Pol II) initiates TCR by triggering the recruitment of downstream repair factors. Yet it remains unknown how transcription factor IIH (TFIIH) is recruited to the intact TCR complex. Combining existing structural data with AlphaFold predictions, we build an integrative model of the initial TFIIH-bound TCR complex. We show how TFIIH can be first recruited in an open repair-inhibited conformation, which requires subsequent CAK module removal and conformational closure to process damaged DNA. In our model, CSB, CSA, UVSSA, elongation factor 1 (ELOF1), and specific Pol II and UVSSA-bound ubiquitin moieties come together to provide interaction interfaces needed for TFIIH recruitment. STK19 acts as a linchpin of the assembly, orienting the incoming TFIIH and bridging Pol II to core TCR factors and DNA. Molecular simulations of the TCR-associated CRL4CSA ubiquitin ligase complex unveil the interplay of segmental DDB1 flexibility, continuous Cullin4A flexibility, and the key role of ELOF1 for Pol II ubiquitination that enables TCR. Collectively, these findings elucidate the coordinated assembly of repair proteins in early TCR.

Paul, Tanmoy↗

Question-answering system extracts information on injection drug use from clinical notes

Background. Injection drug use (IDU) can increase mortality and morbidity. Therefore, identifying IDU early and initiating harm reduction interventions can benefit individuals at risk. However, extracting IDU behaviors from patients’ electronic health records (EHR) is difficult because there is no other structured data available, such as International Classification of Disease (ICD) codes, and IDU is most often documented in unstructured free-text clinical notes. Although natural language processing can efficiently extract this information from unstructured data, there are no validated tools. Methods. Here, to address this gap in clinical information, we design a question-answering (QA) framework to extract information on IDU from clinical notes for use in clinical operations. Our framework involves two main steps: (1) generating a gold-standard QA dataset and (2) developing and testing the QA model. We use 2323 clinical notes of 1145 patients curated from the US Department of Veterans Affairs (VA) Corporate Data Warehouse to construct the gold-standard dataset for developing and evaluating the QA model. We also demonstrate the QA model’s ability to extract IDU-related information from temporally out-of-distribution data. Results. Here, we show that for a strict match between gold-standard and predicted answers, the QA model achieves a 51.65% F1 score. For a relaxed match between the gold-standard and predicted answers, the QA model obtains a 78.03% F1 score, along with 85.38% Precision and 79.02% Recall scores. Moreover, the QA model demonstrates consistent performance when subjected to temporally out-of-distribution data. Conclusions. Our study introduces a QA framework designed to extract IDU information from clinical notes, aiming to enhance the accurate and efficient detection of people who inject drugs, extract relevant information, and ultimately facilitate informed patient care.

60 APPLIED LIFE SCIENCES↗

Efficient generation of grids and traversal graphs in compositional spaces towards exploration and path planning

Abstract Diverse disciplines across science and engineering deal with problems related to compositions, which exist in non-Euclidean simplex spaces, rendering many standard tools inaccurate or inefficient. This work explores such spaces conceptually in the context of materials discovery, quantifies their computational feasibility, and implements several essential methods specific to simplex spaces through a new high-performance open-source library . Most significantly, we derive and implement an algorithm for constructing a novel n-dimensional simplex graph data structure, containing all discretized compositions and possible neighbor-to-neighbor transitions. Critically, no distance or neighborhood calculations are performed, instead leveraging pure combinatorics and order in procedurally generated simplex grids, keeping the algorithm $${\mathcal{O}}(N)$$ O ( N ) , with minimal memory, enabling rapid construction of graphs with billions of transitions in seconds. Additionally, we demonstrate how such graph representations can be combined to homogeneously express complex path-planning problems, while facilitating efficient deployment of existing high-performance gradient descent, graph traversal, and other optimization algorithms.

Krajewski, Adam M. (ORCID:0000000222660099)↗

Delocalization error poisons the density-functional many-body expansion

The many-body expansion is a fragment-based approach to large-scale quantum chemistry that partitions a single monolithic calculation into manageable subsystems. This technique is increasingly being used as a basis for fitting classical force fields to electronic structure data, especially for water and aqueous ions, and for machine learning. Here, we show that the many-body expansion based on semilocal density functional theory affords wild oscillations and runaway error accumulation for ion–water interactions, typified by F − (H 2 O) N with N ≳ 15. We attribute these oscillations to self-interaction error in the density-functional approximation. The effect is minor or negligible in small water clusters, explaining why it has not been noticed previously, but grows to catastrophic proportion in clusters that are only moderately larger. This behavior can be counteracted with hybrid functionals but only if the fraction of exact exchange is ≳50%, whereas modern meta-generalized gradient approximations including ωB97X-V, SCAN, and SCAN0 are insufficient to eliminate divergent behavior. Other mitigation strategies including counterpoise correction, density correction (i.e., exchange–correlation functionals evaluated atop Hartree–Fock densities), and dielectric continuum boundary conditions do little to curtail the problematic oscillations. In contrast, energy-based screening to cull unimportant subsystems can successfully forestall divergent behavior. These results suggest that extreme caution is warranted when the many-body expansion is combined with density functional theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cholesterol-dependent enzyme activity of human TSPO1

The amino acid sequence of the tryptophan-rich sensory proteins (TSPO) is substantially conserved throughout all kingdoms of life. Human mitochondrial TSPO1 (HsTSPO1) binds to porphyrins and steroids, although its interactions with these molecules remains unknown.HsTSPO1 is associated with numerous physiological and pathological disorders, but the underlying molecular mechanisms are unknown. Here, we disclose the finding of human mitochondrial TSPO as a cholesterol-dependent protoporphyrin IX oxygenase. The results of our biochemical characterization are consistent with structural data and evolutionary analysis. The dependence ofHsTSPO1 activity on cholesterol may be the result of the coevolution of this membrane protein with the membrane system. Our study provides a molecular foundation for comprehending the various roles played by mitochondrial TSPO in normal physiological and pathological situations.

Science & Technology - Other Topics↗

Three-dimensional lattice modulations in the charge density wave system Lu 2 Ir 3 Si 5

Using total and resonant x-ray scattering coupled to large-scale computer modeling, we study the lattice modulations in the complex charge density wave (CDW) material Lu 2 ⁢Ir 3 ⁢Si 5 . Here, we find that it is a unique quantum system where periodic lattice modulations related to emergent CDW order occur in three orthogonal atomic planes of the crystal lattice, leading to the emergence of an unusual three-dimensional (3D) pattern of short and long Ir-Ir and Lu-Lu bonds. The 3D character of observed lattice modulations explains the largely isotropic character of the changes in the electronic properties occurring when the CDW order sets in, demonstrating the strong electron-lattice coupling in Lu 2 ⁢Ir 3 ⁢Si 5 . The result is supported by DFT calculations based on the experimental structure data. Altogether, our work provides strong evidence for the presence of a relationship between the dimensionality of emergent lattice distortions and that of concurrent changes in the electronic properties of CDW materials. The relationship may need to be accounted for when these materials are explored for practical applications.

36 MATERIALS SCIENCE↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗

A Survey on Privacy in Graph Neural Networks: Attacks, Preservation, and Applications

Graph Neural Networks (GNNs) have gained significant attention owing to their ability to handle graph-structured data and the improvement in practical applications. However, many of these models prioritize high utility performance, such as accuracy, with a lack of privacy consideration, which is a major concern in modern society where privacy attacks are rampant. To address this issue, researchers have started to develop privacy-preserving GNNs. Despite this progress, there is a lack of a comprehensive overview of the attacks and the techniques for preserving privacy in the graph domain. In this survey, we aim to address this gap by summarizing the attacks on graph data according to the targeted information, categorizing the privacy preservation techniques in GNNs, and reviewing the datasets and applications that could be used for analyzing/solving privacy issues in GNNs. We also outline potential directions for future research in order to build better privacy-preserving GNNs.

97 MATHEMATICS AND COMPUTING↗

Study of electron-induced chemical transformations in polymers

In extreme ultraviolet (EUV) photoresist exposure, the primary and secondary electrons drive chemistry rather than the EUV photons themselves. These electrons have a wide range of energies below approximately 80 eV, which are capable of complex network of reactions during exposure. To better understand the ability of electrons of different energies within the EUV primary and secondary electron range, we want to characterize and compare the chemistry induced in pure polymer films by direct exposure to electrons. Thin films of poly(tert-butyl methacrylate), poly(methyl methacrylate), and poly(4-hydroxystyrene) were exposed to a 20 to 80 eV electron beam. Outgassing during exposure was characterized in-situ using a quadrupole residual gas analyzer. The thickness changes were measured using ellipsometry and chemical bond structure data were collected using Fourier-transform infrared spectroscopy (FTIR) after exposure to compare different exposure conditions. Poly(4-hydroxystyrene) demonstrated stability during electron exposures. Exposures of the other two materials led to outgassing of protecting groups, the intensity of which decayed in time. Outgassing, FTIR, and thickness loss data exhibited approximately linear relationships to each other.

Mueller, Maximillian↗

Deciphering spin-parity assignments of nuclear levels

Spin-parity assignments of nuclear levels are critical for understanding nuclear structure and reactions. However, inconsistent notation conventions and ambiguous reporting in research papers often lead to confusion and misinterpretations. Here, this paper examines the policies of the Evaluated Nuclear Structure Data File (ENSDF) and the evaluations by Endt and collaborators, highlighting key differences in their approaches to spin-parity notation. Sources of confusion are identified, including ambiguous use of strong and weak arguments and the conflation of new experimental results with prior constraints. Recommendations are provided to improve clarity and consistency in reporting spin-parity assignments, emphasizing the need for explicit notation conventions, clear differentiation of argument strengths, community education, and separate reporting of new findings. These steps aim to enhance the accuracy and utility of nuclear data for both researchers and evaluators.

Experimental Nuclear Physics↗

AI-Powered Knowledge Graphs for Neuromorphic and Energy-Efficient Computing

The surge in scientific literature obscures breakthroughs and hinders the discovery of new research paths. We propose an artificial intelligence (AI) powered framework using large language models (LLMs) and knowledge graphs (KGs) to automate parts of scientific discovery, focusing on energy-efficient AI circuits. Our hybrid approach combines LLMs, structured data, and ontology-based reasoning to construct a comprehensive knowledge graph that integrates insights across computational neuroscience, spiking neuron models, learning rules, architectural motifs, and neuromorphic device technologies. This multi-domain representation enables the generation of hypotheses that connect biological function with implementable, energy-efficient hardware architectures. Using KG embeddings and graph neural networks, the framework generates hypotheses for novel circuits, validates them through optimization on exascale HPC systems, and with tools like SuperNeuro and Fugu, the most promising designs will be prototyped in hardware. This open-source system aims to accelerate discoveries and bridging neuroscience with hardware innovation, drive collaboration, and unlock new opportunities in low-power AI computing.

Gautam, Ashish [ORNL]↗

Lowering and Runtime Support for Fortran’s Multi-Image Parallel Features using LLVM Flang, PRIF, and Caffeine

This paper provides an overview of the multi-image parallel features in Fortran 2023 and their implementation in the LLVM flang compiler and the Caffeine parallel runtime library. The features of interest support a Single-Program, Multiple-Data (SPMD) programming model based on executing multiple “images”, each of which is a program instance. The features also support a Partitioned Global Address Space (PGAS) in the form of “coarray” distributed data structures. The paper discusses the lowering of multi-image features to the Parallel Runtime Interface for Fortran (PRIF) and the implementation of PRIF in the Caffeine parallel runtime library. This paper also provides an early view into the design of a new multi-image dialect of the LLVM Multi-Level Intermediate Representation (MLIR). We describe validation and testing of the resulting software stack, and demonstrate that performance compares favorably to another open-source compiler and runtime library: GNU Compiler Collection (GCC) gfortran and OpenCoarrays, respectively.

Bonachea, Dan↗

The ArborX Library: Version 2.0

This article provides an overview of the 2.0 release of the ArborX library, a performance portable geometric search library based on Kokkos. We describe the major changes in ArborX 2.0 including a new interface for the library to support a wider range of user problems, new search data structures (brute force and distributed), support for user functions to be executed on the results (callbacks), and an expanded set of the supported algorithms (ray tracing and clustering).

GPU↗

Accessible Content Optimization for Research Needs (ACORN)

ACORN employs a set of automated processes for informing and/or enforcing defined content schemas to create standardized and highly structured data. Because of its standardized data source, ACORN easily applies computer automation to generate communication assets such as PDFs, Powerpoint presentations, and web pages. Built using the memory-safe Rust programming language, ACORN is portable and accessible for use on any Windows, Mac, or Linux machine.

Wohlgemuth, JasonHoward [Oak Ridge National Labora↗

ATcT — Active Thermochemical Tables Python Interface

SF-25-140 atct is a lightweight, Python client for the ATcT v1 API that enables programmatic access to high-accuracy thermochemical data and turnkey reaction-enthalpy analysis. The package implements full v1 endpoint coverage (species lookup by ATcT ID, name, formula, SMILES, InChI, CAS RN; covariance queries; health checks) with robust error handling, retries, and environment-based configuration for local/production endpoints. Beyond data retrieval, atct provides rigorously implemented reaction calculators that propagate uncertainties via either (i) a conventional independent-errors method (0 K or 298.15 K) or (ii) covariance-aware propagation using provided covariances at 298.15 K. Typed data classes ensure transparent, reproducible data structures and carry ATcT Thermochemical Network (TN) version identifiers for provenance. Dual import paths and comprehensive examples facilitate integration into research pipelines, enabling reproducible thermochemical calculations, automated validation, and downstream method development.

Bross, DavidHamilton [Argonne National Laboratory ↗

Fortran mimetic abstraction language (Formal) v0.1.

The Fortran mimetic abstraction language ("Formal") is a domain-specific language (DSL) embedded in Fortran 202Y [1]. Formal provides novel software abstractions for simulating phenomena governed by the partial differential equations (PDEs) of vector and tensor calculus. Such equations model an extremely broad set of physical phenomena, ranging from atmospheric winds to light propagation. Formal's data structures and algorithms mimic in form and behavior continuous functions and operators. Formal supports these mathematical constructs using mimetic discretizations that define a discrete calculus satisfying various tensor calculus theorems, thereby ensuring high-fidelity representations of the physics being modeled. [2] Formal 0.1.0 also lays a foundation for the future use of Fortran 202Y type-safe templates to facilitate the formal verification of tensor contractions in computational physics and artificial intelligence [3]. [1] "Fortran 202Y" is Fortran standard committee's informal designation for the next Fortran revision, which will likely be "Fortran 2028". [2] Corbino, J. and Castillo, J. (2020) Journal of Computational and Applied Mathematics, https://doi.org/10.1016/j.cam.2019.06.042. [3] Haveraaen, M., Järvi, J., & Rouson, D. (2019). Reflecting on Generics for Fortran. https://j3-fortran.org/doc/year/19/19-188.pdf.

Rouson, Damian [Lawrence Berkeley National Laborat↗

Chimera D-Series Gravitational Wave Emission Sourced from Neutrino Anisotropy

Gravitational wave data sourced from the time-dependent anisotropic neutrino emission, as well as the time-dependent fluid quadrupole motion, in the Chimera D-Series three-dimensional core collapse supernova simulations. Data from three models initiated from three different progenitors are presented: D9.6-3D, D15-3D, and D25-3D. Please see the README for more information about the data structure and progenitors.

79 ASTRONOMY AND ASTROPHYSICS↗