Search NASA⌕ Search

SEARCH · Search NASA

Results for “structured text”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Thermo-Fluid Modeling Framework for Supercomputer Digital Twins: Part 1, Demonstration at Exascale

A thermo-fluid modeling framework is being developed for ExaDigiT---an open-source framework for developing comprehensive digital twins of liquid-cooled supercomputers. The work is being conducted in two parts, and discussion is divided into two companion papers. The work documented in this paper focuses on the development of a cooling system library in Dymola for the Frontier supercomputer at Oak Ridge National Laboratory. The second part, outlined in a companion paper, focuses on a templating structure called Auto-CSM for easily creating model-agnostic, physics-based thermo-fluid cooling system models for liquid-cooled supercomputers using a text-based schema. The cooling model is being developed using primarily the open-source Transient Simulation Framework of Reconfigurable Models (TRANSFORM) library. The library follows the templating architecture developed within the TRANSFORM library for modeling subsystems. A full-system validation was performed to validate a very simple model that is integrated with the system controls, and the results are presented herein.

Kumar, Vineet↗

Conceptual Structural Design and Analysis of a 20 T Hybrid Cos Dipole for Future Particle Colliders

Here, to reach high collision energy for future high-energy particle colliders, like the Future Circular Collider (FCC) or the Muon Collider, it is required to achieve high field strength of the bending dipoles. Currently, the practical limit for Nb$_{\text{3}}$ Sn technology is around 16 T and, in order to further increase the magnetic field, the superconducting magnet community is considering High Temperature Superconductors (HTS), in particular Bi-2212 and REBCO conductors. However, their relevant higher cost has led the community to consider a hybrid approach where HTS materials are used in the high field region of the coils with so-called insert coils, and Low Temperature Superconductors are involved in the lower field part ($< $ 16 T) with so-called outsert coils. This paper describes the conceptual mechanical design of a 20 T hybrid cos$\theta$ dipole configuration. The high stress levels that the structure is facing due to the high magnetic field are discussed. Moreover, it presents the results of the optimization analysis of the shell-based support structure based on the key-and-bladder technology that provides the azimuthal pre-stress during room temperature assembly and cooldown to cryogenic temperatures. The aim of this work is to present a feasible design that satisfies the stress requirements.

D'Addazio, Marika [Politecnico di Torino (Italy); ↗

A Summary of Advances in Document Summarization from 2023-2024

In computer science, Document Summarization is the task of condensing some quantity of text and related content through automated means. In this document, we review recent literature in text summarization. “Hybrid” extractive-abstractive approaches continue to be explored. Some of the latest efforts have also sought to enable users to adjust summaries with queries or other structure and begun to test reinforcement-learning style agentic LLM-based solutions.

97 MATHEMATICS AND COMPUTING↗

Modeling and Simulation of Electrostatics of Ge$_{\text{1-x}}$Sn$_{\text{x}}$ Layers Grown on Ge Substrates

This work introduces a comprehensive simulation tool that provides a robust 1D Schrödinger – Poisson solver for modeling the electrostatics of heterostructures with an arbitrary number of layers, and non-uniform doping profiles along with the treatment of partial ionization of dopants at low temperatures. The effective masses are derived from the first-principles calculations. The solver is used to characterize three Ge 1-x Sn x /Ge heterostructures with non-uniform doping profiles and determine the subband structure at various temperatures. Here, the simulation results of the sheet carrier densities show excellent agreement with the experimentally extracted data, thus demonstrating the capabilities of the solver.

42 ENGINEERING↗

Observational constraints on early dark energy

In this paper, we review and update constraints on the Early Dark Energy (EDE) model from cosmological data sets, in particular Planck PR3 and PR4 cosmic microwave background (CMB) data and large-scale structure (LSS) data sets including galaxy clustering and weak lensing data from the Dark Energy Survey, Subaru Hyper Suprime-Cam and KiDS+VIKING-450, as well as BOSS/eBOSS galaxy clustering and Lyman-[Formula: see text] forest data. We detail the fit to CMB data, and perform the first analyses of EDE using the CAMSPEC and Hillipop likelihoods for Planck CMB data, rather than Plik, both of which yield a tighter upper bound on the allowed EDE fraction than that found with Plik. We then supplement CMB data with LSS data in a series of new analyses. All these analyses are concordant in their Bayesian preference for [Formula: see text]CDM over EDE, as indicated by marginalized posterior distributions. We perform a series of tests of the impact of priors in these results, and compare with frequentist analyses based on the profile likelihood, finding qualitative agreement with the Bayesian results. All these tests suggest prior volume effects are not a determining factor in analyses of EDE. This work provides both a review of existing constraints and several new analyses.

Astronomy & Astrophysics↗

MnRhBi3: A Cleavable Antiferromagnetic Metal

This dataset contains DFT input and output files supporting the theoretical modeling in the associated publication (Chem. Mater. 2024, 36, 11306-11316). The calculations characterize MnRhBi3, an orthorhombic (Cmmm) van der Waals-layered intermetallic compound that cleaves easily between neighboring Bi layers. The dataset is organized into three calculation types: (i) Bulk: Structural relaxations of the periodic MnRhBi3 crystal in antiferromagnetic (AFM) and ferromagnetic (FM) configurations, using the vdW-DF-optB86b functional. These provide the equilibrium lattice constants, magnetic energy differences (AFM is 0.5 meV/f.u. lower than FM), and magnetic moments (4.4 µB/Mn, 0.17 µB/Rh, 0.18 µB/Bi) reported in Table 1 of the main text. (ii) Slab: Same magnetic configurations computed with an 18 Ang vacuum layer introduced between Bi layers, used to calculate the cleavage energy Ec = 0.56 J/m2 (AFM) and 0.57 J/m2 (FM), establishing MnRhBi3 as a van der Waals-layered material comparable to graphite, MoS2, and CrI3. (iii) ELF: Single-point calculation on the relaxed bulk AFM geometry with LELF=.TRUE., producing the ELFCAR file used to generate electron localization function isosurfaces and contour maps (Fig. 2, main text) showing Bi lone pairs directed into the van der Waals gaps. All folders contain CONTCAR, INCAR, KPOINTS, OUTCAR, and POSCAR. The ELF/ folder additionally contains ELFCAR. Calculations were performed using VASP 6.3.2 with PBE + vdW-DF-optB86b, PAW potentials, and an energy cutoff of 800 eV.

36 MATERIALS SCIENCE↗

Unveiling Highly Sensitive Active Site in Atomically Dispersed Gold Catalysts for Enhanced Ethanol Dehydrogenation

Developing a desirable ethanol dehydrogenation process necessitates a highly efficient and selective catalyst with low cost. Herein, we show that the “complex active site” consisting of atomically dispersed Au atoms with the neighboring oxygen vacancies (Vo) and undercoordinated cation on oxide supports can be prepared and display unique catalytic properties for ethanol dehydrogenation. The “complex active site” Au–Vo–Zr 3+ on Au 1 /ZrO 2 exhibits the highest H 2 production rate, with above 37,964 mol H 2 per mol Au per hour (385 g H 2 $\text{g}^{-1}_{\text{Au}} \text{h}^{–1}$) at 350 °C, which is 3.32, 2.94 and 15.0 times higher than Au 1 /CeO 2 , Au 1 /TiO 2 , and Au 1 /Al 2 O 3 , respectively. Combining experimental and theoretical studies, we demonstrate the structural sensitivity of these complex sites by assessing their selectivity and activity in ethanol dehydrogenation. Our study sheds new light on the design and development of cost-effective and highly efficient catalysts for ethanol dehydrogenation. Fundamentally, atomic-level catalyst design by colocalizing catalytically active metal atoms forming a structure-sensitive “complex site”, is a crucial way to advance from heterogeneous catalysis to molecular catalysis. Finally, our study advanced the understanding of the structure sensitivity of the active site in atomically dispersed catalysts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Text-mined dataset of solid-state syntheses with impurity phases using Large Language Model

Solid-state synthesis is widely used to obtain various inorganic materials, such as battery materials and bulk thermoelectrics. Despite its prevalence, the process remains challenging due to the lack of a general theory and well-understood underlying reaction mechanisms. While prior works have successfully extracted structured datasets from literature, they often neglect product phase purity or yield. In this work, we construct a solid-state synthesis dataset consisting of 80,806 syntheses extracted with a large language model (LLM), including 18,869 reactions with impurity phase(s). Our dataset not only validates expected thermodynamic trends for impurity phase formation but also identifies challenging cases where impurity phases emerge even when the target phase is significantly more stable.

Lee, Sanghoon↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials

Following publication of this article, concerns were raised about the unambiguous identification of the compound structures using diffraction as well as the original claims of material novelty. We acknowledge that the original claims of material novelty were subject to misinterpretation—their intention was to indicate that the materials were new to the prediction platform, not necessarily new to science. The article text has been updated to reflect this in the HTML and PDF versions of the article.

Szymanski, Nathan J. [University of California, Be↗

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING↗

First constraints on the nonperturbative gluon Collins-Soper kernel

The gluon Collins-Soper kernel, which encodes the rapidity evolution of transverse-momentum-dependent gluon distributions, is constrained for the first time in the nonperturbative regime, for transverse momentum scales $q_{T} \in [ 300\text{ MeV}, 1.3\text{ GeV}]$. The constraints are determined in lattice QCD at a close-to-physical pion mass $M_π= 172(3)\text{ MeV}$, a single lattice spacing $a=0.15\text{ fm}$, and next-to-next-to-leading logarithmic matching in Large-Momentum Effective Theory. These results represent the first step toward a controlled determination of the gluon Collins-Soper kernel in QCD, with eventual phenomenological import and relevance to present and future experiments sensitive to the gluon structure of hadronic matter.

Avkhadiev, Artur [Argonne; MIT, Cambridge, CTP]↗

Understanding Event Trajectories Across Massive Temporal Datasets with Word Embeddings and Visualization

In collaboration with researchers from Virginia Tech, Savannah River National Laboratory has continued development of a natural language processing pipeline to identify and extract events of interest from massive open data sources in the domain of worldwide state-sponsored civil nuclear energy. The foundation of the pipeline is built on compass aligned temporal word embedding models, whereby contextual shifts are automatically identified by comparing keyword embedding vectors across successive time windows. Within the approach, a contextual shift indicates the occurrence of a potential event of interest. However, in such a broad topical domain that captures events at a global scale, across various life cycle stages, and across numerous different technology types, a user that is monitoring events may have broad interests in capturing many different event types with varying degrees of signal. As such, the quantity of information that may be returned from an automated event extraction pipeline can be substantial, requiring manual effort to sift through the information to identify any relevant bits of information. Therefore, a more streamlined workflow that aids in directing a user toward specific information at different points in time is necessary. The workflow presented here has been developed with this concept in mind, built on top of the initial prototype event extraction pipeline, whereby a user can analyze temporal text-based data sources at multiple different contextual levels to isolate key points in time and key subdomains captured within a data corpus. Using multiple corpuses that consist of approximately 7 million Tweets and 7 million news articles, the team has extended compass aligned temporal word embedding models to establish an interconnected and hierarchical structure that relates known key words of interest to documents, local topics (i.e., within a time window), and global topics across the corpuses. All of this information is packaged into a visual analytics system that is linked to the information extraction pipeline and enables a user to identify contextual information that describes the evolution of a high dimensional embedding space across time to isolate changes of interest and explore associated events. This report demonstrates the use of these analytics and a means to fuse information across multiple datasets.

97 MATHEMATICS AND COMPUTING↗

ZENN: A thermodynamics-inspired computational framework for heterogeneous data–driven modeling

Traditional entropy-based methods—such as cross-entropy loss in classification problems—have long been essential tools for representing the information uncertainty and physical disorder in data and for developing artificial intelligence algorithms. However, the rapid growth of data across various domains has introduced new challenges, particularly the integration of heterogeneous datasets with intrinsic disparities. To address this, we introduce a zentropy-enhanced neural network (ZENN), extending zentropy theory into the data science domain via intrinsic entropy, enabling more effective learning from heterogeneous data sources. ZENN simultaneously learns both energy and intrinsic entropy components, capturing the underlying structure of multisource data. To support this, we redesign the neural network architecture to better reflect the intrinsic properties and variability inherent in diverse datasets. We demonstrate the effectiveness of ZENN on classification tasks and energy landscape reconstructions, showing its superior generalization capabilities and robustness-particularly in predicting high-order derivatives. In image and text classification tasks, ZENN demonstrates superior generalization by introducing a learnable temperature variable that models latent multisource heterogeneity, allowing it to surpass state-of-the-art models on CIFAR-10/100, BBC News, and AG News. As a practical application in materials science, we employ ZENN to reconstruct the Helmholtz energy landscape of Fe3Pt using data generated from density functional theory and capture key material behaviors, including negative thermal expansion and the critical point in the temperature–pressure space. Overall, this work presents a zentropy-grounded framework for data-driven machine learning, positioning ZENN as a versatile and robust approach for scientific problems involving complex, heterogeneous datasets.

36 MATERIALS SCIENCE↗

Knowledge Graph of RB-Tnseq Data from Fitness Browser (KP-DP1)

Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.

59 BASIC BIOLOGICAL SCIENCES↗

LAMP DTL Scoping Studies (Technical Report)

The present studies are based on the preliminary design efforts, and on the Scoping Studies of the “Strawman” design of LAMP front-end upgrade, referred in the text below to as “Feb.2024 Iteration”, presented in. The main accomplishment of present studies was substantial increase of the fidelity of the beam dynamics simulations in the proposed drift tube linac (DTL). The main accent was on development of the methodology for calculation of the longitudinal (synchrotron) and transverse (betatron) oscillations frequencies (phase advance per focusing period) values and providing the accelerating structure focusing lattice that has safe parameters of such oscillations to avoid unwanted emittance growth and possible beam halo formation. The resulting values of the oscillations phase advances are presented in Table 1, and in Figure 5 in the main body of the report. Special efforts were made to achieve the RF power consumption within limits of the existing RF power system and make sure that DTL fits in the existing tunnel. The LAMP scope does not suggest any additional building and/or tunnel construction. Table 1 summarizes some of these results, as well as Table 3 in the main body of the report.

43 PARTICLE ACCELERATORS↗

Jackson, L., Johnson, M.B., Latrach, A., Grimes, D., Martinez, C., and Mclaughlin, J.F., 2024, Multidisciplinary geotechnical data collection, curation, and analysis for conformity with the regulatory framework for geologic carbon storage in Wyoming, USA: Geological Society of America Abstracts with Programs. Vol. 56, No. 5, 2024, doi: 10.1130/abs/2024AM-405024

Title: Multidisciplinary Geotechnical Data Collection, Curation, and Analysis for Conformity with the Regulatory Framework for Geologic Carbon Storage in Wyoming, USA. Text: Construction and operation of wells for geologic sequestration of carbon dioxide necessitate that they are permitted under the Environmental Protection Agency’s Underground Injection Control Class VI requirements. Class VI wells conform to stringent requirements to ensure long-term safety and integrity of the storage site and the protection of Underground Sources of Drinking Water. Entities pursuing Class VI permitting must provide comprehensive geologic site characterization, including regional geologic structure and stratigraphy, aquifer information, reservoir and confining unit geomechanical properties, geochemical analyses, assessment of trapping capacity and mechanisms, and a variety of other of multidisciplinary geotechnical data. The Wyoming Class VI Site Characterization Database Project is focused on developing a geologic site characterization database of geotechnical information, which has been compiled and verified from established, public databases/entities and scientific literature to expedite Class VI permitting in Sweetwater County within the Greater Green River Basin of southern Wyoming. The preliminary suite of compiled data from 14,000 wells includes 8,000 wells with logs and 7,250 wells with formation tops, ~70 wells with core data (e.g., X-Ray diffraction, petrographic, and petrophysical data), ~2,500 water analyses, ~740 seismic events data, and ~520 bottom-hole temperature measurements. Future work on—and stemming from—this project will include new core analyses, calculation and interpolation of subsurface temperature gradients, mechanical earth models, geochemical simulations, storage capacity estimation, stratigraphic column generation and correlation, and construction of subsurface maps. Finally, this work will help to inspire and facilitate subsurface data compilation and curation beyond Sweetwater County, Wyoming.

42 ENGINEERING↗

Disorder-induced magnetoelastic behaviors of MnTexSbyBi1-x-y alloys

This dataset contains input and output files from density functional theory (DFT) simulations used to study the disorder-induced magnetoelastic behaviors of MnTexSbyBi1-x-y (0 ≤ x + y ≤ 1) alloys and their binary end members MnTe, MnSb, and MnBi. The alloys adopt the hexagonal NiAs-type (nickeline) structure and span ternary (MnTexSb1-x, MnTexBi1-x, MnBixSb1-x), and quaternary compositions across the full MnTe–MnSb–MnBi composition triangle. For each alloy composition, the dataset provides DFT calculations in three magnetic configurations: A-type antiferromagnetic (AFM), C-type AFM, and ferromagnetic (FM). Every magnetic configuration folder contains the fully relaxed crystal structure (CONTCAR), VASP input parameters (INCAR), and the main VASP output file (OUTCAR), from which total electronic energies, Mn magnetic moments, lattice parameters, and percent volume changes between magnetic states are extracted. These data are used to construct compositional phase diagrams, evaluate thermodynamic stability (formability), and map magnetoelastic responses across the alloy space. For A-type AFM and FM configurations, additional data are provided as follows: (i) FORCE_CONSTANTS and thermal_properties.yaml files at the top level of A-type_AFM/ and FM/ folders — present only for compositions marked with an asterisk (*) in Table I of the main text. These are derived from Phonopy finite-displacement calculations on full disordered 128-atom supercells and provide vibrational free energy, entropy (Svib)contribution from explicit disorder calculations. (Table I of the associated main manuscript) (ii) A VCA/ subfolder within A-type_AFM/ and FM/, containing FORCE_CONSTANTS and thermal_properties.yaml from Virtual Crystal Approximation phonon calculations (without spin-orbit coupling). VCA data are available for all compositions and are used to estimate vibrational contributions to the Gibbs free energy across the full composition space. (iii) A SOC/ subfolder containing CONTCAR, INCAR, and OUTCAR from spin-orbit coupling calculations, providing relativistic corrections to electronic energies and lattice parameters (Tables S2–S3 of the SM, and Table I of the main manuscript). (iv) A SOC/VCA/ subfolder containing FORCE_CONSTANTS and thermal_properties.yaml from VCA phonon calculations performed within the SOC framework, combining relativistic and vibrational thermodynamic corrections. The computed properties are used to map the AFM–FM magnetic crossover near MnTe0.75Sb0.25, demonstrate disorder- and spin-induced phonon broadening, identify a semiconductor-to-metal crossover, and quantify the pronounced magnetoelastic volume response near the magnetic phase boundary.

36 MATERIALS SCIENCE↗