Search NASASearch

SEARCH · Search NASA

Results for “Feature extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Contaminant Investigation and Pre‐Processing Opportunities for Textile‐To‐Textile Recycling

Millions of metric tons of textiles are landfilled or incinerated each year in the United States, with less than 1% of textiles recycled into new clothing or fabrics. To counter this trend, a growing number of companies and researchers are exploring how a circular economy can be applied to support textile‐to‐textile recycling. A significant barrier they face comes down to quickly and efficiently extracting pure feedstock material from post‐consumer garments that feature a mix of natural and synthetic fibers. Textile recyclers prefer pure feedstocks, as working with mixed sources typically means lower throughput, higher risk of equipment failure, and diminished business margins. To facilitate a circular economy for textiles, methods, and technologies are needed that can efficiently separate out materials and contaminants from end‐of‐life textiles to increase the flow of pure feedstocks to recyclers. This paper summarizes findings from interviews with a cross section of textile recyclers and from a review of literature to define basic feedstock requirements. In addition to our qualitative research, we deconstruct a bale of post‐consumer textiles and analyze them using computer‐vision imaging, Fourier transform infrared spectroscopy (FTIR), and machine learning. The resulting data are used to set system‐level design inputs for an automated contaminant removal system to process post‐consumer clothing into appropriate feedstocks for recycling. To set the system's levels for automated real‐time near‐infrared analysis, we identify the minimum percentage of primary material that any single garment in a load of used clothing must contain for the average of the full output stream to meet the target purity levels of recyclers. Here, the envisioned automated system can also address undesirable trace materials that might contaminate the processed stream by using imaging cameras coupled with artificial intelligence to identify sections of clothing for de‐trimming. Proof‐of‐concept machine learning algorithms are evaluated to locate and identify trims or garment areas with hidden contaminant materials. Integrating these methods into automated textile cutting systems can provide a cost‐effective means for increasing feedstock purity from used clothing, which can advance circularity for textiles by helping recyclers to reach production volumes and quality targets that were not possible solely with manual dismantling operations.

Parsons, Ryan [Rochester Institute of Technology,

Moltensaltpropnet

MoltenSaltPropnet is a physics-informed machine learning framework that aims to predict the thermophysical properties of molten fluoride and chloride salt mixtures, which are crucial for the design and safety of Generation IV molten salt reactors. The code processes data from the Molten-Salt Thermal Properties Database (MSTDB-TP) and the Janz compendium, converting critically evaluated correlations into fast, differentiable surrogate models for density, viscosity, thermal conductivity, and heat capacity across 448 distinct salt systems. The implementation consists of several key components: 1. Data Curation: The code parses and cleans the raw data, normalizing elemental mole fractions and extracting relevant regression coefficients for various thermophysical properties. 2. Feature Engineering: It generates fixed-length numerical descriptors that encapsulate the composition and temperature, incorporating polynomial interaction terms and dimensionality-reduction techniques to optimize model performance. 3. Coefficient Learning: Four different machine learning architectures are employed: a deep residual network (ResNet), a Kolmogorov–Arnold network (KAN), a sparsity-inducing neural network (SNN), and classical regression models. Each model learns to predict coefficients that define the temperature-dependent correlations for the thermophysical properties. 4. Property Reconstruction: The predicted coefficients are used to compute temperature-dependent property values, ensuring positivity and monotonic trends through a composite loss function that enforces physical constraints. 5. User Interface: An open-source web application enables users to filter the database, train task-specific models, and visualize the results, allowing for rapid exploration of candidate salt mixtures. MoltenSaltPropnet bridges the gap between limited experimental data and high-fidelity reactor simulations, providing a powerful tool for researchers in the field of molten salt reactors and advanced nuclear energy systems.

Retamales, Mauricio Eduardo Tano [Idaho National L

Validation of the DESI DR2 Ly$α$ forest full-shape analysis

We present the validation of the Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) Lyman-$α$ (Ly$α$) forest full-shape analysis. This analysis combines three-dimensional Ly$α$ forest auto-correlations and cross-correlations with quasars to extract information from both the baryon acoustic oscillation (BAO) feature and the broadband clustering signal, with primary emphasis on the Alcock-Paczynski (AP) measurement. Compared to the DESI DR1 analysis, the DR2 validation uses substantially larger and more realistic mock datasets, including CoLoRe 2LPT and AbacusSummit Ly$α$ forest simulations. The modeling framework is also improved through analytic marginalization over small scales ($<10$$h^{-1}$Mpc) and the impact of ultraviolet background fluctuations. The validation program was completed prior to unblinding and defines quantitative requirements for the cosmological parameters of interest, which are evaluated using hundreds of mock realizations. We further test the analysis through independent fits to the auto- and cross-correlations, multiple catalog splits, and a broad suite of analysis and modeling variations applied to both mocks and blinded observational data. We find that the BAO and AP parameters satisfy all validation requirements and remain stable across all tests. In contrast, mock studies reveal a significant bias in the inferred growth-rate parameter $fσ_8$, leading us to exclude this measurement from the final analysis. The consistency across mocks, data splits, and robustness tests demonstrates that the DR2 Ly$α$ full-shape analysis provides a reliable and substantially improved broadband AP measurement over previous Ly$α$ forest studies.

Herbold, M. [Chicago U., KICP; Ohio State U.] (ORC

Broadband Light Extraction from Near-Surface NV Centers Using Crystalline-Silicon Antennas

We use crystalline silicon (Si) antennas to efficiently extract broadband single-photon fluorescence from shallow nitrogen-vacancy (NV) centers in diamond into free space. Our design features relatively easy-to-pattern high-index Si resonators on the diamond surface to boost photon extraction by overcoming total internal reflection and Fresnel reflection at the diamond-air interface and providing modest Purcell enhancement, without etching or otherwise damaging the diamond surface. In simulations, ∼17 times more single photons are collected from a single NV center compared to the case without the antenna; in experiments, we observe an enhancement of ∼9 times, limited by spatial alignment between the NV and the antenna. Furthermore, our approach can be readily applied to other color centers in diamond, and more generally to the extraction of light from quantum emitters in wide-bandgap materials.

Antennas

Hydropower Fish Passage Webmap

The National Fish Passage Webmap application provides an environment that allows users to visualize information information on fish passage facility existence, type, and direction at hydropower developments across the conterminous United States. It was developed through collaborative partnerships with fish passage engineers and biologists at both the US Fish and Wildlife Service (USFWS) and the National Marine Fisheries Service (NMFS), and hydropower experts at the Low Impact Hydropower Institute (LIHI). Data on fish passage facilities at hydropower features were compiled from numerous sources including published and non-published datasets, published reports, email communications with federal and state resource managers and hydropower operators, and by extracting information from regulatory documents within the FERC eLibrary. The number of sources for a given feature varied, which occasionally resulted in conflicting information regarding the existence of fish passage facilities or in the type or sub-type of passage technologies. Such discrepancies were reviewed and resolved individually, based on the weight of evidence or, when available, on direct observations from information providers or aerial imagery.

13 HYDRO ENERGY

Development of a bench-scale dissolution concept for the direct extraction of nuclear fuel

Current used nuclear fuel reprocessing efforts utilize a hydrometallurgical approach in which the UNF is dissolved in hot nitric acid followed by solvent extraction into an organic solvent to harvest target nuclides. Previous studies have shown that the dissolution and loading process could be combined into a single organic dissolution/extraction step, producing loaded organic in a single step process. This single step process also includes the advantage of selectively targeting key nuclides in the dissolution while leaving undesirable constituents as part of the undissolved solids. This process is referred to hereafter as direct extraction. Ongoing research from multiple national labs has proven the effectiveness of this technique at research scale. Therefore, potential methods to implement direct extraction at both bench and industrial scale have been developed. The key features of potential dissolver system designs were identified via the team at Pacific Northwest National Laboratory and several designs based on industrial counterparts were assessed for feasibility. This report summarizes the advantages and disadvantages of multiple methods, concluding with a path forward to create multiple unique dissolver designs. The first design will be a single stage recirculating eductor mixer. The second design recommendation is a stator rotor static mixing flow loop design. Each dissolver could be utilized separately, simultaneously, or in series to answer questions surrounding reaction kinetics including residence time, provide a proof of concept for targeted extractions of specific nuclides, and inform needs for industrial scale implementation

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]

Machine Learning Framework for Characterizing Processing–Structure Relationship in Block Copolymer Thin Films

The morphology of block copolymers (BCPs) critically influences material properties and applications. This work introduces a machine learning (ML)-enabled, high-throughput framework for analyzing grazing incidence small-angle X-ray scattering (GISAXS) data and atomic force microscopy (AFM) images to characterize BCP thin film morphology. A convolutional neural network was trained to classify AFM images by surface features, achieving 97% testing accuracy. Classified images were then analyzed to extract 2D grain size measurements from the samples in a high-throughput manner. ML models were trained to predict domain orientation based on processing parameters such as solvent ratio, additive type, and additive ratio. GISAXS-based properties were predicted with strong performances (R 2 > 0.75), while AFM-based property predictions were less accurate (R 2 < 0.60), likely due to the localized nature of AFM measurements compared to the bulk information captured by GISAXS. Beyond model performance, interpretability was addressed using SHapley Additive exPlanations (SHAP). SHAP analysis revealed that the additive ratio had the largest impact on morphological predictions, where additive provides the BCP chains with increased volume to rearrange into thermodynamically favorable morphologies. This interpretability helps validate model predictions and offers insight into parameter importance. Altogether, the presented framework combining high-throughput characterization and interpretable ML offers an approach to exploring and optimizing BCP thin film morphology across a broad processing landscape.

36 MATERIALS SCIENCE

Robust Biaxial Anisotropy and Switchable Néel Vectors in LaFeO 3 Epitaxial Films

Antiferromagnets with highly stable but switchable Néel vectors are desired for antiferromagnetic spintronics with ultrafast speed and terahertz frequencies. Electrical switching of antiferromagnetic insulators has been demonstrated using binary antiferromagnets, while large families of complex antiferromagnets such as perovskites are largely unexplored. Here, we show that epitaxial LaFeO 3 thin films on SrTiO 3 (001) exhibit clear, robust biaxial anisotropy with a spin-flop field of a few tesla. Angular-dependent spin-Hall magnetoresistance (SMR) characterizations of Pt/LaFeO 3 bilayers with the current channel along SrTiO 3 [100] and [110] reveal distinct, intriguing shapes and field dependence. Simulations using a macrospin model accurately describe the main behavior and fine features of the SMR data from which key antiferromagnetic parameters are extracted. Furthermore, remanent SMR measurement confirms the high fidelity of the Néel vector along either easy axis of the biaxial anisotropy, indicating that epitaxial films of LaFeO 3 and potentially other perovskite antiferromagnets offer an attractive platform for antiferromagnetic spintronics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

AI Design Assistant

SAND2025-01930O The AI Design Assistant uses ChatGPT to provide a natural language interface to airfoil analysis tools (XFOIL). Most of the code base is glue code, connecting XFOIL (a tool for analyzing airfoils) to the OpenAI interface. Among the more novel features are an airfoil geometry class and methods on how to extract detailed data from XFOIL. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Karcher, Cody

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics

Forecasting Battery Electrode Performance via Electrochemical Fluorescence Microscopy and Machine-Learning

Predicting lithium-ion battery performance is hindered by microscale electrode heterogeneities invisible to conventional diagnostics. Here, we combine electrochemical fluorescence microscopy (EFM), which maps electronic connectivity by visualizing an electrofluorophore reaction distribution, with a multitask ElasticNet regression to forecast discharge capacity from spatial heterogeneity. Analyzing 196 images from six pilot-scale LiNi 0.5 Mn 0.3 Co 0.2 O 2 cathodes with varying carbon loadings, we extract 62 descriptors that capture morphology and texture. A compact five-feature model predicts capacity across eight discharge rates, achieving a per-target R 2 of up to 0.63 and an overall R 2 of 0.92, with a mean absolute percentage error of less than 2%. This performance rivals impedance-based approaches while avoiding their reliance on postformation data and incomplete electronic network information. Our facile and rapid, image-driven method may enable electrode quality control upstream of costly cell assembly to offer a transformative tool for data-driven battery research and manufacturing.

battery electrodes

Estimation and Visualization of Isosurface Uncertainty from Linear and High-Order Interpolation Methods

Isosurface visualization is fundamental for exploring and analyzing 3D volumetric data. Marching cubes (MC) algorithms with linear interpolation are commonly used for isosurface extraction and visualization. Although linear interpolation is easy to implement, it has limitations when the underlying data is complex and high-order, which is the case for most real-world data. Linear interpolation can output vertices at the wrong location. Its inability to deal with sharp features and features smaller than grid cells can lead to an incorrect isosurface with holes and broken pieces. Despite these limitations, isosurface visualizations typically do not include insight into the spatial location and the magnitude of these errors. We utilize high-order interpolation methods with MC algorithms and interactive visualization to highlight these uncertainties. Our visualization tool helps identify the regions of high interpolation errors. It also allows users to query local areas for details and compare the differences between isosurfaces from different interpolation methods. In addition, we employ high-order methods to identify and reconstruct possible features that linear methods cannot detect. We showcase how our visualization tool helps explore and understand the extracted isosurface errors through synthetic and real-world data.

Ouermi, Timbwaoga

Collision Tracking in OpenMC: Methods and Applications in Neutron Noise, Neutron Imaging, Time-of-Flight, and Multiplicity Counting

We present the development and application of a collision tracking feature within the OpenMC Monte Carlo particle transport code, designed for diverse applications such as neutron spectroscopy, scatter camera system, neutron noise, and multiplicity counting simulations. This feature enables the tracking of individual particle collisions, with potential applications in nuclear nonproliferation, reactor physics, and nuclear security. Additionally, the feature holds potential for the calibration of neutron detectors, specifically in converting light output into energy deposited within the detectors. The implementation consists of a set of filters—such as reaction type, energy, cell, and material—that constrain the set of collisions that are tracked, extensions to the Python API to enable simple input specification, and support for writing either OpenMC’s native HDF5-based format or the Monte Carlo particle list format. This feature was added to the official OpenMC release in version 0.15.3. In this work, the feature will be applied to showcase scenarios such as time-of-flight simulations, scatter-camera imaging for neutron source localization, neutron-noise analysis to extract integral kinetic parameters such as the prompt decay constant α, and multiplicity counting to estimate the mass of special nuclear materials. Ultimately, this feature aims to expand the application scope of open-source Monte Carlo particle transport codes such as OpenMC.

Monte Carlo code

Recovering high-purity uranyl nitrate from simulated used nuclear fuel dissolver solutions by crystallization: rejecting technetium

The separation of U from Tc and other problematic fission product elements like Mo and Ru, along with Sr, Zr, Cs, and Nd, has been achieved via the crystallization of uranyl nitrate hexahydrate (UNH). Rejection of technetium as pertechnetate anion ( 99 TcO 4 – ) is an especially important feature of this system, as it otherwise tends to follow U (VI) in extractive separations. It also raises the salient question regarding why this oxoanion cannot replace nitrate within the crystalline lattice of UNH. Results showed high-yield (>90 %), high-purity (>99 %) recovery of U as UNH from solutions containing 99 TcO 4 – by simple reduction of temperature from 60°C to 20°C. There was no observable interaction of 99 TcO 4 – with UO 2 2+ . The addition of other cations like, Sr 2+ , Zr 4+ , Cs + , and Nd 3+ , also did not form secondary, contaminant solid phases, leaving the > 99 % of the fission product elements in the mother liquor, while the U was recovered at > 90 %. Similarly, Mo and Ru, when added to the mixture, were shown to behave as the other fission-product elements, remaining in the mother liquor during crystallization. As a result, DFT calculations showed that, despite the higher binding strength of TcO 4 – , HMoO 4 – , and BiO 3 – with the UO 2 2+ cation compared to NO 3 – , the hydrogen-bonding network of the two coordinated ions and four waters of hydration in the UNH crystal structure is the driving force for the high specificity of this separation.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Systems, devices, and methods for authenticating millimeter wave devices

Systems, devices, and methods are described for millimeter wave device authentication. A system may include one or more access points. Each access point of the one or more access points is configured to extract, from one or more beam patterns generated via a client device, a beam feature associated with the client device. Each access point may also be configured to transmit the beam feature. The system may also include a server communicatively coupled to the one or more access points and including a database for storing known beam features. The server may be configured to receive the beam feature associated with the client device from at least one access point of the one or more access points. Also, the server may be configured to authenticate the client device in response to the received beam feature matching a known beam feature stored in the at least one database.

Bhuyan, Arupjyoti

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING

MaTableGPT: GPT‐Based Table Data Extractor from Materials Science Literature

Abstract Efficiently extracting data from tables in the scientific literature is pivotal for building large‐scale databases. However, the tables reported in materials science papers exist in highly diverse forms; thus, rule‐based extractions are an ineffective approach. To overcome this challenge, the study presents MaTableGPT, which is a GPT‐based table data extractor from the materials science literature. MaTableGPT features key strategies of table data representation and table splitting for better GPT comprehension and filtering hallucinated information through follow‐up questions. When applied to a vast volume of water splitting catalysis literature, MaTableGPT achieves an extraction accuracy (total F1 score) of up to 96.8%. Through comprehensive evaluations of the GPT usage cost, labeling cost, and extraction accuracy for the learning methods of zero‐shot, few‐shot, and fine‐tuning, the study presents a Pareto‐front mapping where the few‐shot learning method is found to be the most balanced solution owing to both its high extraction accuracy (total F1 score >95%) and low cost (GPT usage cost of 5.97 US dollars and labeling cost of 10 I/O paired examples). The statistical analyses conducted on the database generated by MaTableGPT revealed valuable insights into the distribution of the overpotential and elemental utilization across the reported catalysts in the water splitting literature.

Yi, Gyeong Hoon [Computational Science Research Ce