Search NASA⌕ Search

SEARCH · Search NASA

Results for “instance selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

If We Build Them, They Will Run: Automated HPC Apps Deployment and Profiling with eBPF in Cloud

The high performance computing (HPC) community is in a period of transition. The rise of AI/ML coupled with a changing landscape of resources deems portability a new metric of performance, and methods to move between on-premises and cloud environments and assess compatibility are paramount. Here we design and test a strategy for bridging the gap between traditional HPC and Kubernetes environments – first containerizing applications, providing automated orchestration to run studies, and packaging the setup with automated means to assess performance using low overhead eXtended Berkeley Packet Filter (eBPF) programs. We first assess different designs for eBPF collection, demonstrating a tradeoff between number of programs deployed on a node and overhead added. We develop 5 low overhead eBPF programs that combine with streaming ML models to assess CPU, futex, TCP, shared memory, and file access across four different builds of an HPC application for CPU and GPU. We use eBPF data to generate insights into the possible underlying etiology of scaling issues. We then assess compatibility of a well-known benchmark, HPCG, across matrices of micro-architectures and optimization levels (217 containers across 24 instance types and over 7500 runs). We provide to the community 30 applications to deploy in our automated setup and perform a scaling study from 4 to a maximum of 256 nodes for both CPU and GPU applications. Finally, we use our gained knowledge about performance to generate compatibility artifacts that are used by a newly developed Kubernetes controller to intelligently select instance type based on optimizing a figure of merit. Along with insights to scaling in this environment with a collection of applications and templates to work from, we provide an overall strategy for approaching HPC application deployment and image selection based on compatibility in cloud.

Computer science↗

Enzymatic Routes to Designer Hemicelluloses for Use in Biobased Materials

Various enzymes can be used to modify the structure of hemicelluloses directly in vivo or following extraction from biomass sources, such as wood and agricultural residues. Generally, these enzymes can contribute to designer hemicelluloses through four main strategies: (1) enzymatic hydrolysis such as selective removal of side groups by glycoside hydrolases (GH) and carbohydrate esterases (CE), (2) enzymatic cross-linking, for instance, the selective addition of side groups by glycosyltransferases (GT) with activated sugars, (3) enzymatic polymerization by glycosynthases (GS) with activated glycosyl donors or transglycosylation, and (4) enzymatic functionalization, particularly via oxidation by carbohydrate oxidoreductases and via amination by amine transaminases. Thus, this Perspective will first highlight enzymes that play a role in regulating the degree of polymerization and side group composition of hemicelluloses, and subsequently, it will explore enzymes that enhance cross-linking capabilities and incorporate novel chemical functionalities into saccharide structures. These enzymatic routes offer a precise way to tailor the properties of hemicelluloses for specific applications in biobased materials, contributing to the development of renewable alternatives to conventional materials derived from fossil fuels.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uniform-in-phase-space data selection with iterative normalizing flows

Improvements in computational and experimental capabilities are rapidly increasing the amount of scientific data that are routinely generated. In applications that are constrained by memory and computational intensity, excessively large datasets may hinder scientific discovery, making data reduction a critical component of data-driven methods. Datasets are growing in two directions: the number of data points and their dimensionality. Whereas dimension reduction typically aims at describing each data sample on lower-dimensional space, the focus here is on reducing the number of data points. A strategy is proposed to select data points such that they uniformly span the phase-space of the data. The algorithm proposed relies on estimating the probability map of the data and using it to construct an acceptance probability. An iterative method is used to accurately estimate the probability of the rare data points when only a small subset of the dataset is used to construct the probability map. Instead of binning the phase-space to estimate the probability map, its functional form is approximated with a normalizing flow. Therefore, the method naturally extends to high-dimensional datasets. The proposed framework is demonstrated as a viable pathway to enable data-efficient machine learning when abundant data are available.

97 MATHEMATICS AND COMPUTING↗

Phase engineering of layered anode materials during ion-intercalation in Van der Waal heterostructures

Transition metal dichalcogenides (TMDs) are a class of 2D materials demonstrating promising properties, such as high capacities and cycling stabilities, making them strong candidates to replace graphitic anodes in lithium-ion batteries. However, certain TMDs, for instance, MoS 2 , undergo a phase transformation from 2H to 1T during intercalation that can affect the mobility of the intercalating ions, the anode voltage, and the reversible capacity. In contrast, select TMDs, for instance, NbS 2 and VS 2 , resist this type of phase transformation during Li-ion intercalation. This manuscript uses density functional theory simulations to investigate the phase transformation of TMD heterostructures during Li-, Na-, and K-ion intercalation. The simulations suggest that while stacking MoS 2 layers with NbS 2 layers is unable to limit this 2H → 1T transformation in MoS 2 during Li-ion intercalation, the interfaces effectively stabilize the 2H phase of MoS 2 during Na- and K-ion intercalation. However, stacking MoS 2 layers with VS 2 is able to suppress the 2H → 1T transformation of MoS 2 during the intercalation of Li, Na, and K-ions. The creation of TMD heterostructures by stacking MoS 2 with layers of non-transforming TMDs also renders theoretical capacities and electrical conductivities that are higher than that of bulk MoS 2 .

25 ENERGY STORAGE↗

MLP-NN vs Gauss-Newton files

-Input files for randomly selected subset data (30 instances) from primary dipole-dipole forward modeling data.-Inversion results in surfer grid format as well as in .DAT format.-Scatterplots for MLP-NN vs Gauss-Newton present in the excel sheet.

58 GEOSCIENCES↗

Theoretical assessments of Pd–PdO phase transformation and its impacts on H 2 O 2 synthesis and decomposition pathways

The direct synthesis of H 2 O 2 from O 2 and H 2 provides a green pathway to produce H 2 O 2 , a popular industrial oxidant. Here, in this study, we theoretically investigate the effects of Pd oxidation states, coordination environments, and particle sizes on primary H 2 O 2 selectivities, assessed by calculating the ratio of rate constants for the formation of H 2 O 2 (via OOH* reduction; k O–H ) and the decomposition of OOH* (via O–O cleavage; k O–O ). For Pd metals, the k O–H /k O–O ratio decreased from 10 -4 for Pd(111) to 10 -10 for the Pd 13 cluster at 300 K, indicating poorer H 2 O 2 selectivity as Pd particle size decreases and low primary selectivities for H 2 O 2 overall. As the oxygen chemical potential increases and metals form surface and bulk oxides, the perturbation of Pd–Pd ensemble sites by lattice O atoms results in selectivities that become dramatically higher than unity. For instance, at 300 K, the k O–H /k O–O ratio increases significantly from 10 -4 to 10 9 to 10 16 as Pd(111) oxidizes to Pd 5 O 4 /Pd(111) and to PdO(100), respectively. In contrast, such selectivity enhancements are not observed for surface and bulk oxides that persistently contain rows of more metallic, undercoordinated Pd–Pd ensemble sites, such as PdO(101)/Pd(100) and PdO(101). These Pd–Pd ensembles are also absent when smaller Pd nanoparticles fully oxidize, indicating that smaller PdO clusters can be more selective for H 2 O 2 synthesis. These trends for primary H 2 O 2 selectivities were found to inversely correlate with trends for H 2 O 2 decomposition rates via O–O bond cleavage, demonstrating that catalysts with high primary H 2 O 2 selectivity can also hinder H 2 O 2 decomposition. Ab initio thermodynamic calculations are used to estimate the thermodynamically favored phase among Pd, PdO/Pd and PdO in O 2 , H 2 O 2 /H 2 O, and O 2 /H 2 environments. These results are combined to show that smaller Pd nanoparticles are more prone to be oxidized at lower oxygen chemical potentials, upon which they become more selective than larger Pd particles for H 2 O 2 synthesis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ninterp: N-dimensional interpolation Rust crate [SWR-25-25]

The ninterp crate provides multivariate interpolation over a regular, sorted, nonrepeating grid of any dimensionality. A variety of interpolation strategies are implemented, however more are likely to be added. Extrapolation beyond the range of the supplied coordinates is supported for 1-D linear interpolators, using the slope of the nearby points. There are hard-coded interpolators for lower dimensionalities (up to N = 3) for better runtime performance. All interpolation is handled through instances of the Interpolator enum, with the selected tuple variant containing relevant data. Interpolation is executed by calling Interpolator::interpolate.

Carow, Kyle↗

Benchmark of numerical modeling approaches on the systematic performance evaluation of wave energy converters

Different numerical modeling methods have been developed and applied to evaluate a variety of performance indicators of wave energy converters (WECs), including the power performance, structural loads, levelized cost of energy, etc. Based on the modeling fidelity, the commonly used numerical modeling approaches can be classified as linear modeling, weakly nonlinear modeling and fully nonlinear modeling approaches. Each method differs in accuracy and computational efficiency, making them suitable for different stages of WEC design. However, the selection of modeling approach could significantly impact evaluation outcomes. For instance, simplified linear models may underestimate structural loads or overestimate energy production in some operational conditions, potentially leading to less cost-effective designs. Given the widespread utilization of these models, it is essential to understand the uncertainties brought by them in performance evaluations. This work is dedicated to benchmarking different linear-potential-flow-based numerical models for evaluating the systematic performance of WECs. Three representative numerical modeling approaches are considered in this work, including linear frequency-domain modeling, statistically linearized spectral-domain modeling and Cummins equation-based nonlinear time-domain modeling. A generic point absorber WEC is considered as the research reference in this work, and different sea sites are taken into account. The numerical models are utilized to predict critical performance indicators, including power performance, the annual energy production, the capacity factor, the levelized cost of energy and the PTO fatigue loads. By comparing the results, this work identifies the uncertainties associated with different modeling approaches in evaluating WEC performance.

Fatigue↗

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology↗

Polyhedral Relaxations for Optimal Pump Scheduling of Potable Water Distribution Networks

The classic pump scheduling or optimal water flow (OWF) problem for water distribution networks (WDNs) minimizes the cost of power consumption for a given WDN over a fixed time horizon. In its exact form, the OWF is a computationally challenging mixed-integer nonlinear program (MINLP). It is complicated by nonlinear equality constraints that model network physics, discrete variables that model operational controls, and intertemporal constraints that model changes to storage devices. To address the computational challenges of the OWF, this paper develops tight polyhedral relaxations of the original MINLP, derives novel valid inequalities (or cuts) using duality theory, and implements novel optimization-based bound tightening and cut generation procedures. The efficacy of each new method is rigorously evaluated by measuring empirical improvements in OWF primal and dual bounds over 45 literature instances. The evaluation suggests that our relaxation improvements, model strengthening techniques, and a thoughtfully selected polyhedral relaxation partitioning scheme can substantially improve OWF primal and dual bounds, especially when compared with similar relaxation-based techniques that do not leverage these new methods.

bound tightening↗

Quasi-static to Dynamic Mechanical Response and Microstructure Development of Tantalum-Tungsten Alloys

Lawrence Livermore National Laboratory (LLNL) is interested the quasi-static to dynamic mechanical response and microstructure evolution of tantalum-tungsten (Ta-W) alloys (Ta-2.5W, Ta-5W, and Ta-10W, wt.%) made by conventional wrought processing and additive manufacturing (AM). This work scope was performed at the Colorado School of Mines (Mines) and included quasi-static (e.g., 10 -3 s -1 ) mechanical testing in tension and compression, along with selected high strain rate (Kolsky) pressure bar testing in compression (e.g., 10 3 s -1 ), with and without temperature variations in some instances. Complementary microstructure characterization was performed on undeformed and deformed samples to understand the role of processing on microstructural evolution and the deformation mechanisms that impact the mechanical response with variations in strain rate, temperature, and strain state (e.g., tension versus compression in selected examples). The wrought material provided by LLNL from Viridis Materials was found to have unrecrystallized regions within the microstructure, which led to unexpected results relative to previously reported properties for Ta-W. The AM Ta-2.5W (wt.%) material provided by LLNL was found to have higher compressive strength than wrought Ta-2.5W (wt.%), which is hypothesized to be due to differences in the crystallographic texture between the two materials. This project partially supported several postdocs and graduate students at Mines.

36 MATERIALS SCIENCE↗

Characterize and control the multi-range structure of solution-phase systems with resonant x-ray scattering

This work aims at demonstrating the feasibility and impact of Anomalous X-ray Scattering (AXS) in characterizing the multi-range structures of solution-state systems. We will show preliminary investigation of long-range correlations in concentrated aqueous electrolytes with a combination of AXS and molecular dynamics (MD) simulations. We will also start exploring the capability of AXS to capture intra- and inter-molecular structure of dilute molecular systems, and specifically its sensitivity to the chemical environment surrounding active metal sites. We anticipate that the development of this method will help us to understand ion solvation and transport, which affect the performances of, for instance, electrocatalytic cells, as well as to identify/control the intrinsic structural factors that lead to selectivity, efficiency, and stability control during catalysis.

36 MATERIALS SCIENCE↗

A fast computational framework for the design of solvent-based plastic recycling processes

Multicomponent plastics cannot be processed using mechanical recycling technologies, hindering efforts to deal with plastic waste. Multicomponent plastics include multilayer plastic films, which are widely used for food and healthcare packaging. Multilayer films combine several layers (potentially dozens) of different polymers to protect products from external factors (e.g., oxygen, water, temperature, shock, and light). Solvent-based separation processes have emerged as a promising alternative to recycle these complex materials. For instance, the Solvent-Targeted Recovery and Precipitation (STRAP TM ) process uses sequential solvent washes to selectively dissolve and separate constituent polymers from multicomponent plastic waste, including films. STRAP TM process design (separation sequence, type of solvents, and operating conditions) changes significantly depending on the design of the multilayer plastic film (e.g., number, types, and proportions of polymers). The ability to quickly quantify the economic and environmental benefits of diverse STRAP TM process designs is essential to accelerate the development of sustainable recycling processes and more recyclable multilayer film products. In this work, we present a fast computational framework that integrates molecular-scale models, process modeling, and techno-economic and life cycle analysis to quickly evaluate STRAP TM designs. The computational framework is general and can be used to study the processing of complex multilayer plastic waste streams that contain many layers. Furthermore, we highlight the different uses of the framework via targeted case studies.

Computational framework↗

Rapid Evaluation Framework for the CMIP7 Assessment Fast Track

As Earth system models (ESMs) grow in complexity and in volume of output data, there is an increasing need for rapid, comprehensive evaluation of their scientific performance. The upcoming Assessment Fast Track for the Seventh Phase of the Coupled Model Intercomparison Project (CMIP7) will require expeditious response for model analyses designed to inform and drive integrated Earth system assessments. To meet this challenge, the Rapid Evaluation Framework (REF), a community-driven platform for benchmarking and performance assessment of ESMs, was designed and developed. The initial implementation of the REF, constructed to meet the near-term needs of the CMIP7 Assessment Fast Track, builds upon four disparate community evaluation and benchmarking tools that are coupled together using the Coordinated Model Evaluation Capabilities (CMEC) framework. The REF runs within a containerized workflow for portability and reproducibility and is aimed at generating and organizing diagnostics covering a variety of model variables. The REF leverages well documented observational datasets to provide assessments of model fidelity across a collection of diagnostics. All diagnostics were identified and selected with community involvement and consultation. Operational integration with the Earth System Grid Federation (ESGF) will permit automated execution of the REF for selected diagnostics as soon as model output data are published on ESGF by the originating modeling centers. The REF is designed to be portable across a range of current computational platforms to facilitate use by modeling centers for assessing the evolution of model versions or gauging the relative performance of CMIP simulations before being published on ESGF. When integrated into production simulation workflows, results from the REF provide immediate quantitative feedback that allows model developers and scientists to quickly identify model biases and performance issues. After the REF is released to the community, its subsequent development and support will be prioritized by an international consortium of scientists and engineers, enabling a broader impact across Earth science disciplines. For instance, the REF will facilitate improvements to models and will enhance confidence in model projections through process-based selection of models based on their performance with respect to observations. Production of reproducible diagnostics and community-based assessments are key features of the REF. Furthermore, providing interoperability with existing evaluation packages assures that contributions from previous community efforts will be available for use in future model intercomparison projects.

Hoffman, Forrest [ORNL] (ORCID:0000000158024134)↗

Neutrino Physics with Deep Learning on NOvA

The NOvA experiment has made both νμ \nu_\mu disappearance and νe \nu_e appearance measurements in Fermilab's NuMI beam, and is working on cross section measurements using near detector data. At the core of NOvA's measurements is the use of deep learning algorithms for identification and reconstruction of the neutrino flavor and energy. These algorithms, used for the first time on NOvA in 2016, yielded large improvements in selection efficiency, and will be applied to our first anti-neutrino results to be released this year. Presented here is the extension of our deep learning efforts for identification of neutrino signal events, final state identification, single particle tagging, and reconstruction using instance segmentation techniques. We will describe the new implementations of modified Convolutional Neural Networks for anti-neutrino events, single particles and their performance for analysis final states selection, standard candle measurements, and reconstruction.

Psihas, Fernanda [Indiana U.]↗

Instantiation of the Damara Tern Platform for Advanced Materials and Manufacturing Technologies (AMMT) Program Collaborative Data Management

This work package focused on deploying an instance of the Damara Tern platform to support AMMT collaborative research activities. The objectives were to provide selected AMMT collaborators with access to a shared environment for capturing operations, trackables, and associated metadata, and to implement data entry functionalities that reflect site-specific procedures. Key activities included creating configurable, schema-driven entry forms and validating the data collection process. The report details the deployment process, the platform infrastructure, and the implemented data entry workflows, providing a reference for end users and establishing a foundation for future production-scale deployments.

36 MATERIALS SCIENCE↗

Segmentation method comparison for residual fiber length measurement across tiled microscopy images

Fiber length distribution (FLD), in part, governs mechanical properties in discontinuous fiber composites, yet manual measurement methods limit the high-throughput characterization needed for materials design optimization. This study compares deep learning segmentation approaches for automated FLD measurement in large-field microscopy, evaluating how method choice affects the microstructural descriptors used in structure-property-processing relationships. A critical challenge is that high-resolution microscopy images (10,000×10,000 pixels) must be tiled for deep learning analysis, fragmenting fibers at boundaries. We demonstrate that segmentation method proves crucial for measurement accuracy. For example, instance segmentation with Slicing Aided Hyper Inference (SAHI) preserves individual fiber integrity across tiles while semantic segmentation prioritizes speed. Comparing against manual measurement of extracted carbon fibers, YOLOv11-SAHI matched manual ground truth (238 μm weighted mean) with 40x speedup (4.5 vs 167 minutes per image). U-Net provides rapid quantification although it is at the cost of reduced accuracy due only reliably measuring stand-alone fibers. Our comparative analysis reveals that instance segmentation with SAHI better preserves length measurements while semantic segmentation prioritizes speed, providing empirical guidance for method selection. The characterization provides essential inputs for mechanical property prediction models and inverse design workflows, accelerating composite materials development cycles.

Additive manufacturing↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗