Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scalable Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Impact of vertical and seasonal variation in leaf traits on simulating soybean canopy photosynthesis via 1D and 3D modeling

Accurate modeling of photosynthesis is crucial for predicting crop productivity and quantifying the carbon cycle in agroecosystems. Leaf traits are essential inputs for modeling canopy photosynthesis. Yet, many existing models still use fixed plant functional type (PTF)-based values to parameterize leaf traits under a big-leaf or two-big-leaf assumption, neglecting their vertical profiles and seasonal changes. This simplification may introduce significant uncertainties in estimating gross primary productivity (GPP). In this study, we simulated soybean GPP and tested the effects of vertical and seasonal variation in three key leaf photosynthetic traits: the maximum carboxylation rate at 25 °C (Vcmax 25 ), leaf chlorophyll content (LCC), and leaf mass per area (LMA) in the 1D-SCOPE and 3D-Helios models. Weekly field measurements were conducted during the growing season of 2024 to support the simulation. We designed ten leaf trait parameterization schemes by incorporating different combinations of vertical profiles and seasonal changes, while assuming homogeneous canopy architecture in both models. Our results revealed that Vcmax 25 vertical and seasonal variation had the strongest influence on simulated GPP in both 1D and 3D models, while LCC and LMA effects were minimal. Particularly, the scheme with an empirically parameterized Vcmax 25 profile achieved comparable performance to the scheme with the measured Vcmax 25 profile. Both 1D-SCOPE and 3D-Helios accurately modeled GPP (SCOPE: R 2 = 0.87, Bias = 0.55 µmol m⁻² s⁻¹; Helios: R 2 = 0.9, Bias = 0.22 µmol m⁻² s⁻¹) under the most complex scheme, and their responses to vertical and seasonal variation in leaf traits were consistent, demonstrating the robustness of our findings. Based on our findings, we propose a scalable framework for parameterizing leaf traits to improve GPP simulations. This study contributes to improving the representation of leaf trait dynamics in canopy-level photosynthesis models, potentially enhancing our ability to predict crop productivity and understand agroecosystem carbon dynamics.

54 ENVIRONMENTAL SCIENCES↗

Single-Photon Counting Detector Scalability for High Photon Efficiency Optical Communications Links

For high photon-efficiency deep space or low power optical communications links, such as the Orion Artemis-2 Optical Communications System (O2O) project, the received optical signal is attenuated to the extent that single- photon detectors are required. For direct-detection receivers operating at 1.55 µm wavelength, single-photon detectors including Geiger-mode InGaAs avalanche photon diodes (APDs), and in particular superconducting nanowire single-photon detectors (SNSPDs) offer the highest sensitivity and fastest detection speeds. However, these photon detectors exhibit a recovery time between registered input pulses, effectively reducing the detection efficiency over the recovery interval, resulting in missed photon detections, reduced count rate, and ultimately limiting the achievable data rate. A method to overcome this limitation is to divide the received optical signal into multiple detectors in parallel. Here we analyze this approach for a receiver designed to receive a high photon efficiency serially concatenated pulse position modulation (SCPPM) input waveform. From measured count rate and efficiency data using commercial SNSPDs, we apply a model from which we determine the effective detection efficiency, or blocking loss, for different input signal rates. We analyze the scalability of adding detectors in parallel for different modulation orders and background levels to achieve desired data rates. Finally we show tradeoffs between the number of detectors and the required received optical power, useful for real link design considerations.

Vyhnalek, Brian E.↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Acoustic space occupancy: Combining ecoacoustics and lidar to model biodiversity variation and detection bias across heterogeneous landscapes

There is global interest in quantifying changing biodiversity in human-modified landscapes. Ecoacoustics may offer a promising pathway for supporting multi-taxa monitoring, but its scalability has been hampered by the sonic complexity of biodiverse ecosystems and the imperfect detectability of animal-generated sounds. The acoustic signature of a habitat, or soundscape, contains information about multiple taxa and may circumvent species identification, but robust statistical technology for characterizing community-level attributes is lacking. Here, we present the Acoustic Space Occupancy Model, a flexible hierarchical framework designed to account for detection artifacts from acoustic surveys in order to model biologically relevant variation in acoustic space use among community assemblages. We illustrate its utility in a biologically and structurally diverse Amazon frontier forest landscape, a valuable test case for modeling biodiversity variation and acoustic attenuation from vegetation density. We use complementary airborne lidar data to capture aspects of 3D forest structure hypothesized to influence community composition and acoustic signal detection. Our novel analytic framework permitted us to model both the assembly and detectability of soundscapes using lidar-derived estimates of forest structure. Our empirical predictions were consistent with physical models of frequency-dependent attenuation, and we estimated that the probability of observing animal activity in the frequency channel most vulnerable to acoustic attenuation varied by over 60%, depending on vegetation density. There were also large differences in the biotic use of acoustic space predicted for intact and degraded forest habitats, with notable differences in the soundscape channels predominantly occupied by insects. This study advances the utility of ecoacoustics by providing a robust modeling framework for addressing detection bias from remote audio surveys while preserving the rich dimensionality of soundscape data, which may be critical for inferring biological patterns pertinent to multiple taxonomic groups in the tropics. Our methodology paves the way for greater integration of remotely sensed observations with high-throughput biodiversity data to help bring routine, multi-taxa monitoring to scale in dynamic and diverse landscapes.

Airborne lidar↗

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

General Aviation Citizen Science Study to Help Tackle Remote Sensing of Harmful Algal Blooms (HABs)

We present a new, low-cost approach, based on volunteer pilots conducting high-resolution aerial imaging, to help document the onset, growth, and outbreak of harmful algal blooms (HABs) and related water quality issues in central and western Lake Erie. In this model study, volunteer private pilots acting as citizen scientists frequently flew over 200 mi of Lake Erie coastline, its islands, and freshwater estuaries, taking high-quality aerial photographs and videos. The photographs were taken in the nadir (vertical) position in red, green, and blue (RGB) and near-infrared (NIR) every 5 s with rugged, commercially available built-in Global Positioning System (GPS) cameras. The high-definition (HD) videos in 1080p format were taken continuously in an oblique forward direction. The unobstructed, georeferenced, high-resolution images, and HD videos can provide an early warning of ensuing HAB events to coastal communities and freshwater resource managers. The scientists and academic researchers can use the data to compliment a collection of in situ water measurements, matching satellite imagery, and help develop advanced airborne instrumentation, and validation of their algorithms. This data may help develop empirical models, which may lead to the next steps in predicting a HAB event as some watershed observed events changed the water quality such as particle size, sedimentation, color, mineralogy, and turbidity delivered to the Lake site. This paper shows the efficacy and scalability of citizen science (CS) aerial imaging as a complimentary tool for rapid emergency response in HABs monitoring, land and vegetation management, and scientific studies. This study can serve as a model for monitoring/management of freshwater and marine aquatic systems.

Ansari, Rafat R.↗

The search for high-entropy fuel-cell catalysts using disorder descriptors

The transition to a hydrogen economy depends on efficient, affordable catalysts for fuel cells. Platinum—the industry standard for fuel-cell electrodes—is costly and scarce, highlighting the need for practical alternatives. High-entropy alloys offer vast compositional diversity and tunable properties that can mitigate these issues, yet their chemical complexity and configurational disorder have hindered rational discovery. Here, we introduce a data-driven framework that couples machine learning with first-principles disorder descriptors—including the entropy forming ability, disordered enthalpy-entropy descriptor, and electronic-structure similarity metrics to platinum—to predict alloy synthesizability and catalytic performance. These descriptors are applied for the first time in the context of fuel-cell catalyst discovery. The workflow rapidly screens more than 20 000 compositions and identifies several platinum-free candidates that are economically viable, readily scalable, and exhibit promising predicted activity. These results demonstrate that disorder descriptors are reliably predicted by machine learning models and can be effectively integrated into materials-discovery pipelines, accelerating innovation across complex compositional spaces.

fuel-cell catalysts↗

Federated learning for 2D synchrotron x-ray diffractometry: a cross-institutional approach for phase quantification of Ti–6Al–4V alloy

High-energy Two dimensional (2D) synchrotron x-ray diffractometry provides important insights into the atomistic structure and phase evolution of materials, yet traditional analysis methods remain complex, knowledge-intensive, and computationally demanding. Deep-learning models offer a powerful alternative for automating their analysis. Institutions that hold these datasets may be unwilling to share their data due to privacy and security policies, as well as the challenges associated with large-scale data transfer. As a result, models trained on local datasets often perform well only on their own data but exhibit bias and poor generalization across different instruments or facilities. To overcome these limitations, we explore federated learning (FL) for 2D synchrotron diffractograms, enabling collaborative model training without exchanging raw data. In this study, 2D synchrotron diffractograms of Ti–6Al–4V alloy collected from two independent facilities are used to train convolutional neural networks for predicting the β-phase volume fraction. Experimental results show that federated global models significantly outperform locally trained models in terms of generalization and achieve accuracy comparable to centralized trained models. These findings demonstrate the potential of FL to enable secure, cross-institutional collaboration and enhance the scalability of deep-learning-based materials characterization.

36 MATERIALS SCIENCE↗

BoBa

BoBa is a C++ software library for working with large matrices, tensors, and tensor decompositions. The library provides tools for dense matrix and tensor operations, tensor decompositions, and tensor decomposition methods that support modern CPU and GPU architectures. It includes portable abstractions for linear algebra, tensor algebra, and multidimensional computation. BoBa is intended for scientific computing applications that involve large multidimensional data sets or high dimensional mathematical models. Its capabilities support tasks such as data compression, linear algebra, efficient numerical computation, and the development of scalable algorithms for heterogeneous hardware. Tutorials, tests, and example applications are included to help users learn and apply the library.

Yao, Jin [Lawrence Livermore National Laboratory (↗

Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs

This paper presents a methodology for debugging the performance of message-passing programs on both tightly coupled and loosely coupled distributed-memory machines. The AIMS (Automated Instrumentation and Monitoring System) toolkit, a suite of software tools for measurement and analysis of performance, is introduced and its application illustrated using several benchmark programs drawn from the field of computational fluid dynamics. AIMS includes (i) Xinstrument, a powerful source-code instrumentor, which supports both Fortran77 and C as well as a number of different message-passing libraries including Intel's NX Thinking Machines' CMMD, and PVM; (ii) Monitor, a library of timestamping and trace -collection routines that run on supercomputers (such as Intel's iPSC/860, Delta, and Paragon and Thinking Machines' CM5) as well as on networks of workstations (including Convex Cluster and SparcStations connected by a LAN); (iii) Visualization Kernel, a trace-animation facility that supports source-code clickback, simultaneous visualization of computation and communication patterns, as well as analysis of data movements; (iv) Statistics Kernel, an advanced profiling facility, that associates a variety of performance data with various syntactic components of a parallel program; (v) Index Kernel, a diagnostic tool that helps pinpoint performance bottlenecks through the use of abstract indices; (vi) Modeling Kernel, a facility for automated modeling of message-passing programs that supports both simulation -based and analytical approaches to performance prediction and scalability analysis; (vii) Intrusion Compensator, a utility for recovering true performance from observed performance by removing the overheads of monitoring and their effects on the communication pattern of the program; and (viii) Compatibility Tools, that convert AIMS-generated traces into formats used by other performance-visualization tools, such as ParaGraph, Pablo, and certain AVS/Explorer modules.

Yan, Jerry C.↗

Development of High-Performance Graphene-HgCdTe Detector Technology for Mid-Wave Infrared Applications

A high-performance graphene-based HgCdTe detector technology is being developed for sensing over the mid-wave infrared (MWIR) band for NASA Earth Science, defense, and commercial applications. This technology involves the integration of graphene into HgCdTe photodetectors that combines the best of both materials and allows for higher MWIR(2-5 m) detection performance compared to photodetectors using only HgCdTe material. The interfacial barrier between the HgCdTe-based absorber and the graphene layer reduces recombination of photogenerated carriers in the detector. The graphene layer also acts as high mobility channel that whisks away carriers before they recombine, further enhancing the detector performance. Likewise, HgCdTe has shown promise for the development of MWIR detectors with improvements in carrier mobility and lifetime. The room temperature operational capability of HgCdTe-based detectors and arrays can help minimize size, weight, power and cost for MWIR sensing applications such as remote sensing and earth observation, e.g., in smaller satellite platforms. The objective of this work is to demonstrate graphene-based HgCdTe room temperature MWIR detectors and arrays through modeling, material development, and device optimization. The primary driver for this technology development is the enablement of a scalable, low cost, low power, and small footprint infrared technology component that offers high performance, while opening doors for new earth observation measurement capabilities.

Sood, Ashok K.↗

MESA: Scalable Runtime Verification Tool Using Actors

This work presents our runtime verification approach implemented by the tool MESA (MEssage-based System Analysis) which allows for using concurrent monitors to check for properties specified in linear temporal logic and finite state machines.We employ the actor programing model to implement MESA where monitors are captured by concurrent actors that communicate via messaging. The paper also presents a case study where MESA is used to monitor flights in National Airspace System of United States using live air traffic data stream. The case study which motivated this work in the first place shows that our approach is effective.We also perform empirical study by conducting experiments using monitoring systems with different numbers of concurrent monitors and different layers of indexing.This paper describes our experiments, evaluates our results,and discusses challenges faced during the study. The evaluation shows our approach is scalable.

runtime verification, concurrency, actor programin↗

A Simple Method for High-Lift Propeller Conceptual Design

In this paper, we present a simple method for designing propellers that are placed upstream of the leading edge of a wing in order to augment lift. Because the primary purpose of these "high-lift propellers" is to increase lift rather than produce thrust, these props are best viewed as a form of high-lift device; consequently, they should be designed differently than traditional propellers. We present a theory that describes how these props can be designed to provide a relatively uniform axial velocity increase, which is hypothesized to be advantageous for lift augmentation based on a literature survey. Computational modeling indicates that such propellers can generate the same average induced axial velocity while consuming less power and producing less thrust than conventional propeller designs. For an example problem based on specifications for NASA's Scalable Convergent Electric Propulsion Technology and Operations Research (SCEPTOR) flight demonstrator, a propeller designed with the new method requires approximately 15% less power and produces approximately 11% less thrust than one designed for minimum induced loss. Higher-order modeling and/or wind tunnel testing are needed to verify the predicted performance.

Patterson, Michael↗

Transplatformer: translating toxicogenomic profiles between generations of platforms

Background Transcriptomic profiling technologies have advanced the analysis of biological and toxicological responses. However, substantial differences in probe design, dynamic range, gene coverage, and preprocessing pipelines across platforms introduce artifacts that limit cross-study integration and hinder the reuse of historical datasets. We aim to develop computational methods for accurate cross-platform translation to maximize the value of legacy resources. Results We present TransPlatformer a deep learning framework for translating gene expression profiles across heterogeneous toxicogenomics platforms. TransPlatformer employs a novel attention-based architecture to map high-dimensional fold-change vectors from legacy microarray technologies to current platforms. Models are trained and evaluated using DrugMatrix, spanning three technological generations. We investigate mixed-tissue, single-tissue, and cross-tissue training paradigms and benchmark performance against multilayer perceptron and matrix-completion baselines. In mixed-tissue training, TransPlatformer achieves a greater than 50% reduction in mean absolute error (0.043 vs. 0.09) and nearly doubles Pearson correlation ( ≈ 0.71 vs. 0.37) relative to baseline methods. Importantly, TransPlatformer preserves rare but biologically meaningful over- and under-expressed signals, with mean absolute error below 0.22. Single-tissue models yield further improvements for well-represented organs, such as a 10% reduction in liver mean absolute error, while underscoring the need for data augmentation strategies in low-sample tissues.ra Conclusions TransPlatformer provides an effective and scalable computational solution for cross-platform transcriptomic translation. By enabling biologically faithful harmonization of gene expression data, the proposed approach facilitates the reuse of legacy toxicogenomics datasets, enhances downstream biomarker discovery, and supports more reproducible predictive modeling in toxicology.

59 BASIC BIOLOGICAL SCIENCES↗

Cybersecurity Center for Offshore Wind Energy (Final Project Report)

This project establishes a Cybersecurity Center for Offshore Wind Energy with the objective of designing and operating a cyber-physical testbed for wind energy farms (WEFs) that enables comprehensive cybersecurity research. The testbed incorporates a Supervisory Control and Data Acquisition (SCADA) system connected to turbine models via industrial-grade programmable logic controllers (PLCs) and remote terminal units (RTUs). It supports side-channel data acquisition, implementation and analysis of various cyberattack scenarios, and development of attack detection, mitigation, and best-practice guidance tailored to wind energy systems. During the project, the team expanded the number and fidelity of mathematical turbine models (MTMs), integrated these models with SCADA infrastructure, and deployed a scaled physical turbine and associated sensors. High-resolution operational and side-channel data streams were collected and used to refine machine-learning (ML)-based attack detection systems and to extend the WindCRAFT framework to multi-turbine threat scenarios. The project demonstrated a realistic, scalable environment for evaluating cyber threats, validated attack detection approaches using enriched datasets, and identified new multi-turbine and inter-turbine communication attack vectors. The resulting testbed, models, and security mechanisms provide a foundation for ongoing R&D and deployment of cyber-resilient offshore wind energy systems.

17 WIND ENERGY↗