Fluxion: A Scalable Graph-Based Resource Model for HPC Scheduling Challenges
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
ORBIT-2 is a scalable foundation model for global, hyper-resolution climate and weather downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with 𝑅2 scores in range of 0.98–0.99 against observation data.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
This paper describes a scalable structural model suitable for Hybrid Wing Body (HWB) centerbody analysis and optimization. The geometry of the centerbody and primary wing structure is based on a Vehicle Sketch Pad (VSP) surface model of the aircraft and a FLOPS compatible parameterization of the centerbody. Structural analysis, optimization, and weight calculation are based on a Nastran finite element model of the primary HWB structural components, featuring centerbody, mid section, and outboard wing. Different centerbody designs like single bay or multi-bay options are analyzed and weight calculations are compared to current FLOPS results. For proper structural sizing and weight estimation, internal pressure and maneuver flight loads are applied. Results are presented for aerodynamic loads, deformations, and centerbody weight.
The rising number of small unmanned aerial vehicles (UAVs) expected in the next decade will enable a new series of commercial, service, and military operations in low altitude airspace as well as above densely populated areas. These operations may include on-demand delivery, medical transportation services, law enforcement operations, traffic surveillance and many more. Such unprecedented scenarios create the need for robust, efficient ways to monitor the UAV state in time to guarantee safety and mitigate contingencies throughout the operations. This work proposes a generalized monitoring and prediction methodology that utilizes realtime measurements of an autonomous UAV following a series of way-points. Two different methods, based on sinusoidal acceleration profiles and high-order splines, are utilized to generate the predicted path. The monitoring approach includes dynamic trajectory re-planning in the event of unexpected detour or hovering of the UAV during flight. It can be further extended to different vehicle types, to quantify uncertainty affecting the state variables, e.g., aerodynamic and other environmental effects, and can also be implemented to prognosticate safety-critical metrics which depend on the estimated flight path and required thrust. The proposed framework is implemented on a simplified, scalable UAV modeling and control system traversing 3D trajectories. Results presented include examples of real-time predictions of the UAV trajectories during flight and a critical analysis of the proposed scenarios under uncertainty constraints.
High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.
This poster presents a novel modular, scalability receiver design for various Concentrating Solar Thermal (CST)-Concentrating Solar Power (CSP) applications and scale of economics.
Improving our understanding of hurricane inter-annual variability and the impact of climate change (e.g., doubling CO2 and/or global warming) on hurricanes brings both scientific and computational challenges to researchers. As hurricane dynamics involves multiscale interactions among synoptic-scale flows, mesoscale vortices, and small-scale cloud motions, an ideal numerical model suitable for hurricane studies should demonstrate its capabilities in simulating these interactions. The newly-developed multiscale modeling framework (MMF, Tao et al., 2007) and the substantial computing power by the NASA Columbia supercomputer show promise in pursuing the related studies, as the MMF inherits the advantages of two NASA state-of-the-art modeling components: the GEOS4/fvGCM and 2D GCEs. This article focuses on the computational issues and proposes a revised methodology to improve the MMF's performance and scalability. It is shown that this prototype implementation enables 12-fold performance improvements with 364 CPUs, thereby making it more feasible to study hurricane climate.
Societal level macro models of social behavior do not sufficiently capture nuances needed to adequately represent the dynamics of person-to-person interactions. Likewise, individual agent level micro models have limited scalability - even minute parameter changes can drastically affect a model's response characteristics. This work presents an approach that uses agent-based modeling to represent detailed intra- and inter-personal interactions, as well as a system dynamics model to integrate societal-level influences via reciprocating functions. A Cognitive Network Model (CNM) is proposed as a method of quantitatively characterizing cognitive mechanisms at the intra-individual level. To capture the rich dynamics of interpersonal communication for the propagation of beliefs and attitudes, a Socio-Cognitive Network Model (SCNM) is presented. The SCNM uses socio-cognitive tie strength to regulate how agents influence--and are influenced by--one another's beliefs during social interactions. We then present experimental results which support the use of this network analytical approach, and we discuss its applicability towards characterizing and understanding human information processing.
Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronizat and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.
This work will focus on developing the capabilities and validating the models for a sub transmission network with multiple feeders and microgrids. To achieve this scale of Hardware-in-the-loop (HitL) simulation, it is necessary to federate and collaborate. The work aims to design the large-scale feeder models to allow federation with complementary testbeds in the future. The feeder would be designed to be reconfigurable to put the system into a variety of modes. Aggregators models will be included in each distribution network’s federate to take control actions and interact with the management systems. Lastly, the feeder model will support large scale resilience studies involving complex Distributed Energy Resources (DER) controls, microgrid studies and emulation of complex data flows in future grid architectures.
Explore the source record for details and available documents.
A comprehensive understanding of the meteorological and microphysical nature of Mediterranean storms requires a combination of in situ data analysis, radar data analysis, and satellite data analysis, effectively integrated with numerical modeling studies at various scales. An important aspect of understanding microphysical controls of severe storms, is first understanding the meteorological controls under which a storm has evolved, and then using that information to help characterize the dominant microphysical processes. For hazardous Mediterranean storms, highlighted by the October 5-6, 1998 Friuli flood event in northern Italy, a comprehensive microphysical interpretation requires an understanding of the multiple phases of storm evolution. This involves intense convective development, Sratiform decay, orographic lifting, and sloped frontal lifting processes, as well as the associated vertical motions and thermodynamical instabilities governing physical processes that effect details of the size distributions and fall rates of the various types of hydrometeors found within the storm environment. This talk overviews the microphysical elements of a severe Mediterranean storm in such a context, investigated with the aid of TRMM satellite and other remote sensing measurements, but guided by a nonhydrostatic mesoscale model simulation of the Friuli flood event. The data analysis for this paper was conducted by my research groups at the Global Hydrology and Climate Center in Huntsville, AL and Florida State University in Tallahassee, and in collaboration with Dr. Alberto Mugnai's research group at the Institute of Atmospheric Physics in Rome. The numerical modeling was conducted by Professor Oreg Tripoli and Ms. Giulia Panegrossi at the University of Wisconsin in Madison, using Professor Tripoli's nonhydrostatic modeling system (NMS). This is a scalable, fully nested mesoscale model capable of resolving nonhydrostatic circulations from regional scale down to cloud scale and below.
In this work, we extend the superstructure model proposed by Wamble et al. (2022) that considers feed input locations, recycling strategies, split fractions, stage numbers, and membrane area. We include the total number of stages as a decision variable, which might be particularly useful when there is cost as- sociated with adding additional stages. We propose a Generalized Disjunctive Programming (GDP) superstructure model that integrates all the design variables of the system. We also investigate the scalability of the model by varying the number of stages and the number of finite elements per stage to determine the impact on recovery and solution time.
The development of a parallel blade-element rotor model and its implementation into an adaptive Cartesian method is described. The unsteady version of the rotor model applies a body force to all cells contained in the swept space-time volume at each timestep and special care is taken to maintain axisymmetry on the Cartesian grid. Mesh convergence of rotor thrust and torque is obtained with around 10000 cells in the disk for the steady model. Parallelization is accomplished using OpenMP and the rotor force computation is distributed across all available nodes. Simulations of an isolated XV-15 rotor in hover show good correlation with experimental data and predictions of multi-rotor thrust variation closely match previous high fidelity simulations. The final paper will also include results from the unsteady rotor model and parallel scaling tests.
Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.