Search NASASearch

SEARCH · Search NASA

Results for “multiple computing resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Portability and Cross-Platform Performance of an MPI-Based Parallel Polygon Renderer

Visualizing the results of computations performed on large-scale parallel computers is a challenging problem, due to the size of the datasets involved. One approach is to perform the visualization and graphics operations in place, exploiting the available parallelism to obtain the necessary rendering performance. Over the past several years, we have been developing algorithms and software to support visualization applications on NASA's parallel supercomputers. Our results have been incorporated into a parallel polygon rendering system called PGL. PGL was initially developed on tightly-coupled distributed-memory message-passing systems, including Intel's iPSC/860 and Paragon, and IBM's SP2. Over the past year, we have ported it to a variety of additional platforms, including the HP Exemplar, SGI Origin2OOO, Cray T3E, and clusters of Sun workstations. In implementing PGL, we have had two primary goals: cross-platform portability and high performance. Portability is important because (1) our manpower resources are limited, making it difficult to develop and maintain multiple versions of the code, and (2) NASA's complement of parallel computing platforms is diverse and subject to frequent change. Performance is important in delivering adequate rendering rates for complex scenes and ensuring that parallel computing resources are used effectively. Unfortunately, these two goals are often at odds. In this paper we report on our experiences with portability and performance of the PGL polygon renderer across a range of parallel computing platforms.

Crockett, Thomas W.

Processing Spacecraft Data Without Confusion

Producing multiple versions of the same data product for the same time frame with the same remotely sensed inputs can be a recipe for disaster. Yet, amidst the commotion of satellite launch and early operations (LEO), such data processing is needed. After LEO, the situation gets worse. Processing newly arriving data ("forward processing") is augmented with reprocessing and algorithm development, comparison, evaluation, and testing -- often happening all at the same time. The problem can be analyzed in three main parts -- maintaining multiple versions of algorithms and data so that end-product users are not overwhelmed,allocating computer resources efficiently, and simplifying production operations so that va st amounts of data can be processed with minimal staff and fewer errors. OMIDAPS provides a framework for execution of algorithms that transform lower level data acquired by OMI on NASA's Aura satellite into higher level science data products. In contrast to traditional science data processing systems, we address all parts of the problem with an innovative approach allowing multiple data processing to run within a single physical system. The data products, imports, exports, and execution planning are all segregated into distinct "ArchiveSets." This paper describes reasons for multiple concurrent productions on a typical satellite data processing project using OMI as an example. It describes the virtual data processing system concept and its advantages over separate physical processing strings. It explores the specific implementation of the virtual systems within OMIDAPS and discusses some of the implications of our approach and describes how virtual processing is used to accomplish the overall mission of OMI data processing.

Tilmes, Curt

Development of typical solar years and typical wind years for efficient assessment of renewable energy systems across the U.S.

Weather data plays a critical role in renewable energy analysis. Compared to using multiple Actual Meteorological Years, simulations using a single typical year require significantly fewer computational resources. Previous efforts to create typical weather datasets for renewable energy analysis either lack justified or optimized strategies for selecting and weighting different weather parameters or are limited to a few specific locations. Here, in this study, we developed a dataset comprising Typical Solar Years (TSYs) and Typical Wind Years (TWYs) for over 2000 locations across the U.S., based on data from NASA's POWER project. The strategies for creating TSYs and TWYs were optimized based on the simulated outputs of various PV systems and wind turbines in 16 representative cities. This dataset provides an efficient means for the rapid evaluation and optimization of renewable energy systems throughout the entire U.S. Additionally, the optimal strategies identified in this study can be directly applied to create near-optimal TSYs and TWYs for most locations worldwide. However, readers can also employ the optimization approach presented in this work to develop optimal strategies tailored for particular regions.

NASA POWER

An Efficient Storage-Driven Machine Learning Model for Performance in the Era of Multimodal Scientific Data

Scientific workflows are increasingly relying on machine learning (ML), simulation, and hybrid techniques to predict, understand, and optimize the behavior of complex experiments. High-performance computing has greatly improved researchers’ ability to acquire diverse data modalities in these workflows. Recent studies suggest that the performance of machine learning models can be improved by integrating data from various sources. Unfortunately, these workloads pose unprecedent pressure on the network storage to meet the demands associated with accessing these multimodal data. To mitigate the impact of intensive IO, we propose a solution that utilizes a multi-tier High-Performance Computing (HPC) distributed storage and data processing framework, placing computation where the data resides for better performance. By adopting this project, the scientific community will gain new opportunities to explore multimodal storage-driven possibilities, integrating multiple scientific data sources with advanced streaming frameworks. Additionally, our framework effectively utilizes computing resources and bridges the gaps identified by HPC experts. Our proposed approach tackles scalability and persistence challenges by leveraging native persistency, which has posed difficulties in traditional approaches. Furthermore, we seek to enhance fault-tolerance and load-balance of computations by leveraging real-time streaming in diverse scientific computing environments, thereby propelling advanced scientific computing research into the next generation.

97 MATHEMATICS AND COMPUTING

Economical Implementation of a Filter Engine in an FPGA

A logic design has been conceived for a field-programmable gate array (FPGA) that would implement a complex system of multiple digital state-space filters. The main innovative aspect of this design lies in providing for reuse of parts of the FPGA hardware to perform different parts of the filter computations at different times, in such a manner as to enable the timely performance of all required computations in the face of limitations on available FPGA hardware resources. The implementation of the digital state-space filter involves matrix vector multiplications, which, in the absence of the present innovation, would ordinarily necessitate some multiplexing of vector elements and/or routing of data flows along multiple paths. The design concept calls for implementing vector registers as shift registers to simplify operand access to multipliers and accumulators, obviating both multiplexing and routing of data along multiple paths. Each vector register would be reused for different parts of a calculation. Outputs would always be drawn from the same register, and inputs would always be loaded into the same register. A simple state machine would control each filter. The output of a given filter would be passed to the next filter, accompanied by a "valid" signal, which would start the state machine of the next filter. Multiple filter modules would share a multiplication/accumulation arithmetic unit. The filter computations would be timed by use of a clock having a frequency high enough, relative to the input and output data rate, to provide enough cycles for matrix and vector arithmetic operations. This design concept could prove beneficial in numerous applications in which digital filters are used and/or vectors are multiplied by coefficient matrices. Examples of such applications include general signal processing, filtering of signals in control systems, processing of geophysical measurements, and medical imaging. For these and other applications, it could be advantageous to combine compact FPGA digital filter implementations with other application-specific logic implementations on single integrated-circuit chips. An FPGA could readily be tailored to implement a variety of filters because the filter coefficients would be loaded into memory at startup.

Kowalski, James E.

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]

Collectives for Multiple Resource Job Scheduling Across Heterogeneous Servers

Efficient management of large-scale, distributed data storage and processing systems is a major challenge for many computational applications. Many of these systems are characterized by multi-resource tasks processed across a heterogeneous network. Conventional approaches, such as load balancing, work well for centralized, single resource problems, but breakdown in the more general case. In addition, most approaches are often based on heuristics which do not directly attempt to optimize the world utility. In this paper, we propose an agent based control system using the theory of collectives. We configure the servers of our network with agents who make local job scheduling decisions. These decisions are based on local goals which are constructed to be aligned with the objective of optimizing the overall efficiency of the system. We demonstrate that multi-agent systems in which all the agents attempt to optimize the same global utility function (team game) only marginally outperform conventional load balancing. On the other hand, agents configured using collectives outperform both team games and load balancing (by up to four times for the latter), despite their distributed nature and their limited access to information.

Tumer, K.

Evaluating local indirect addressing in SIMD proc essors

In the design of parallel computers, there exists a tradeoff between the number and power of individual processors. The single instruction stream, multiple data stream (SIMD) model of parallel computers lies at one extreme of the resulting spectrum. The available hardware resources are devoted to creating the largest possible number of processors, and consequently each individual processor must use the fewest possible resources. Disagreement exists as to whether SIMD processors should be able to generate addresses individually into their local data memory, or all processors should access the same address. The tradeoff is examined between the increased capability and the reduced number of processors that occurs in this single instruction stream, multiple, locally addressed, data (SIMLAD) model. Factors are assembled that affect this design choice, and the SIMLAD model is compared with the bare SIMD and the MIMD models.

Middleton, David

NASA Tech Briefs, February 2004

Topics include: Simulation Testing of Embedded Flight Software; Improved Indentation Test for Measuring Nonlinear Elasticity; Ultraviolet-Absorption Spectroscopic Biofilm Monitor; Electronic Tongue for Quantitation of Contaminants in Water; Radar for Measuring Soil Moisture Under Vegetation; Modular Wireless Data-Acquisition and Control System; Microwave System for Detecting Ice on Aircraft; Routing Algorithm Exploits Spatial Relations; Two-Finger EKG Method of Detecting Evasive Responses; Updated System-Availability and Resource-Allocation Program; Routines for Computing Pressure Drops in Venturis; Software for Fault-Tolerant Matrix Multiplication; Reproducible Growth of High-Quality Cubic-SiC Layers; Nonlinear Thermoelastic Model for SMAs and SMA Hybrid Composites; Liquid-Crystal Thermosets, a New Generation of High-Performance Liquid-Crystal Polymers; Formulations for Stronger Solid Oxide Fuel-Cell Electrolytes; Simulation of Hazards and Poses for a Rocker-Bogie Rover; Autonomous Formation Flight; Expandable Purge Chambers Would Protect Cryogenic Fittings; Wavy-Planform Helicopter Blades Make Less Noise; Miniature Robotic Spacecraft for Inspecting Other Spacecraft; Miniature Ring-Shaped Peristaltic Pump; Compact Plasma Accelerator; Improved Electrohydraulic Linear Actuators; A Software Architecture for Semiautonomous Robot Control; Fabrication of Channels for Nanobiotechnological Devices; Improved Thin, Flexible Heat Pipes; Miniature Radioisotope Thermoelectric Power Cubes; Permanent Sequestration of Emitted Gases in the Form of Clathrate Hydrates; Electrochemical, H2O2-Boosted Catalytic Oxidation System; Electrokinetic In Situ Treatment of Metal-Contaminated Soil; Pumping Liquid Oxygen by Use of Pulsed Magnetic Fields; Magnetocaloric Pumping of Liquid Oxygen; Tailoring Ion-Thruster Grid Apertures for Greater Efficiency; and Lidar for Guidance of a Spacecraft or Exploratory Robot.

Source record

An Integrated Economics Model for ISRU in Support of a Mars Colony - Initial Results Report

This database, the Mars Colony Architecture Model (MCAM), is then linked to a variety of “downsteam” analytic models. In particular, we integrated an Extraction Process (i.e., “Mining”) Model, an Infrastructure and Integrated Logistics Support (ILS) Model, and an Economics Integration Model. The Extraction Process Model focuses on the technologies associated with in situ resource extraction, processing, storage and handling, and delivery. For each mined resource, which may involve multiple cooperating In Situ Resource Utilization (ISRU) systems in a given architecture, the Extraction Process Model computes the production rate as a function of the systems’ technical parameters and the local Mars environment. As with our earlier work, this model focuses on the extraction and processing of Mars water/ice.

Shishko, Robert

Extreme-scale workflows: A perspective from the JLESC international community

The Joint Laboratory for Extreme-Scale Computing (JLESC) focuses on software challenges in high-performance computing systems to meet the needs of today’s science campaigns, which often require large resources, consist of multiple tasks, and generate vast amounts of data. In this context, extreme-scale workflows have been the key factor in enabling scientific discoveries by helping scientists automate the dependencies and data exchanges between workflow tasks, instead of managing those manually. Here, in this paper, we present representative extreme-scale workflows and feature workflow systems developed by JLESC participating institutions. We present lessons learned while developing these tools, alongside with the open challenges and future research directions in the field of extreme-scale workflows.

97 MATHEMATICS AND COMPUTING

An Automatic Medium to High Fidelity Low-Thrust Global Trajectory Toolchain; EMTG-GMAT

Solving the global optimization, low-thrust, multiple-flyby interplanetary trajectory problem with high-fidelity dynamical models requires an unreasonable amount of computational resources. A better approach, and one that is demonstrated in this paper, is a multi-step process whereby the solution of the aforementioned problem is solved at a lower-fidelity and this solution is used as an initial guess for a higher-fidelity solver. The framework presented in this work uses two tools developed by NASA Goddard Space Flight Center: the Evolutionary Mission Trajectory Generator (EMTG) and the General Mission Analysis Tool (GMAT). EMTG is a medium to medium-high fidelity low-thrust interplanetary global optimization solver, which now has the capability to automatically generate GMAT script files for seeding a high-fidelity solution using GMAT's local optimization capabilities. A discussion of the dynamical models as well as thruster and power modeling for both EMTG and GMAT are given in this paper. Current capabilities are demonstrated with examples that highlight the toolchains ability to efficiently solve the difficult low-thrust global optimization problem with little human intervention.

low thrust

(ODIN): An Open Source, Low-Latency Data Integration & Visualization Framework for the NASA System Wide Safety Project's Disaster Response Safety Demonstration Series

The Open Data Integration Framework (ODIN) is an open source, low latency data integration and visualization framework (https://github.com/NASARace/race-odin) developed under NASA’s System WideSafety Program to demonstrate new safety capabilities designed to improve US airspace operations. Safety demonstrations are a set of increasingly complex (from public safety perspective) disaster response scenarios under which air systems must operate with increased capacity and include: 1) Wildland fire response, 2) Hurricane relief and recovery, 4) Emergency medical delivery via UAS and 4) Urban disaster relief. To accommodate disaster response, ODIN is field deployable and can scale on one or more multi-core, commodity laptops operating with full to limited or intermittent internet connectivity, conditions likely encountered during operations. ODIN runs as webserver with local, persistent data storage to serve either public or a secured, ad hoc network (e.g., an incident command post). The current released ODIN, ODIN-Fire is tailored for wildland fire management incorporating information on satellite overpasses with links to the near real-time data and imagery from the respective agencies. Included are winds data, an important variable for emergency responders and airspace operations, and high-resolution wind forecasts generated by super-computing resources and ingested into ODIN. As an open-source project, ODIN has attracted interest from multiple entities. We will show how 1) a commercial field instrument and data provider uses ODIN to help users visualize, publish and integrate their in-situ sensor network data and 2) ODIN’s capabilities to ingest, integrate and display near-real time satellite data with air traffic and a USFS winds forecast model used in fire response and post-fire assessment. Within NASA ODIN demonstrated novel, near terminal airspace safety capabilities for a project close-out event and previously it monitored the national airspace in real-time to meet an agency milestone. ODIN is presently under development for the anticipated hurricane relief and response demonstration notionally scheduled for the 2025-27 time frame and is available from NASA's github at the above link.

Aeronautics

A Generic Multibody Parachute Simulation Model

Flight simulation of dynamic atmospheric vehicles with parachute systems is a complex task that is not easily modeled in many simulation frameworks. In the past, the performance of vehicles with parachutes was analyzed by simulations dedicated to parachute operations and were generally not used for any other portion of the vehicle flight trajectory. This approach required multiple simulation resources to completely analyze the performance of the vehicle. Recently, improved software engineering practices and increased computational power have allowed a single simulation to model the entire flight profile of a vehicle employing a parachute.

Neuhaus, Jason Richard

Nuclear Space System Analysis and Modelling (NSSAM): A Software Tool to Efficiently Analyze the Design Space of Space Reactor Systems

Space reactors have the potential to play a key role in future NASA exploration activities due to their capability to enable sustainable power and advanced propulsion systems. To enable assessment of the space reactor design space, the nuclear space system analysis and modelling (NSSAM) software was developed by Analytical Mechanics Associates. NSSAM leverages a scalable and extensible software architecture which automates reactor analysis to perform coupled engine-reactor and reactor physics-thermal hydraulics calculations. This allows space reactor systems to be evaluated by a wider number of users with a consistent analysis approach to compare designs. NSSAM has been developed with multiple use cases to tailor the analysis to the level of detail desired by the user and computing resources. This summary overviews the NSSAM architecture and development approach, current capabilities (including design variants and use cases) and analysis approach for reactor and system component models.

nuclear thermal propulsion

Boosting Noise2Inverse via enhanced model selection for denoising computed tomography data

Synchrotron-based x-ray tomographic imaging enables the examination of the internal structure of materials at high spatial and temporal resolution. Experimental constraints can impose dose and time limits on the measurements, introducing a higher level of noise and artifacts in the reconstructed images. Deep learning has emerged as a powerful tool to remove noise from reconstructed images. Recently, the Noise2Inverse method was designed specifically for denoising reconstructed images without requiring paired noisy and clean images. This method creates multiple statistically independent reconstructions used to pair the data in which training involves transforming one reconstruction into the other, and vice versa. Originally designed to be used after a fixed number of epochs, we see in practice that this approach may not produce the optimal model and may unnecessarily waste computational resources. Therefore, we propose an alternative method of identifying the best model during training that aligns with the Noise2Inverse method. During validation, we compare the model output of the multiple reconstructions among each other. We hypothesize that the best model is the one that produces images with the highest similarity, implying a convergence in the predicted material properties and absorption values. To compare model outputs, we consider the absolute error, square error, structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), and cosine similarity. We evaluate our method on two simulated tomography datasets and two, real-world, low-contrast, high-energy x-ray tomography datasets. We show our approach is more effective at determining the best model, up to an increase of 12.50% and 12.53% in SSIM and PSNR, respectively, while only requiring a fifth of the training time compared to the original approach.

CT

Computational capacity in hydrodynamic real-time hybrid simulation applied to simulate the dynamic response of floating offshore wind turbines

Real-time hybrid simulation (RTHS) mitigates similitude distortions in model-scale tests of floating offshore wind turbines (FOWTs) by coupling physical experiments with numerical models in real time. The coupling requires faster-than-real-time numerical computations to satisfy temporal similitude with the physical experiment, presenting a bottleneck for using more complex numerical models in RTHS. This paper presents a hydrodynamic-RTHS (hydro-RTHS) framework for FOWTs that simulates the hydrodynamics physically and the aerodynamics numerically with sensor feedback from the physical testing. The framework adapts the three-loop hardware architecture to leverage greater computational resources and mitigate strict temporal requirements, enabling more computationally demanding numerical analyses in hydro-RTHS. The three-loop hardware architecture integrates multiple machines, each dedicated to either numerical analysis or RTHS controls, with a rate-transition algorithm to synchronize the tasks executed across the different machine processors. Virtual and physical tests verified and validated the hydro-RTHS framework, respectively. The ”virtual” tests, which approximates the physical domain numerically, verified the RTHS framework with respect to a numerical full-scale complete FOWT model simulated in the open-source software, OpenFAST. The virtual tests were able to maintain comparable control signals while enabling greater computational resources for the numerical calculations. Real-world physical tests demonstrated that the hydro-RTHS framework computes aerodynamic forces similar to the complete OpenFAST model, validating the hydro-RTHS framework using the three-loop hardware architecture. Findings show that the hydro-RTHS framework with the three-loop hardware architecture is computationally efficient, with reserve capacity to simulate more complex problems due to the customized software, hardware, and rate-transition algorithm.

17 WIND ENERGY

Enabling end-to-end secure federated learning in biomedical research on heterogeneous computing environments with APPFLx

Facilitating large-scale, cross-institutional collaboration in biomedical machine learning (ML) projects requires a trustworthy and resilient federated learning (FL) environment to ensure that sensitive information such as protected health information is kept confidential. Specifically designed for this purpose, this work introduces APPFLx - a low-code, easy-to-use FL framework that enables easy setup, configuration, and running of FL experiments. APPFLx removes administrative boundaries of research organizations and healthcare systems while providing secure end-to-end communication, privacy-preserving functionality, and identity management. Furthermore, it is completely agnostic to the underlying computational infrastructure of participating clients, allowing an instantaneous deployment of this framework into existing computing infrastructures. Experimentally, the utility of APPFLx is demonstrated in two case studies: (1) predicting participant age from electrocardiogram (ECG) waveforms, and (2) detecting COVID-19 disease from chest radiographs. Here, ML models were securely trained across heterogeneous computing resources, including a combination of on-premise high-performance computing and cloud computing facilities. By securely unlocking data from multiple sources for training without directly sharing it, these FL models enhance generalizability and performance compared to centralized training models while ensuring data remains protected. In conclusion, APPFLx demonstrated itself as an easy-to-use framework for accelerating biomedical studies across organizations and healthcare systems on large datasets while maintaining the protection of private medical data.

Biomedical Research