Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

A scalable framework for efficient coupling of thermal and microstructural simulations in additive manufacturing

Predicting microstructure evolution in metal additive manufacturing (AM) is important for process optimization, but spatiotemporal scale disparities between thermal transport and microstructure evolution create significant challenges for efficient data transfer between simulation codes. To address this, we present Stork, a scalable framework for coupling thermal and microstructural simulations. Stork uses a sparse data representation to identify and store active solidification sub-volumes, enabling highly parallel quad-linear interpolation from coarse thermal grids to fine microstructure grids without large intermediate storage. We demonstrate the framework by coupling the semi-analytic heat transfer code 3DThesis with the time-parallel cellular automata code Toucan. This approach achieves over two orders of magnitude reduction in data generation time and file size compared to prior workflows. Numerical studies show that quad-linear interpolation preserves grain morphology and crystallographic texture in laser powder bed fusion (LPBF) simulations for coarsening ratios up to 16. Overall, Stork provides a scalable pathway for high-throughput, component-scale AM simulations on modern high-performance computing systems.

36 MATERIALS SCIENCE↗

Coupling Metabolic Source Isotopic Pair Labeling and Genome Wide Association for Metabolite and Gene Annotation in Plants (Final Technical Report)

In this project, we applied our labeling pipeline to Arabidopsis and sorghum by feeding tissues with isotopically labeled versions of commercially available amino acids to identify all metabolite features that incorporate the label. In sorghum, we fed five accessions, sampled across the diversity of sorghum, to identify the precursor-of-origin for metabolites that vary between accessions as well as those that may be missing from a single reference genotype. This provided us with precursor-of-origin annotation for thousands of unknown metabolites. We then used GWA to map genes responsible for the synthesis of precursor-of-origin classified metabolites. For sorghum leaf and root ducible metabolites, we performed untargeted metabolomics on leaf and root tissues from 300 diverse genotyped sorghum inbred lines. The amino acid precursor-of-origin metabolite library were then used to identify the corresponding metabolites in the GWA data sets and to identify novel gene-metabolite associations. Finally, we utilized existing and newly generated sequenced EMS mutants of sorghum to validate the predicted gene-metabolite relationships that our labelling analysis identified. In parallel, we conducted similar feeding experiments in Arabidopsis to categorize metabolites based on precursor-of-origin, identify those that vary across our existing Arabidopsis metabolite GWA dataset, and identify genes required for the synthesis of each metabolite. To provide an independent test of gene annotation and pathway involvement, we tested the GWA gene-metabolite associations in Arabidopsis by analyzing the metabolic phenotypes of gene knockouts. Genes of particular interest from both sorghum and Arabidopsis were studied in detail by directly measuring the activity of the corresponding enzymes following heterologous expression. In summary, this work classified as-yet-unknown amino acid-derived metabolites and identified genes involved in their production generated through “omics” technologies. This information was used to validate gene function and identify new metabolism in Arabidopsis and sorghum.

09 BIOMASS FUELS↗

Studies of Quark Transport and Hadronization in Nuclei

In this project, we conducted the first measurement of di‑hadron azimuthal correlations in deep inelastic scattering (DIS) off nuclei using the CLAS detector at Jefferson Lab. Using 5 GeV electron‑beam data collected on deuterium, carbon, iron, and lead targets, we extracted di‑pion correlation functions over a broad kinematic range. The results show a monotonic broadening of the correlation peak with increasing nuclear mass, along with pronounced dependencies on the pions’ kinematics. Separately, we implemented an algorithm based on the Kalman filter that achieved the first complete alignment of the CLAS12 central tracking system. In parallel, we developed simulations, algorithms, and performance studies that informed the conceptual designs of the forward hadronic calorimeter Insert and the Zero Degree Calorimeter, both of which are now included in the ePIC detector baseline for the forthcoming Electron Ion Collider.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Aerodynamic stability and drag characteristics of a parallel burn/SRM ascent configuration (M equals 0.6 to 4.96)

Experimental aerodynamic investigations were conducted in the NASA/MSFC 14-inch trisonic wind tunnel during April 1972 on a 0.004-scale model of a solid rocket motor version of the space shuttle ascent configuration. The configuration consisted of a parallel burn solid rocket motor booster on an external HO centerline tank orbiter. Six component aerodynamic force and moment data were recorded over an angle of attack range from -10 deg to +10 deg at zero degrees sideslip and over a sideslip range from -10 deg to +10 deg at zero degrees angle of attack. Mach numbers ranged from 0.6 to 4.96. The purpose of the test was to determine the performance and stability characteristics of the complete ascent configuration and buildup, and to determine the effects of variations in HO tank and SRM nose shaping, orbiter incidence and position, and position of the solid rocket motors.

Sims, F.↗

Optimizing Distributed Training on Frontier for Large Language Models

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Loss scaling studies have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable computational resources. For example, training a one trillion parameter GPT-style model on 20 trillion tokens requires a staggering 120 million exaflops. This research explores efficient distributed training strategies to extract this computation from Frontier, the world's first exascale supercomputer. We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning. For 22 Billion, 175 Billion, and 1 Trillion parameters, we achieved GPU throughputs of 38.38%, 36.14%, and 31.96%, respectively. For the training of the 175 Billion parameter model and the 1 Trillion parameter model, we achieved 100% weak scaling efficiency on 1024 and 3072 Mi250X GPUs, respectively. We also achieved strong scaling efficiencies of 89% and 87% for these two models. We trained these models only tens of iterations instead of training till completion.

Yin, Junqi↗

Application of an onboard processor to the OAO C spacecraft

The design of a stored program computer for spacecraft use and its application on the fourth Orbiting Astronomical Observatory (OAO) is reported. The computer is a medium scale, parallel machine with a memory capacity of 16384 words of 18 bits each. It possesses a comprehensive instruction repertoire and operates on 45 W of power (including the dc-to-dc converter). The machine operates at a 500-kHz rate and executes an add instruction in 10 microseconds. Its primary functions on OAO C will be auxiliary command storage, spacecraft monitoring and malfunction reporting, data compression and status summary, and possible performance of emergency corrective action for certain anomalous situations.

Stewart, W. N.↗

Analysis of gamma prime shape changes in a single crystal Ni-base superalloy

The microstructural evolution of a commercial single crystal superalloy, NASAIR 100, is analyzed using the existing high-temperature lattice mismatch data and high-temperature moduli obtained from tests on single crystals of gamma and gamma prime. A multiparticle analysis of the microstructural evolution is performed using a novel microstructural lattice simulation technique, MCFET. Under a uniaxial stress, a regular array of gamma prime particles in the simulated microstructure is predicted to coalesce and form a plate morphology, with the broad faces of the plates and stress axis perpendicular in tension but parallel in compression. These results are consistent with changes in gamma prime shape observed in NASAIR 100 following creep testing at 1000 C.

Gayda, J.↗

SSME Streamtube Evaluation Program (SSTEP)

An analytical model of the Space Shuttle Main Engine (SSME) called SSME Streamtube Evaluation Program (SSTEP) has been developed based upon the assumption that the propellant flows through the main combustion chamber can be represented by a bundle of parallel streamtubes. The motivation for the development of SSTEP lies in the desire to gain a basic understanding of the engine performance effects of several common SSME hardware modifications. Specifically, this model has been used to evaluate the changes in performance due to boundary layer coolant hole enlargement, LOX post plugging, acoustic cavity elimination, baffle removal, and main combustion chamber coolant leakage. The results show a good general agreement with the available test data suggesting at least a qualitative agreement between SSTEP modeling and actual engine performance. Through the use of several adjustment factors, which represent relaxations of the SSTEP formulation assumptions, it is shown that the test data can be very closely matched and that SSTEP can be used as a performance prediction tool.

Greene, William D.↗

Implementation of ADI: Schemes on MIMD parallel computers

In order to simulate the effects of the impingement of hot exhaust jets of High Performance Aircraft on landing surfaces a multi-disciplinary computation coupling flow dynamics to heat conduction in the runway needs to be carried out. Such simulations, which are essentially unsteady, require very large computational power in order to be completed within a reasonable time frame of the order of an hour. Such power can be furnished by the latest generation of massively parallel computers. These remove the bottleneck of ever more congested data paths to one or a few highly specialized central processing units (CPU's) by having many off-the-shelf CPU's work independently on their own data, and exchange information only when needed. During the past year the first phase of this project was completed, in which the optimal strategy for mapping an ADI-algorithm for the three dimensional unsteady heat equation to a MIMD parallel computer was identified. This was done by implementing and comparing three different domain decomposition techniques that define the tasks for the CPU's in the parallel machine. These implementations were done for a Cartesian grid and Dirichlet boundary conditions. The most promising technique was then used to implement the heat equation solver on a general curvilinear grid with a suite of nontrivial boundary conditions. Finally, this technique was also used to implement the Scalar Penta-diagonal (SP) benchmark, which was taken from the NAS Parallel Benchmarks report. All implementations were done in the programming language C on the Intel iPSC/860 computer.

Vanderwijngaart, Rob F.↗

A Versatile Planetary Radio Science Microreceiver

We have developed a low-power. programmable radio "microreceiver" that combines the functionality of two science instruments: a Relative Ionospheric Opacity Meter (riometer) and a swept-frequency, VTF/HF radio spectrometer. The radio receiver, calibration noise source, data acquisition and processing, and command and control functions are all contained on a single circuit board. This design is suitable for miniaturizing as a complete flight instrument. Several of the subsystems were implemented in a field-programmable gate array (FPGA), including the receiver detector, the control logic, and the data acquisition and processing blocks. Considerable efforts were made to reduce the power consumption of the instrument, and eliminate or minimize RF noise and spurious emissions generated by the receiver's digital circuitry. A prototype instrument was deployed at McMurdo Station, Antarctica, and operated in parallel with a traditional riometer instrument for approximately three weeks. The attached paper (accepted for publication by Radio Science) describes in detail the microreceiver theory of operation, performance specifications and test results.

Fry, Craig D.↗

Liquid Oxygen/Liquid Methane Test Results of the RS-18 Lunar Ascent Engine at Simulated Altitude Conditions at NASA White Sands Test Facility

Tests were conducted with the RS-18 rocket engine using liquid oxygen (LO2) and liquid methane (LCH4) propellants under simulated altitude conditions at NASA Johnson Space Center White Sands Test Facility (WSTF). This project is part of NASA's Propulsion and Cryogenics Advanced Development (PCAD) project. "Green" propellants, such as LO2/LCH4, offer savings in both performance and safety over equivalently sized hypergolic propulsion systems in spacecraft applications such as ascent engines or service module engines. Altitude simulation was achieved using the WSTF Large Altitude Simulation System, which provided altitude conditions equivalent up to ~122,000 ft (~37 km). For specific impulse calculations, engine thrust and propellant mass flow rates were measured. LO2 flow ranged from 5.9 - 9.5 lbm/sec (2.7 - 4.3 kg/sec), and LCH4 flow varied from 3.0 - 4.4 lbm/sec (1.4 - 2.0 kg/sec) during the RS-18 hot-fire test series. Propellant flow rate was measured using a coriolis mass-flow meter and compared with a serial turbine-style flow meter. Results showed a significant performance measurement difference during ignition startup due to two-phase flow effects. Subsequent cold-flow testing demonstrated that the propellant manifolds must be adequately flushed in order for the coriolis flow meters to give accurate data. The coriolis flow meters were later shown to provide accurate steady-state data, but the turbine flow meter data should be used in transient phases of operation. Thrust was measured using three load cells in parallel, which also provides the capability to calculate thrust vector alignment. Ignition was demonstrated using a gaseous oxygen/methane spark torch igniter. Test objectives for the RS-18 project are 1) conduct a shakedown of the test stand for LO2/methane lunar ascent engines, 2) obtain vacuum ignition data for the torch and pyrotechnic igniters, and 3) obtain nozzle kinetics data to anchor two-dimensional kinetics codes. All of these objectives were met with the RS-18 data and additional testing data from subsequent LO2/methane test programs in 2009 which included the first simulated-altitude pyrotechnic ignition demonstration of LO2/methane.

Melcher, John C., IV↗

Mars Cube One (MarCO) Shifting the Paradigm in Relay Deep Space Operations

A very significant challenge in the planetary mission design and operations is communications with ground control teams during critical events and highly risky maneuvers. These include entry, descent, and landing (EDL) and orbit insertion, which should not be carried out in the blind. Although vast planetary distances and long round-trip light times disallow real-time intervention from controllers, acquiring the relevant event performance parameters in near real-time can be imperative for determining the corrective actions needed immediately following or, in the case of significant anomalies, aid in the diagnostic analysis. During several previous Mars missions landing events, and the Huygens probe landing on Titan, the communications strategy relied on proximity links to planetary orbiters, which then relayed the data to the Deep Space Network (DSN). In addition, attempts were successfully made in parallel to receive the signal carrier directly at Earth often using large radio telescopes when the wavelength was outside the DSN’s reception bands. This Direct-to-Earth (DTE) back-up method was only possible due to special techniques utilizing the DSN’s open-loop Radio Science Receivers. In every case, it was very challenging since the link budget of a landing vehicles were designed for proximity orbiter relays and not for distances across the solar system. A new method is introduced since not all future missions can rely on the presence of pre-existing orbiters at their planetary targets to relay their critical data ad, furthermore, most missions would not likely have the resources to implement a reliable DTE link at acceptable data rates, bypassing a need for a relay asset. With the advent CubeSat form-factor spacecraft, one or more, for added reliability, CubeSats can be launched with the primary mission, travel to the target, and be positioned to view the critical event, such as EDL, and carry out real-time relay of the data to the DSN at higher rates. CubeSats have flown in the Earth environment but never flown or been operated in deep space or planetary environment so careful design as well as flight experience are needed. The relay function requires the development of radio and antenna systems to meet challenging specifications. After initial technical demonstration of the concept and operational experience, the cost can decrease as systems become more standardized with increased reliability. This paper describes the invention of the “carry your own relay” concept and the formulation of the mission likely to be the first planetary CubeSat mission called Mars Cube One (MarCO). It also describes the operational concept of relay small spacecraft and their role reducing mission risk as well as overall mission cost.

Asmar, Sami↗

Cassini Power Subsystem

The electrical output of Cassini’s power system has been decaying consistently as predicted during its 20-year mission between October 1997 and September 2017. The power telemetry data is presented for the entire Cassini mission, including launch, cruise, and Saturn tour, up to the most recent available data. The spacecraft has been powered by three independent Radioisotope Thermoelectric Generators (RTGs) connected in parallel, which were able to generate 882 W at the beginning of the mission shortly after launch. The decrease in power energy output has mainly been driven by the heat reduction of the hot side of the RTGs due to the natural radioactive decay of its heat source plutonium (mostly plutonium-238 [238Pu]), degradation of thermoelectric material performance, and interface degradation.

Carr, Gregory A.↗

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton↗

Improved CDMA Performance Using Parallel Interference Cancellation

This paper considers a general parallel interference cancellation scheme that significantly reduces the degradation effect of user interference but with a lesser implementation complexity than the maximum-likelihood technique. The scheme operates on the fact that parallel processing simultaneously removes from each user the total interference produced by the remaining most reliably received users accessing the channel. The parallel processing can be done in multiple stages. The proposed scheme uses tentative decision devices with different optimum thresholds at the multiple stages to produce the most reliably received data for generation and cancellation of user interference.

CDMA↗

Conceptual Designs for Irradiation Creep Testing of SiC in HFIR

Understanding irradiation creep of nuclear fuel cladding is important to properly size the initial fuel-cladding gap and understand when pellet-cladding contact is expected to occur due to a combination of fuel swelling and cladding creep-down. Irradiation creep also plays a role in relaxing stresses that develop in-pile. Silicon carbide fiber–reinforced silicon carbide matrix (SiC/SiC) composites are the leading long-term accident-tolerant fuel cladding concept for light-water reactors (LWRs). Although some limited data are available regarding irradiation creep of the individual constituents (fibers, matrix), data regarding irradiation creep of SiC/SiC composites are currently insufficient. Additional data regarding irradiation creep compliance and the rupture lifetime (combination of creep and slow crack growth) are needed to understand material limitations. This work describes the design and development of two irradiation vehicles that are being pursued for testing SiC/SiC concepts in the High Flux Isotope Reactor (HFIR). The first is a passive experiment, referred to as the PRECISE experiment, that leverages the constant coolant pressure of HFIR to compress a metallic bellows and provide a well-characterized load to drive creep in a SiC/SiC dog bone specimen. The total creep strain would be quantified post-irradiation by measuring dimensional changes of the specimen length as well as local dimensional changes within the gauge region. Non-stressed specimens would also be irradiated under the same conditions to provide an indication of dimensional changes due to radiation-induced swelling in the absence of creep. A second, more complex experiment, referred to as the INSITE experiment, is being designed in parallel that would use pneumatics to pressurize a metal bellows and linear variable differential transformers (LVDTs) to measure the specimen displacement in situ during irradiation. Such an experiment would provide significantly more data regarding the evolution of the creep compliance as a function of dose and applied stress within a single experiment but would require significantly more development time and cost to execute. The primary concern with the INSITE experiment is the accuracy, reliability, and expected lifetime of the LVDTs during irradiation at elevated temperatures. Efforts are being made to adjust the experiment design and operating procedure to limit LVDT temperatures and mitigate or otherwise compensate for uncertainties due to factors such as temperature fluctuations, creep in the surrounding structural materials, and drift of the LVDTs. This work describes the experiment designs, thermal and structural analysis that were performed to ensure that the desired temperature and stress conditions can be achieved, some initial sensitivity analyses to predict the evolution of the radiation-induced specimen displacements, and potential sources of uncertainty in the measurements. Out-of-pile testing is being performed in parallel to confirm that the test trains achieve the expected stress states in the specimens and do not result in prohibitive stress concentrators (e.g., in the grip regions) that might risk pre-mature failure. The PRECISE experiments are proceeding toward fabrication and assembly with HFIR insertion planned during fiscal year 2026. The INSITE experiment is progressing toward out-of-pile demonstrations, which will provide more conclusive evidence regarding the feasibility of executing these tests in HFIR or whether alternative displacement monitoring techniques may need to be considered.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗

A Survey on Error-Bounded Lossy Compression for Scientific Datasets

Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.

Error-Bounded Lossy Compression↗