Search NASA⌕ Search

SEARCH · Search NASA

Results for “big data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

AmeriFlux FLUXNET-1F US-xPU NEON Pu'u Maka'ala Natural Area Reserve (PUUM)

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-xPU NEON Pu'u Maka'ala Natural Area Reserve (PUUM). This is the FLUXNET version of the carbon flux data for the site US-xPU NEON Pu'u Maka'ala Natural Area Reserve (PUUM) produced by applying the standard ONEFlux (1F) software. Site Description - NEON's PUUM field site is located in the Pu'u Maka'ala Natural Area Reserve (NAR) on the eastern side of Hawaii’s “Big Island,” managed by the Hawaii Division of Forestry and Wildlife (DOFAW). More than 18,000 acres in size, the NAR is home to a rainforest with many native species, some of them endangered. It was established to protect some of the Big Island’s best wet native forest and unique geologic features.

Network), NEON (National Ecological Observatory [N↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the cumulative effects of nearshore habitat restoration actions for multiple populations of juvenile salmon in Whidbey Basin, Washington: foundation and approach for synthesis and evaluation

Ecosystem restoration is a common tool for re-establishing ecosystem processes, structures, and functions to improve biodiversity and services in coastal and estuarine ecosystems. In the Salish Sea, salmon habitats have been fragmented, reduced in size, and diminished in quality, and the ecosystem processes that form and sustain these habitats have been degraded and disrupted as well. This loss is especially prevalent in estuaries, where up to 90% of former salmon habitat has been lost or compromised. Salmon species are integral to the identities and cultures of people in the Pacific Northwest, yet salmon abundances remain at historic lows, especially in urbanized areas. Recent investments in restoration are creating rearing habitat and repairing lost ecosystem function. However, restoration efforts in this region have largely proceeded at the site scale, with less attention to big-picture thinking regarding how restoration will effectively recover degraded or lost habitats for target species. As a result, no landscape-scale evaluation program exists, and the cumulative benefits of multiple interventions are unknown. We describe innovative methods for science synthesis related to the evaluation of cumulative effects of ecosystem restoration for Pacific salmon, using years of existing, but disparate data. Building from previous work on cumulative effects evaluation and incorporating a hierarchy of hypotheses approach, we propose using causal inference across numerous hypotheses in a framework to assess the cumulative benefits to Pacific salmon from multiple estuarine restoration projects. We present the framework as a method that can be used to address many complex questions and provide examples from the Salish Sea where the approach is being implemented. The framework draws on science synthesis from numerous fields and uses a hierarchy of hypotheses, causal analysis at multiple scales, and a new hierarchy of synthesis for assessing multiple lines of evidence documenting restoration effects on Pacific salmon. We propose causal inference to synthesize dissimilar data streams, in our case, to identify various manifestations of cumulative effects of restoration and benefits to salmon, and to further inform restoration and recovery planning. A unifying framework would allow for the detection of thresholds at which restoration provides measurable improvement and would greatly advance understanding of the effects of restoration on ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

Release of ENDF81SaB: ENDF/B-VIII.1-Based ACE Data Files for Thermal Scattering

On August 30, 2024, the National Nuclear Data Center (NNDC) released the ENDF/B-VIII.1 nuclear data library. The library was released in the standard Evaluated Nuclear Data File (ENDF) format. These files can be accessed on the NNDC's website (www.nndc.bnl.gov). The files provided in the thermal neutron scattering sublibrary were processed into A Compact ENDF (ACE)-formatted files, verified, and validated by the XCP-5 Nuclear Data Team, resulting in the ENDF81SaB application library. This report details the processing of these files and the quality assurance approach taken. This is not intended to be a full validation effort; rather, this library is intended to simply reproduce the released files for further validation testing by the community. The validation basis and details of the evaluations are documented in the forthcoming ``Big Paper''.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Enhanced matter power spectrum from axion kination after Big Bang nucleosynthesis

Despite stringent constraints from Big Bang Nucleosynthesis (BBN) and cosmic microwave background (CMB) observations, it is still possible for well-motivated particle physics models to substantially alter the cosmic expansion history between BBN and recombination. In this work we consider two different axion models that can realize a period of first matter domination, then kination, in this epoch. We perform fits to both primordial element abundances as well as CMB data and determine that up to a decade of late axion domination is allowed by these probes of the early universe. We establish the implications of late axion domination for the matter power spectrum on the scales 1/Mpc ≲ k ≲ 10 3 /Mpc. Our 'log' model predicts a relatively modest bump-like feature together with a small suppression relative to the standard ΛCDM predictions on either side of the enhancement. Our 'two-field' model predicts a larger, plateau-like feature that realizes enhancements to the matter power spectrum of up to two orders of magnitude. These features have interesting implications for structure formation at the forefront of current detection capabilities.

79 ASTRONOMY AND ASTROPHYSICS↗

Testing Protocol Development for the Fracture Toughness of Parts Built with Big Area Additive Manufacturing

The mechanical testing of additively manufactured parts has largely relied on the existing standards developed for traditional manufacturing. While this approach leverages the investment made in current standards development, it inaccurately assumes that the mechanical response of additive manufacturing (AM) parts is identical to that of parts manufactured through traditional processes. When considering thermoplastic, material extrusion AM, the differences in response can be attributed to an AM part’s inherent inhomogeneity caused by porosity, interlayer zones, and surface texture. Additionally, the interlayer bonding of parts printed with large-scale AM is difficult to adequately assess, as much testing is performed such that stress is distributed across many layer interfaces; therefore, the lack of AM-specific standards to assess interlayer bonding is a significant research gap. To quantify interlayer bonding via fracture toughness, double cantilever beam (DCB) testing has been used for some AM materials, and DCB has been generally used for a variety of materials including metal, wood, and laminates. Mode I DCB testing was performed on thermoplastic matrix composites printed with Big Area Additive Manufacturing (BAAM). Of particular interest was the notch shape and deflection speed during testing. The results examine the differences when using two notch types and three deflection speeds. The testing method introduced by the following paper differentiates itself from the ones described in the standards used by modernizing the methodology. This was conducted with the introduction of Digital Image Correlation (DIC) to gather displacement and load data simultaneously without human intervention.

Polymer Science↗

Limits on non-relativistic matter during Big-bang nucleosynthesis

Big-bang nucleosynthesis (BBN) probes the cosmic mass-energy density at temperatures ~10 MeV to ~100 keV. Here, we consider the effect of a cosmic matter-like species that is non-relativistic and pressureless during BBN. Such a component must decay; doing so during BBN can alter the baryon-to-photon ratio, η, and the effective number of neutrino species. We use light element abundances and the cosmic microwave background (CMB) constraints on η and N ν to place constraints on such a matter component. We find that electromagnetic decays heat the photons relative to neutrinos, and thus dilute the effective number of relativistic species to N eff < 3 for the case of three Standard Model neutrino species. Intriguingly, likelihood results based on Planck CMB data alone find N ν = 2.800 ± 0.294, and when combined with standard BBN and the observations of D and 4 He give N ν = 2.898 ± 0.141. While both results are consistent with the Standard Model, we find that a nonzero abundance of electromagnetically decaying matter gives a better fit to these results. Our best-fit results are for a matter species that decays entirely electromagnetically with a lifetime τ X = 0.89 sec and pre-decay density that is a fraction ξ = (ρ X /ρ rad |10 MeV = 0.0026 of the radiation energy density at 10 MeV; similarly good fits are found over a range where ξτ X 1/2 is constant. On the other hand, decaying matter often spoils the BBN+CMB concordance, and we present limits in the (τ X ,ξ) plane for both electromagnetic and invisible decays. For dark (invisible) decays, standard BBN (i.e. ξ = 0) supplies the best fit. We end with a brief discussion of the impact of future measurements including CMB-S4.

79 ASTRONOMY AND ASTROPHYSICS↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evaluating E3SM Global Storm‐Resolving Model Simulations of Deep Convection: Insights From DP‐SCREAM During TRACER

Global Storm-Resolving Models (GSRMs) are becoming increasingly vital for advancing climate modeling and improving the prediction of extreme weather events. Houston, a coastal region frequently affected by deep convective storms, offers an ideal setting to evaluate the ability of GSRMs to simulate deep convection. This study assesses the performance of the Doubly Periodic Simple Cloud-Resolving E3SM (Energy Exascale Earth System Model) Atmosphere Model (DP-SCREAM) using observations from the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign. DP-SCREAM effectively reproduces the diurnal cycles of clouds and precipitation, demonstrating much greater skill than the E3SM single column model. The DP-SCREAM is demonstrated to be applicable to coastal regions, partially due to the forcing data sets already capturing the influence of breezes. DP-SCREAM also replicates biases persistent in the global version of SCREAM: the underrepresentation of boundary layer shallow clouds, a lack of mid-level congestus clouds, and the popcorn convection, characterized by small and disorganized convective cells generating the strongest precipitation. To investigate these issues, two sensitivity experiments were conducted: increasing the mixing length and scaling up the buoyancy flux within the Simplified Higher Order Closure scheme. Increasing the mixing length improved mid-level congestus representation and reduced unrealistic early morning fog occurrence. Enhancing buoyancy flux only marginally improved the bias of underproduced big convective cells. In conclusion, an additional resolution sensitivity test at 0.5 km grid spacing demonstrated that a refined horizontal resolution alone is insufficient to resolve these biases.

54 ENVIRONMENTAL SCIENCES↗

Performance Evaluation of Intelligent Solar Control Software Through Hardware-in-the-Loop (CRADA Final Report)

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been developed by Latimer Controls, Inc. to estimate the headroom of large PV plants for grid operation and control; however, these technologies lack comprehensive validation under real-world application scenarios. Latimer Controls, Inc. received two voucher awards for research at a national laboratory from the Department of Energy American Made Solar Prize Round 6. The National Renewable Energy Laboratory (NREL) was selected to collaborate with Latimer staff to conduct a performance evaluation of Latimer PV control software. The NREL team will develop a hardware-in-the-loop (HIL) testbed to perform testing and validation of the Latimer PV control technology in a de-risked yet realistic testbed environment. Latimer and NREL worked together to analyze the test data, draw conclusions from the results, and disseminate the resulting scientific findings. In this CRADA work, we propose to test and validate the real-world application of the Latimer Control solution in an HIL environment. We evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. In particular, a data-driven potential high limit (PHL) estimation is developed for large solar plants to accurately estimate their headroom so that they have fast and short-time regulation and control capability to participate in grid services and respond to grid signals in real time (e.g., AGC). This PHL estimation algorithm is embedded in a hardware power plant controller (PPC) and tested with an IEEE-39 bus system model developed in RTDS. To account for the varying cloud conditions and diverse inverter dispatches, we developed a 135-MW PV plant with detailed modeling of 27 individual PV modules and inverters using RTDS. The real-world communications used in such big plants, such as ModBus TCP/IP for inverter level and DNP3 for plant level, were developed to emulate the real-world applications in big PV plants. The ML-based PHL estimation method is tested under nine separate weather scenarios against the ‘reference-control’ solution, hereafter referred to as the baseline solution. The baseline method reserves a subset of inverters (reference group) to operate at their PHL at all times and dispatches only the remaining inverters (control group) at curtailed levels to fulfill the flexibility need. Despite being successfully piloted by NREL in California in 2017 and Chile in 2020, there exist two gaps in the state of the art to fully unlock the flexibility of PV plants: a. There is a trade-off between the PHL estimation accuracy and the flexibility range. b. There lacks granularity in the PHL estimation to capture the variation across inverters. The Latimer solution seeks to address these gaps by applying machine learning methods to improve PHL estimation accuracy while accounting for variability at every inverter. Performance metrics were taken from the 2023 Georgia Power CARES utility-scale RFP. The results demonstrate that the ML-based approach outperforms the traditional baseline method in PHL estimation accuracy for 7 of 9 scenarios. The average PHL error across the nine scenarios was 7.40% for the ML-based method, 2.06% less than the 9.46% PHL error average across scenarios that was exhibited by the baseline method. Additionally, the PHL error was below 5% for at least 95% of the testing interval for 3 of 9 tested intervals with the ML approach, whereas it did not achieve this metric for any of the baseline tests. Overall, simulation results indicate the superior performance of an ML-based approach compared to the conventional baseline reference-control approach, showcasing its potential to support grid stability and operational efficiency. This laboratory HIL testing using real PPC, representative power system simulation models in real-time with detailed PV plant and inverter models, and real-world communication protocols gives us confidence that this machine learning based PHL estimation algorithm works well in the hardware PPC and therefore de-risks future field commissioning. The end goal of this project is to advance grid technology to address the grid operation challenges brought by solar plant’s variability and uncertainties in power generation.

14 SOLAR ENERGY↗

Improved constraints on hematite refractive index for estimating climatic effects of dust aerosols

Abstract Uncertainty in desert dust composition poses a big challenge to understanding Earth’s climate across different epochs. Of particular concern is hematite, an iron-oxide mineral dominating the solar absorption by dust particles, for which current estimates of absorption capacity vary by over two orders of magnitude. Here, we show that laboratory measurements of dust composition, absorption, and scattering provide valuable constraints on the absorption potential of hematite, substantially narrowing its range of plausible values. The success of this constraint is supported by results from an atmospheric transport model compared with station-based measurements. Additionally, we identify substantial bias in simulating hematite abundance in dust aerosols with current soil mineralogy descriptions, underscoring the necessity for improved data sources. Encouragingly, the next-generation imaging spectroscopy remote sensing data hold promise for capturing the spatial variability of hematite. These insights have implications for enhancing dust modeling, thus contributing to efforts in climate change mitigation and adaptation.

Environmental Sciences & Ecology↗

Joint Modeling of GD-1 and C-19 as Old Streams

DESI observational data for the GD-1 and C-19 streams are compared to stream simulations in a common evolving multi-halo potential of a Milky Way-like galaxy based on a cosmological simulation. The goal is to find the best match of the stream velocity spread and the density power spectrum stream density to simulations having either CDM or WDM subhalos. The cocoon velocity width integrated over the length of the stream is independent of orbital blurring along the stream and the power spectrum integrates over the width of the stream, sidestepping the geometric details of the streams. Streams develop from star clusters inserted at $\simeq$1 Gyr after the Big Bang and evolved for 13 Gyr to their current orbital positions. Streams in a CDM subhalo population provide the best match to the velocity width, with streams younger than 10 Gyr ruled out as insufficiently hot. The progenitor star cluster masses, which determine the fraction of stars released at late times which comprise the stream core, are found to be $\simeq 8\times 10^4 M_\odot$ for GD-1 and $\simeq 4\times 10^4 M_\odot$ for C-19, although the mass depends on the star cluster half mass radius. Stream heating leads to stream lumpiness which is measurable for the relatively large and clean GD-1 dataset. The stream density power spectrum measured along the length of the DESI GD-1 sample is in good agreement with CDM simulations, with 1.7 to 1.9 times more power than WDM 7 keV and 5.5 keV simulations.

Carlberg, Raymond G. [Toronto U.] (ORCID:000000027↗

North America’s Potential for an Environmentally Sustainable Nickel, Manganese, and Cobalt Battery Value Chain

The Detroit Big Three General Motors (GMs), Ford, and Stellantis predict that electric vehicle (EV) sales will comprise 40–50% of the annual vehicle sales by 2030. Among the key components of LIBs, the LiNixMnyCo1−x−yO2 cathode, which comprises nickel, manganese, and cobalt (NMC) in various stoichiometric ratios, is widely used in EV batteries. This review reveals NMC cathodes from laboratory research. Furthermore, this study examines the environmental effect of NMC cathode production for EV batteries (including coating technologies), encompassing aspects such as energy consumption, water usage, and air emissions. Although gaps persist in NMC cathode environmental assessments (NMC111, NMC532, NMC622, and NMC811), limited life cycle assessments “(LCA)” have been conducted. Most available data originate from Asia (primarily China), accounting for 85% of the production of EV LIB cathode materials. The concept of battery passports for data collection on LIB components has been proposed to facilitate material traceability as a system for ensuring a sustainable supply chain for critical minerals. The automotive industry’s shift to electrification necessitates a sustainable supply chain from mine to vehicle end-of-life. As the critical mineral supply moves from Asia to North America, environmentally friendly industrial methods must be studied to provide this supply chain direction.

25 ENERGY STORAGE↗

Comparing Compressed and Full-Modeling analyses with FOLPS: implications for DESI 2024 and beyond

The Dark Energy Spectroscopic Instrument (DESI) will provide unprecedented information about the large-scale structure of our Universe. In this work, we study the robustness of the theoretical modelling of the power spectrum of F OLPS , a novel effective field theory-based package for evaluating the redshift space power spectrum in the presence of massive neutrinos. We perform this validation by fitting the AbacusSummit high-accuracy N -body simulations for Luminous Red Galaxies, Emission Line Galaxies and Quasar tracers, calibrated to describe DESI observations. We quantify the potential systematic error budget of F OLPS finding that the modelling errors are fully sub-dominant for the DESI statistical precision within the studied range of scales. Additionally, we study two complementary approaches to fit and analyse the power spectrum data, one based on direct Full-Modelling fits and the other on the ShapeFit compression variables, both resulting in very good agreement in precision and accuracy. In each of these approaches, we study a set of potential systematic errors induced by several assumptions, such as the choice of template cosmology, the effect of prior choice in the nuisance parameters of the model, or the range of scales used in the analysis. Furthermore, we show how opening up the parameter space beyond the vanilla ΛCDM model affects the DESI observables. These studies include the addition of massive neutrinos, spatial curvature, and dark energy equation of state. We also examine how relaxing the usual Cosmic Microwave Background and Big Bang Nucleosynthesis priors on the primordial spectral index and the baryonic matter abundance, respectively, impacts the inference on the rest of the parameters of interest. This paper pathways towards performing a robust and reliable analysis of the shape of the power spectrum of DESI galaxy and quasar clustering using F OLPS .

79 ASTRONOMY AND ASTROPHYSICS↗

Development of a multi-layer canopy model for E3SM Land Model with support for heterogeneous computing

The vertical structure of vegetation canopies creates micro-climates. However, the land components of most Earth System Models, including the Energy Exascale Earth System Model (E3SM), typically neglect vertical canopy structure by using a single layer big-leaf representation to simulate water, CO 2 , and energy exchanges between the land and the atmosphere. In this study, we developed a Multi-Layer Canopy Model for the E3SM Land Model to resolve the micro-climate created by vegetation canopies. The model developed in this study re-implements the CLM-ml_v1 to support heterogeneous computing architectures consisting of CPUs and GPUs and includes three additional optimization-based stomatal conductance models. The use of Portable, Extensible Toolkit for Scientific Computation provides a speedup of 25–50 times on a GPU relative to a CPU. The numerical implementation of the model was verified against CLM-ml_v1 for a month-long simulation using data from the Ameriflux US-University of Michigan Biological Station site. Model structural uncertainty was explored by performing control simulations for five stomatal conductance models that exclude and include the control of plant hydrodynamics (PHD) on photosynthesis. The bias in simulated sensible and latent heat fluxes was lower when PHD was accounted for in the model. Additionally, six idealized simulations were performed to study the impact of three environmental variables (i.e. air temperature, atmospheric CO 2 , and soil moisture) on canopy processes (i.e. net CO 2 assimilation, leaf temperature, and leaf water potential). Increasing air temperature reduced net CO 2 assimilation and increased air temperature. Net CO 2 assimilation increased at higher atmospheric CO 2 , while decreasing soil moisture resulted in lower leaf water potential.

54 ENVIRONMENTAL SCIENCES↗

AmeriFlux FLUXNET-1F US-xSP NEON Soaproot Saddle (SOAP)

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-xSP NEON Soaproot Saddle (SOAP). This is the FLUXNET version of the carbon flux data for the site US-xSP NEON Soaproot Saddle (SOAP) produced by applying the standard ONEFlux (1F) software. Site Description - The Soaproot Saddle is a complex terrain of coarse hills, steep slopes and narrow drainages. With an elevation of 3274 - 4537’ this site encompasses 1438 acres of mixed conifer forests that are experiencing high levels of mortality due to native Pine beetles. At the core of this site stands a 171’ tall flux tower that collects physical and chemical properties of atmosphere and related process. Soaproot Saddle also hosts an array of sensor measurements along with field observations collected by highly trained NEON staff. The automated instrument measurements and some of the terrestrial observational safor this field site are colocted with NEON's aquatic site, Upper Big Creek, which is located just north of Soaproot Saddle's site boundaries.

Network), NEON (National Ecological Observatory [N↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗