Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Assessment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Constraining the impact of chlorine as a neutron absorber in next-gen fast reactor designs

The role of chlorine as a neutron poison and as a seed for producing radioactive waste in nuclear systems has driven a renewed interest to improve its nuclear data uncertainties. Additionally, basic and applied science programs that use CLYC (Cs 2 LiYCl 6 :Ce) detectors for neutron spectroscopy and monitoring are also very sensitive to any change in chlorine nuclear data for simulations of the detector response. In this work, sensitivities relevant for these different applications are addressed through simulations of the efficiency of CLYC detectors in a fast fission spectrum when applying new chlorine nuclear data as input. These simulations are validated by an experimental measurement using CLYC detectors coupled to an ionization chamber loaded with a 252 Cf spontaneous fission source. The results are then used to obtain the first reliable direct measurement of the 35 Cl(n,p 0 ) and summed Cl(n,p+n,α) fission spectrum average cross sections, found to be 54.7(32) and 105.0(98) mb, respectively. The results are within uncertainty of calculated fission spectrum averaged cross sections based on recently re-evaluated chlorine nuclear data, which confirm recent impact studies performed for the Molten Chloride Reactor Experiment. Meanwhile, there currently exists only one published criticality benchmark experiment that is sufficiently sensitive to chlorine nuclear data. Discrepancies are found with this set of criticality safety benchmarks, which are more sensitive to thermal and epithermal neutron energies than the energies, above 100 keV, tested in this current work. Hence, there is still a need to re-evaluate the chlorine nuclear data at lower energies to assess these discrepancies. Interpretation of the data from future “faster” criticality benchmarks, which are needed for next-gen fast reactor designs, benefit from the improved constraints on the chlorine nuclear data validated in this work.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins↗

From Machine Learning to Machine Reasoning: A Model-based Approach to Analyze Equipment Reliability Data

In current nuclear power plants (NPPs) a large amount of condition-based data which can be used to assess and monitor component health and performance. Assessing component health from such data can be performed with a large variety of methods. While the analysis of numeric data can be performed with several methods, the extraction of information from textual data remains a challenge. Currently employed natural language processing (NLP) methods do not really provide quantitative information that might be contained in IRs. In addition, the integration of numeric and textual data to identify possible causal relationships between data elements is still an unresolved challenge. This paper presents an approach to extract information from textual (e.g., incident or maintenance reports) and numeric data that relies on model based system engineer (MBSE) models. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence while semantic analysis is designed to analyze the logic structure of a sentence. An innovative element of our approach is that semantic analysis uses MBSE models to identify links between textual elements. Similarly, numeric data is directly linked to elements of the MBSE models in order to map which functions are being monitored.

97 - MATHEMATICS AND COMPUTING↗

Integrating State Data Assimilation and Innovative Model Parameterization Reduces Simulated Carbon Uptake in the Arctic and Boreal Region

Model representation of carbon uptake and storage is essential for accurate projection of the response of the arctic-boreal zone to a rapidly changing climate. Land model estimates of LAI and aboveground biomass that can have a marked influence on model projections of carbon uptake and storage vary substantially in the arctic and boreal zone, making it challenging to correctly evaluate model estimates of Gross Primary Productivity (GPP). To understand and correct bias of LAI and aboveground biomass in the Community Land Model (CLM), we assimilated the 8-day Moderate Resolution Imaging Spectroradiometer (MODIS) LAI observation and a machine learning product of annual aboveground biomass into CLM using an Ensemble Adjustment Kalman Filter (EAKF) in an experimental region including Alaska and Western Canada. Assimilating LAI and aboveground biomass reduced these model estimates by 58% and 72%, respectively. The change of aboveground biomass was consistent with independent estimates of canopy top height at both regional and site levels. The International Land Model Benchmarking system assessment showed that data assimilation significantly improved CLM's performance in simulating the carbon and hydrological cycles, as well as in representing the functional relationships between LAI and other variables. Here, to further reduce the remaining bias in GPP after LAI bias correction, we re-parameterized CLM to account for low temperature suppression of photosynthesis. The LAI bias corrected model that included the new parameterization showed the best agreement with model benchmarks. Combining data assimilation with model parameterization provides a useful framework to assess photosynthetic processes in LSMs.

58 GEOSCIENCES↗

Quality Ranking of Unary Chloride Salt Property Data Included in MSTDB-TP

Molten salt reactor developers rely on thermal property data to design, license and operate the reactors. The Molten Salt Thermal Database-Thermophysical Properties (MSTDB-TP) was established under the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program and is managed by Oak Ridge National Laboratory to serve as a single source of thermophysical property values measured for a wide variety of molten salt systems for use by researchers, molten salt reactor developers, and regulators. These properties include density, viscosity and thermal diffusivity and conductivity. Published measurements of molten salt properties are lacking for many salts of interest and the data that are available are often inconsistent. This creates a challenge for MSR developers when determining which property values to use when designing their reactors. It is the purpose of this work to apply a consistent ranking system to all data entries that indicates the quality of property values listed in the database. These rankings will be the technical basis for down-selections by the database developers and alert users about the quality of the available property values. MSTDB-TP collects all available property data and indicates preferred data sets or correlations. However, all available data sets are included in the database. Quality assessments and rankings are being applied to data in MSTDB-TP to provide an indication of the quality of each data set independent of consistency with other data. Previous reports detailed the ranking system that was followed and assessments of unary fluoride data sets. Documentation of the quality of data in MSTDB-TP was continued by reviewing and assessing all available sources of density, viscosity and thermal diffusivity or conductivity values for unary chloride salts in MSTDB-TP V3.0 using the same criteria.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Challenges of standard halo models in constraining galaxy properties from cosmic infrared background anisotropies

The halo model, combined with halo occupation distribution (HOD) prescriptions, is widely used to interpret cosmic infrared background (CIB) anisotropies and extract physical information about star-forming galaxies and their connection to large-scale structures. Recent CIB-specific implementations of the halo model have adopted more physical parameterizations. However, the extent to which these models can reliably recover meaningful physical parameters remains uncertain. We assessed whether the current parameterization of CIB halo models is sufficient to recover astrophysical quantities, such as star formation efficiency, η(M h , z), and halo mass at which the peak of star formation efficiency occurs, M max , when fit to mock data. We also assessed whether discrepancies arise from assumptions about galaxy emission (the HOD ingredients) or from more fundamental components in the halo model, such as bias and matter clustering. We fit the M21 CIB HOD model, implemented within the halo model framework, to mock CIB power spectra and star formation rate density (SFRD) data generated from the SIDES-Uchuu simulation, and compared the best-fit parameters to the known simulation inputs. We then repeated the analysis using a simplified version of the simulation (SSU), explicitly designed to match the HOD assumptions. A detailed comparison of model and simulation outputs was carried out to trace the origin of observed discrepancies. While the M21 HOD model provides a good fit to the mock data, it failed to recover the intrinsic parameters accurately, particularly the halo mass at which star formation efficiency peaks. This mismatch persists even when fitting data generated with the same model assumptions. We find strong agreement (within 5%) in the emission-related components (SFRD, emissivity), but observe a scale- and redshift-dependent offset exceeding 20% in the two-halo term of the CIB power spectrum. This likely arises from limitations in the treatment of halo bias and matter clustering within the linear approximation. Additionally, incorporating scatter in the SFR–halo mass relation and the spectral energy distribution (SED) templates significantly affects the shot noise (∼50%), but has only a modest impact (less than 10%) on the clustered component. These results suggest that recovering physical parameters from CIB clustering requires improvements to the cosmological ingredients of the halo model framework, such as adopting scale-dependent halo bias and nonlinear matter power spectra in addition to careful modeling of emission physics.

cosmic background radiation↗

Measuring the Stress Factors for Photovoltaic (PV) Backsheet Degradation

Back sheet failure has resulted in power loss and large-scale recall of photovoltaic modules, resulting in billions of dollars in lost revenue. The light exposure on the backside of a photovoltaic module comes primarily from reflected light which alters the distribution of natural sunlight. Because of this, modelling the backside exposure and duplicating the exposure is much more difficult than modeling the frontside exposure. This project aims to study how various back sheets and junction box materials degrade under different conditions and to develop Python code to help model and predict degradation. The stress factors for back-sheet degradation must be quantified to extrapolate accelerated stress tests to the field. Test samples were placed in the A3, A4, and A5 conditions, as defined in IEC 62788-7-2, to assess the temperature and humidity dependence of ultraviolet (UV) induced degradation. We are utilizing a custom chamber with exposure from 0.5 UV-suns to 5 UV-suns to understand the dependence of degradation on light intensity. A group of samples put in the A3 condition had glass filters with 50% UV cut-offs of 320 nm, 335 nm, and 360 nm to assess the wavelength dependence of UV degradation. All this data is necessary to assess the impact of non-standard UV light exposure. The material evaluation tests include gloss measurements, attenuated total internal reflectance Fourier transform infrared spectroscopy (ATR-FTIR), UV-visible reflectance/transmittance utilizing a Cary Ci7000 spectrophotometer, and a nano-indenter for surface hardness and modulus measurements. Alongside the experimental work, there is a computational effort using raytracing and Python open-source tools in PVDeg , PVLib, and Bifacial_Radiance. This code will create specific exposure scenarios and enable the evaluation of chamber degradation relative to field degradation. Equation 1 is a strawman equation used to model degradation on the backside of a PV module. We will create simplified code, based on the results of ray-tracing calculations, which uses a view factor approach to provide fast calculations for the most common exposure scenarios.

14 SOLAR ENERGY↗

Experimental design for the Marine Ice Sheet–Ocean Model Intercomparison Project – phase 2 (MISOMIP2)

The Marine Ice Sheet–Ocean Model Intercomparison Project – phase 2 (MISOMIP2) is a natural progression of previous and ongoing model intercomparison exercises that have focused on the simulation of ice-sheet and ocean processes in Antarctica. The previous exercises motivate the move towards realistic configurations, as well as more diverse model parameters and resolutions. The main objective of MISOMIP2 is to investigate the performance of existing ocean and coupled ice-sheet–ocean models in a range of Antarctic environments through comparisons to observational data. We will assess the status of ice-sheet–ocean modelling as a community and identify common characteristics of models that are best able to capture observed features. As models are highly tuned based on present-day data, we will also compare their sensitivity to prescribed abrupt atmospheric perturbations leading to either very warm or slightly warmer ocean conditions compared to the present day. The approach of MISOMIP2 is to welcome contributions of models as they are, including global and regional configurations, but we request standardized variables and common grids for the outputs. We target the analysis at two specific regions, the Amundsen Sea and the Weddell Sea, since they describe two different ocean environments and have been relatively well observed compared to other areas of Antarctica. An observational “MIPkit” synthesizing existing ocean and ice-sheet observations for a common period is provided to evaluate ocean and ice-sheet models in these two regions.

58 GEOSCIENCES↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Landscape analysis of environmental data sources for linkage with SEER cancer patients database

Abstract One of the challenges associated with understanding environmental impacts on cancer risk and outcomes is estimating potential exposures of individuals diagnosed with cancer to adverse environmental conditions over the life course. Historically, this has been partly due to the lack of reliable measures of cancer patients’ potential environmental exposures before a cancer diagnosis. The emerging sources of cancer-related spatiotemporal environmental data and residential history information, coupled with novel technologies for data extraction and linkage, present an opportunity to integrate these data into the existing cancer surveillance data infrastructure, thereby facilitating more comprehensive assessment of cancer risk and outcomes. In this paper, we performed a landscape analysis of the available environmental data sources that could be linked to historical residential address information of cancer patients’ records collected by the National Cancer Institute’s Surveillance, Epidemiology, and End Results Program. The objective is to enable researchers to use these data to assess potential exposures at the time of cancer initiation through the time of diagnosis and even after diagnosis. The paper addresses the challenges associated with data collection and completeness at various spatial and temporal scales, as well as opportunities and directions for future research.

60 APPLIED LIFE SCIENCES↗

DICER: Data Intensive Computing Environment and Runtime for Evaluating Unprecedented Scale of Geospatial-Temporal Human Mobility Data

With the significant increase in sources and volume of human mobility data through commercial data vendors as well as microsimulation of cities, the scale of geospatial-temporal data to analyze and assess for mobility characterization has grown to the level of Big Data. There are mobility related commercial organizations deploying scalable computing, but often the system architecture, workflow, and intermediate processing components are not fully disclosed in relevant scope. Current research literature has a notable lack of studies demonstrating architectures and workflows for human mobility analytics that are implemented on a TeraByte scale of geospatial-temporal data. In this context, this paper presents a hyperscale-level system solution named DICER (Data Intensive Computing Environment and Runtime) for processing and analytics of geospatial-temporal data at big data scale. Although the cluster computing architecture of DICER with Apache Spark job running on Kubernetes cluster is not new, there are innovations in the workflow, hierarchical processing logic, and a wide range of intermediate preprocessing and mobility metrics calculation. We have performed case studies to validate the effectiveness of DICER system solution by performing detailed analytics and assessment of human mobility microsimulation output at three different scopes and scale, including a usecase with 16.97 TeraByte and 259.2 Billion rows of data. In addition, we have presented another case study of utilizing DICER to perform the same mobility processing and comparative analytics on large-scale commercially available geospatial-temporal data. All these case studies validate the efficiency and usefulness of DICER in computing population mobility characteristics from geospatial-temporal trajectory data at an unprecedented scale (not only just data volume, but also combination of: number of user entities, temporal frequency, spatial resolution, data duration).

De, Debraj↗

New Matrix Framework to Determine Carbon Storage Technical Viability

Carbon storage is an integral component of reducing CO2 emissions. Research over the past two decades has developed workflows for volume assessments and economic project feasibility, providing necessary and useful tools to progress geologic carbon storage (GCS) projects. These workflows and assessments have focused primarily on determining the in-situ storage resource based on geologic and engineering parameters and do not integrate subsurface characterization with surface conditions, social factors, and environmental factors that may pose a benefit or impediment to the implementation of GCS. Furthermore, the data required to assess the technical viability of GCS are myriad and disparate. There is no current methodology that identifies technical viability criteria and systematically informs how to aggregate these factors for spatial assessments. To address this gap, the National Energy Technology Laboratory has developed a Carbon Storage Technical Viability Approach (CS TVA) Matrix that incorporates a wide variety of factors to inform and accelerate screening for GCS site selection in the United States.

Mulhern, Julia↗

Small Nuclear Reactors for Maritime Ports

Maritime ports are entering a period of sharply rising electricity demand as operations are electrified, shore power adoption grows, and energy-intensive industries expand near ports. This report evaluates the feasibility of small nuclear reactors at U.S. ports, which could offer stable baseload power and reduced dependence on regional grids. The report describes the current state of advanced reactor technologies and other factors important for future deployment such as safety and operations, siting requirements, regulatory environments and policy, and other co-benefit considerations. A structured evaluation across a diverse group of U.S. ports highlights the wide variation in feasibility. Fifteen ports were analyzed through port interviews and data collection to assess the port-specific conditions that most impact advanced reactor deployment. The analysis found that the most feasible ports for deployment have favorable geotechnical stability, supportive regulatory environments, and substantial energy loads in a grid-constrained environment. Additionally, an initial reactor size matching was done for the ports based on their reported energy consumption across different modes of operation. However, further analysis and detailed data will be required in the future to determine the optimal reactor-port matching. The economic viability to adopt this technology will be dictated by competitive cost of nuclear-generated electricity compared to other alternatives. Overall, small nuclear reactors present a promising but highly site-dependent pathway for ports seeking resilient, high-capacity energy solutions that support future energy demands. This report identifies key areas for future research and highlights ports that warrant further evaluation to assess the viability of small nuclear reactor implementation. By establishing these priorities, the analysis provides support for the development of informed strategies for potential advanced nuclear technology deployment. The inclusion of ports in this report does not imply their endorsement of nuclear energy generation, and the authors are grateful for the information, expertise, and perspectives they contributed.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Polymer Size–Catalytic Activity Relationships in Solution by Fluorescence Correlation Spectroscopy

Measuring the catalytic activity of specific sizes of polymers with active catalysts in solution is typically challenging, due to limited instrument detection sensitivity and/or dynamic range. Here, a fluorescence correlation spectroscopy (FCS) method is developed to determine the catalytic activity of living polymers of a specific apparent size in solution. Deviation from a single-component FCS data fitting, as assessed by χ2, is also introduced and developed as a “speciation index”—a method to evaluate and track changes in the relative amount of distinct polymer sizes with reaction progress. These methods are enabled by incorporating a selectively reactive fluorescent monomer into growing polydicyclopentadiene or polynorbornene during ring-opening metathesis polymerization (ROMP). Compared to polynorbornene, data showed that catalysts in aggregates of polyDCPD retained higher activity for longer—outcomes not directly inferable from simple diffusional-access predictions. Here, the ability to assign catalytic activity to polymers of specific sizes, and then to determine how this activity evolves with reaction progress, support long-term goals in the development and measurement of nano-objects that possess size-dependent catalytic activity.

Active catalyst↗

The Carbon Storage Technical Viability Approach (CS TVA)

The Carbon Storage Technical Viability Approach (CS TVA) StoryMap provides an in-depth overview of the products created during the CS TVA research effort. In detail, the StoryMap addresses the CS TVA Matrix, Database Version 2.0, Database Catalog, Workflow, and Data Availability Result Database, discussing how each was developed and implemented. Information on how the matrix, database, and database catalog are interconnected, and their usage is also explained. The workflow section provides information on the CS TVA product development from the data-gathering stage to the final data availability results. A section on an expansion of the CS TVA workflow that utilizes Natural Language processing (NLP) section was included. Finally, the Data Availability Results Database is discussed. These results provide data science-informed insights into potential data gaps when assessing the viability of carbon storage in a given area or region.

Carbon Storage↗

3D Deep Learning Joint Inversion of Active Seismic Full Waveform and Passive Seismic Traveltime Data for Reservoir Imaging and Uncertainty Quantification

Here, we present deep learning (DL) networks for three-dimensional (3D) joint inversion of active seismic full waveform and passive seismic traveltime data to image reservoirs and their properties and quantify imaging uncertainties. Active seismic full-waveform data can provide high-resolution monitoring images but are collected only intermittently because of their high acquisition cost. In contrast, passive seismic data can be gathered at relatively low cost between regular active surveys, although their imaging quality can be compromised by factors such as low signal-to-noise ratios and limited ray coverage of the target. Although these datasets are routinely acquired together at CO 2 storage sites, their combined inversion within a 3D DL framework has not been previously demonstrated. To our knowledge, this is the first study to address this gap, combining the strength of both data types. For efficient data storage and DL training with large 3D seismic datasets, we use a 3D data matrix in which a random number of passive seismic traveltime data are stored as parabolic envelopes using one-hot encoding and a 3D full-waveform data matrix in which multiple shot gathers are summed. Two network architectures are evaluated: a single-encoder U-Net for single-data type inversion and a dual-encoder U-Net for joint inversion of active and passive seismic data. We also evaluate the single-encoder U-Net for joint inversion by concatenating full-waveform data and traveltime data. We propose a systematic approach for selecting an optimal dropout rate that balances regularization during training and Monte Carlo dropout-based uncertainty quantification during prediction by examining the correlation coefficient between standard deviation and prediction error, along with the training misfit, across a range of dropout rates. 3D DL inversion experiments include five different network configurations, with evaluations under ideal, noisy and dropout-enabled conditions. Both model and data uncertainties are assessed, as well as their combined effects. Across all conditions, the networks consistently predict accurate CO 2 saturation models with low prediction errors, such as a structural similarity index of 0.993 and CO 2 difference of 1.1%. Uncertainty estimates show strong spatial correlation with prediction errors, confirming the effectiveness of the proposed dropout selection approach. The results demonstrate that our DL approach, utilizing compact data representations and appropriate uncertainty quantification, yields accurate subsurface images under various inversion conditions and provides valuable insights into the reliability of predictions.

Um, Evan Schankee [Lawrence Berkeley National Labo↗

A Data-driven approach to Core Power distribution reconstruction in a Nuclear Reactor

This report presents the initial development of a data-driven approach for reconstructing the core power distribution in a nuclear reactor (power shape synthesis) using ex-core sensors. Traditional techniques rely on deploying a large number of detectors throughout the reactor core. However, this approach is not feasible for innovative reactor concepts like Advanced Reactors and Microreactors. First, the tight lattice pitch, designed to maximize power density, limits the space available for sensors. Secondly, the harsh operating conditions are not compatible with commercially available detectors. The method proposed in this work integrates high-fidelity modeling with data-driven techniques to accurately reconstruct power distribution across various reactor types, thereby reducing the reliance on in-core sensors. Purdue University Reactor One (PUR-1) was selected as the test case. The CAD model representing the latest configuration of the PUR-1 core was imported into the OpenMC simulation framework, and the model was built. Additionally, the previously developed MCNP6 model was updated. The two models were assessed against the data collected during an experimental campaign conducted in July 2024. Thirty gold foils were placed in three Irradiation Assemblies in PUR-1 core. Using the measured activity of the irradiated foils, the neutron flux at different core locations was estimated.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Tractometry of the Human Connectome Project: resources and insights

The Human Connectome Project (HCP) has become a keystone dataset in human neuroscience, with a plethora of important applications in advancing brain imaging methods and an understanding of the human brain. We focused on tractometry of HCP diffusion-weighted MRI (dMRI) data. We used an open-source software library (pyAFQ; https://yeatmanlab.github.io/pyAFQ) to perform probabilistic tractography and delineate the major white matter pathways in the HCP subjects that have a complete dMRI acquisition (n = 1,041). We used diffusion kurtosis imaging (DKI) to model white matter microstructure in each voxel of the white matter, and extracted tract profiles of DKI-derived tissue properties along the length of the tracts. We explored the empirical properties of the data: first, we assessed the heritability of DKI tissue properties using the known genetic linkage of the large number of twin pairs sampled in HCP. Second, we tested the ability of tractometry to serve as the basis for predictive models of individual characteristics (e.g., age, crystallized/fluid intelligence, reading ability, etc.), compared to local connectome features. To facilitate the exploration of the dataset we created a new web-based visualization tool and use this tool to visualize the data in the HCP tractometry dataset. Finally, we used the HCP dataset as a test-bed for a new technological innovation: the TRX file-format for representation of dMRI-based streamlines. We released the processing outputs and tract profiles as a publicly available data resource through the AWS Open Data program's Open Neurodata repository. We found heritability as high as 0.9 for DKI-based metrics in some brain pathways. We also found that tractometry extracts as much useful information about individual differences as the local connectome method. We released a new web-based visualization tool for tractometry—“Tractoscope” (https://nrdg.github.io/tractoscope). We found that the TRX files require considerably less disk space-a crucial attribute for large datasets like HCP. In addition, TRX incorporates a specification for grouping streamlines, further simplifying tractometry analysis.

59 BASIC BIOLOGICAL SCIENCES↗