Search NASA⌕ Search

SEARCH · Search NASA

Results for “data platform”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

From Data to Insights: A Covariate Analysis of the IARPA BRIAR Dataset for Multimodal Biometric Recognition Algorithms at Altitude and Range

This paper examines covariate effects on fused whole body biometrics performance in the IARPA BRIAR dataset, specifically focusing on UAV platforms, elevated positions, and distances up to 1000 meters. The dataset includes outdoor videos compared with indoor images and controlled gait recordings. Normalized raw fusion scores relate directly to predicted false accept rates (FAR), offering an intuitive means for interpreting model results. A linear model is developed to predict biometric algorithm scores, analyzing their performance to identify the most influential covariates on accuracy at altitude and range. Weather factors like temperature, wind speed, solar loading, and turbulence are also investigated in this analysis. The study found that resolution and camera distance best predicted accuracy and findings can guide future research and development efforts in long-range/elevated/UAV biometrics and support the creation of more reliable and robust systems for national security and other critical domains.

Bolme, David↗

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE↗

Development and Validation of a Process Model and Open-Source Process Simulator for Microalgae-Based Tertiary Phosphorus Recovery

Microalgae-based tertiary wastewater treatment has the potential to meet stringent effluent phosphorus limits, with the added benefit of producing a marketable feedstock. However, the lack of validated mechanistic models and their implementation in process simulators have limited the adoption of this technology. In this study, an updated lumped pathway metabolic model (Phototrophic-Mixotrophic Process Model, PM 2 ), including both photoautotrophic and heterotrophic metabolisms of microalgae, was developed to predict effluent phosphorus concentration and biomass yield in response to dynamic influent and varying environmental conditions. The model was implemented in QSDsan – an open-source, Python-based design and simulation platform – for robust simulation under uncertainty. A global sensitivity analysis was performed to prioritize model parameters for calibration. The model was then calibrated and validated using batch experimental data and 45 days of continuous online monitoring data from a full-scale (568 m 3 ·d -1 ) microalgae-based tertiary wastewater treatment plant (EcoRecover process). In particular, along with dynamic influent composition, temperature and light intensity data with diel variation were provided as model inputs to reflect the microalgal behavior under day-night cycling. Overall, the QSDsan-based microalgae process simulator was able to predict effluent phosphorus within 0.02–0.04 mg-P·L -1 , while also capturing the general trends of state variables according to nutrient availability.

Lumped pathway metabolic model↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

Rapid Evaluation Framework for the CMIP7 Assessment Fast Track

As Earth system models (ESMs) grow in complexity and in volume of output data, there is an increasing need for rapid, comprehensive evaluation of their scientific performance. The upcoming Assessment Fast Track for the Seventh Phase of the Coupled Model Intercomparison Project (CMIP7) will require expeditious response for model analyses designed to inform and drive integrated Earth system assessments. To meet this challenge, the Rapid Evaluation Framework (REF), a community-driven platform for benchmarking and performance assessment of ESMs, was designed and developed. The initial implementation of the REF, constructed to meet the near-term needs of the CMIP7 Assessment Fast Track, builds upon four disparate community evaluation and benchmarking tools that are coupled together using the Coordinated Model Evaluation Capabilities (CMEC) framework. The REF runs within a containerized workflow for portability and reproducibility and is aimed at generating and organizing diagnostics covering a variety of model variables. The REF leverages well documented observational datasets to provide assessments of model fidelity across a collection of diagnostics. All diagnostics were identified and selected with community involvement and consultation. Operational integration with the Earth System Grid Federation (ESGF) will permit automated execution of the REF for selected diagnostics as soon as model output data are published on ESGF by the originating modeling centers. The REF is designed to be portable across a range of current computational platforms to facilitate use by modeling centers for assessing the evolution of model versions or gauging the relative performance of CMIP simulations before being published on ESGF. When integrated into production simulation workflows, results from the REF provide immediate quantitative feedback that allows model developers and scientists to quickly identify model biases and performance issues. After the REF is released to the community, its subsequent development and support will be prioritized by an international consortium of scientists and engineers, enabling a broader impact across Earth science disciplines. For instance, the REF will facilitate improvements to models and will enhance confidence in model projections through process-based selection of models based on their performance with respect to observations. Production of reproducible diagnostics and community-based assessments are key features of the REF. Furthermore, providing interoperability with existing evaluation packages assures that contributions from previous community efforts will be available for use in future model intercomparison projects.

Hoffman, Forrest [ORNL] (ORCID:0000000158024134)↗

Using AI to Reproduce Neutrino Cross Section Analysis - Prototyping the Neutrino Discovery Platform

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]↗

CatTestHub: A benchmarking database of experimental heterogeneous catalysis for evaluating advanced materials

The ability to quantitatively compare newly evolving catalytic materials and technologies is hindered by the widespread availability of catalytic data collected in a consistent manner. While certain catalytic chemistries have been widely studied across decades of scientific research, quantitative comparisons based on literature information is hindered by variability in reaction conditions, types of reported data, and reporting procedures. Here, we present CatTestHub, an open-access database dedicated to benchmarking experimental heterogeneous catalysis data. Combining systematically reported catalytic activity data for selected probe chemistries, with relevant material characterization and reactor configuration information, the database provides a collection of catalytic benchmarks for distinct classes of active site functionality. Through key choices in data access, availability, and traceability, CatTestHub seeks to balance the fundamental information needs of chemical catalysis and the FAIR data design principles. Details of the database architecture and the means through which to navigate it are presented, highlighting examples of catalytic insights readily drawn from the available benchmarking data. In its current iteration, CatTestHub spans over 250 unique experimental data points, collected over 24 solid catalysts, that facilitated the turnover of 3 distinct catalytic chemistries. Here, a roadmap is presented through which to expand the open-access platform that serves as a community wide benchmark, primarily through continuous addition of kinetic information on select catalytic systems by members of the heterogeneous catalysis community at large.

Benchmark↗

Analog-to-digital converter based on voltage-controlled superconducting devices

The increasing demand for cryogenic electronics in superconducting and quantum computing systems calls for ultra-energy-efficient data conversion architectures that remain functional at deep cryogenic temperatures. Here, in this work, we present the first design of a voltage-controlled superconducting flash analog-to-digital converter (ADC) based on a voltage-controlled quantum-enhanced Josephson junction field-effect transistor (JJFET). Exploiting its strong gate tunability and transistor-like behavior, the JJFET offers a scalable alternative to conventional current-controlled superconducting devices while aligning naturally with CMOS-style design methodologies. Building on our previously developed Verilog-A compact model calibrated to experimental data, we design and simulate a three-bit JJFET-based flash ADC targeted for integration within cryogenic control and readout circuitry in quantum computing. The core comparator block is realized through careful bias current selection and augmented with a three-terminal nanocryotron to precisely define reference voltages. Cascaded JJFET comparators ensure robust voltage gain, cascadability, and logic-level restoration across stages. Simulation results demonstrate accurate quantization behavior with ultra-low power dissipation, underscoring the feasibility of voltage-driven superconducting mixed-signal circuits. This work establishes a critical step toward unifying superconducting logic and data conversion, paving the way for scalable cryogenic architectures in quantum–classical co-processors, low-power artificial intelligence accelerators, and next-generation energy-constrained computing platforms.

Analog-to-digital converter↗

ACTIVE

The Automated Control Testbed for Integration, Verification, and Emulation (ACTIVE) framework is a software platform designed to support the optimized operation and management of a wide range of building types. It enables the development, testing, and validation of diverse control strategies, including AI-based, rule-based, and model-based approaches. The platform facilitates a seamless transition from simulation-based evaluation of control strategies to real-world field validation and deployment. ACTIVE supports the full building management lifecycle, encompassing data acquisition and management, system monitoring, optimized control, adaptive learning services, device dispatch and coordination, as well as advanced analytics and visualization. Together, these capabilities provide an integrated environment for improving building performance, operational efficiency, reducing energy cost, and reliability.

Smith, Robert [Oak Ridge National Laboratory (ORNL↗

Water Observations of Flow/No-Flow for the East-Taylor Watershed, Colorado (June-July 2025 and 2026)

This dataset provides multi-year, ground-truth visual observations of surface water flow/no-flow conditions within the East-Taylor Watershed, Colorado, collected during June and July of 2025 and 2026. In June and July 2025, on-the-ground visual observations of flow/no-flow were collected as part of the Watershed Function Scientific Focus Area (SFA) and Rocky Mountain Biological Laboratory (RMBL) Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign (further details are provided within the CHESS Project Description). We obtained 377 water observations of flow/no-flow within the East-Taylor Watershed, Colorado. These ground-truth observations were collected to validate classification maps from remote sensing data and model results within the East-Taylor Watershed. In 2025, flow/no-flow measurements were collected using a field-based app for the CHESS Campaign (Zerion iForm). Within the field app, a water observation form was created to collect coordinates and metadata about the observation. Information collected for the water observation points included information about visually-assessed streamflow presence/absence (standard question obtained from Colorado State University’s StreamTracker project), flow estimate, stream or ponded area width, canopy cover, manganese films, iron seeps, and beaver activity. For 2025 water observations, this dataset contains: (1) a data file with the water observations and coordinates (2025_Water_Observations.csv); (2) a Keyhole Markup Language Zipped (KMZ) with the water observation locations and metadata (2025_Water_Observations_Locations.kmz); (3) photos (.jpg and .jpeg) of the water observation points, organized by location, contained within 2025_Water_Observations_FieldPhotographs.zip file; and (4) water observation protocols and figures (2025_Water_Observation_Protocols.pdf). In June and July 2026, on-the-ground visual observations of flow/no-flow were collected as part of the Watershed Function SFA project. We obtained 365 water observations of flow/no-flow within the East-Taylor Watershed, Colorado. The 2026 observations focused on collecting repeat measurements at the 2025 flow/no-flow observation locations conducted as part of the CHESS campaign. These ground-truth observations were collected to understand differences in flow/no-flow in 2026, given the unprecedented 2026 drought in Colorado. In 2026, flow/no-flow measurements were collected using ArcGIS (Geographic Information System) Survey123. Within the field app, a water observation form was created to collect coordinates and metadata about the observation. Information collected for the water observation points included repeat information from the 2025 water observation effort, including visually-assessed streamflow presence/absence (standard question obtained from Colorado State University’s StreamTracker project), flow estimate, stream or ponded area width, canopy cover, manganese films, iron seeps, beaver activity, and a new metadata component of estimated stream depth (for select locations). For 2026 water observations, this dataset contains: (1) a data file with the water observations and coordinates (2026_Water_Observations.csv); (2) a Keyhole Markup Language Zipped (KMZ) with the water observation locations and metadata (2026_Water_Observations_Locations.kmz); (3) photos (.jpg) of the water observation points, organized by location, contained within 2026_Water_Observations_FieldPhotographs.zip file; and (4) water observation protocols and figures (2026_Water_Observation_Protocols.pdf). For 2025 and 2026 water observations, this dataset contains: (1) a location metadata file (locations.csv); (6) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and (7) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. 2026-09-02: This dataset was updated to include 2026 water observation measurements. The 2025 observation files were also updated to ensure a consistent file naming convention across water observation years.

2018 NEON and 2025 CHESS Campaigns↗

Facile Tensile Testing Platform for In Situ Transmission Electron Microscopy of Nanomaterials

In situ tensile testing using transmission electron microscopy (TEM) is a powerful technique to probe structure-property relationships of materials at the atomic scale. In this work, a facile tensile testing platform for in situ characterization of materials inside a transmission electron microscope is demonstrated. The platform consists of: 1) a commercially available, flexible, electron-transparent substrate (e.g., TEM grid) integrated with a conventional tensile testing holder, and 2) a finite element simulation providing quantification of specimen-applied strain. The flexible substrate (carbon support film of the TEM grid) mitigates strain concentrations usually found in free-standing films and enables in situ straining experiments to be performed on materials that cannot undergo localized thinning or focused ion beam lift-out. The finite element simulation enables direct correlation of holder displacement with sample strain, providing upper and lower bounds of expected strain across the substrate. The tensile testing platform is validated for three disparate material systems: sputtered gold-palladium, few-layer transferred tungsten disulfide, and electrodeposited lithium, by measuring lattice strain from experimentally recorded electron diffraction data. The results show good agreement between experiment and simulation, providing confidence in the ability to transfer strain from holder to sample and relate TEM crystal structural observations with material mechanical properties.

2D materials↗

DuctGPT: A Generative Transformer for Forward Screening of Ductile Refractory Multi-Principal Element Alloys

Designing ductile materials for extreme environments such as fusion reactors requires a deep understanding of the complex interplay between electronic structure, mechanical stability, and wide compositional space. Here, in this work, we introduce DuctGPT, a physics-informed, GPT-powered machine learning platform that enables rapid and accurate prediction of ductility across a wide range of refractory multi-principal element alloys (MPEAs). Trained on both experimental and high-fidelity computational data, DuctGPT integrates descriptors such as density of states at the Fermi level, elastic constants, and valence electron concentration to capture the fundamental mechanisms governing ductile versus brittle behavior. Using this framework, we screen over 1000 compositions in of body-centered cubic (BCC) MPEAs, including two new alloy classes, i.e., NbTa-rich (NbTa $>$ 50 at.%) NbTa-Ti-V and W-rich ($>$ 50 at.%) W-Ti-V MPEAs, to rapidly identify promising alloy compositions with enhanced ductility. Validation against experimental data confirms the model's ability to predict ductility with high fidelity and low uncertainty. By leveraging conversational AI and robust physical modeling, DuctGPT provides a blueprint for the next generation of alloy design assistants, enabling human-AI collaboration in the accelerated discovery of ductile, high-performance materials for fusion, aerospace, and advanced manufacturing.

AI/ML↗

EPICS for small-scale laboratories with Python soft IOCs

While the Experimental Physics and Industrial Control System (EPICS) is widely used at large laboratories for slow controls and instrumentation, the deployment of a full EPICS installation can be difficult, with a steep learning curve to new users. Taking advantage of the pythonSoftIOC module, we developed an EPICS slow controls implementation for Jefferson Lab's Hall B cryotarget written entirely in Python and based on software IOCs that communicate with instruments over Ethernet. Here, this system ran successfully, interfacing with Jefferson Lab's full EPICS network, and we offer it as an example of the capabilities of pythonSoftIOC to build lightweight, yet robust and flexible instrumentation platforms that would be easily adapted for use at a small-scale laboratory. University groups can use these examples to build complete slow controls systems, from device communication to data archiving and display, using open-source, mature EPICS tools and student-friendly Python as an alternative to expensive and proprietary systems such as LabVIEW.

Computing↗

Synapse v1.0

Synapse (SYNergistic software platform for AI, Physics Simulations, and Experiments) is a software package meant to deploy real-time guidance from simulations during experimental campaigns, The software package contains functionalities to collect data from simulations (e.g. running at NERSC) and experiments (e.g. from the BELLA facility at LBNL) into a database, train ML surrogate models from this data, and display the predictions of the surrogate model in the control room of an experimental facility, so as to guide on-going experimental campaign. This software was developed as part of an on-going LDRD.

Lehe, Remi [Lawrence Berkeley National Laboratory ↗

Integrating Immersive Visualization in Molten-Salt Reactor Waste Management for Experimental Design and Planning

The Molten Salt Reactor (MSR) represents a significant innovation in nuclear technology, offering several operational and safety benefits over traditional solid-fuel reactors. However, MSRs face uncertainties in waste management due to their flexible designs and variable waste compositions. To address these challenges, we propose a visualization platform that illustrates solutions and performance predictions for various waste management strategies, enhancing user experience and improving strategy and communication. Immersive visualizations are widely used in the nuclear industry for training, simulation, and safety enhancement. Our project aims to develop a visualization platform incorporating virtual reality (VR) technologies to illustrate MSR characteristics immediately following reactor shutdown. This immersive simulation will allow users to interact, explore, and understand different waste management strategies. The platform will display MSR reactor characterizations, including nuclide decay, salt solidification, and corrosion, which are crucial for assessing and selecting backend management strategies. Using the Meta Quest 3 VR headset with Unity software, our platform will provide real-scale visualizations, enabling users to experience and evaluate designs and plans as if they were physically present. This user-friendly interface will make complex data accessible and understandable for non-domain experts, aiding in decision-making for MSR waste management. Our proposed visualization workflow can be applied to other nuclear reactors, assisting in the design and planning of waste management strategies.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Soil and groundwater environmental sensor data, Wax Lake Delta, Louisiana, March 2023 - March 2024

This study evaluates how environmental parameters that integrate biogeochemical processes vary with water table fluctuations in the freshwater Wax Lake Delta (WLD) in Louisiana, U.S.A. This data package contains seven *.csv files and one Excel file that compiles all the data from the individual .csv files. This dataset reports high frequency (15-min) observations of water level, soil redox potential, specific conductance, and pH made for one year along elevation transects located on the older, proximal (OT) and younger, distal (YT) ends of a deltaic island. Water depth relative to the ground surface (cm; HOBO U20L-04; error ± 0.4 cm), water pH and temperature (HOBO MX2501), and specific conductance and temperature (HOBO U24-001) sensors were installed in March 2023. Water depth was corrected for barometric pressure recorded by a separate logger secured to a platform above the highest water level. Soil redox probes (SWAP ORP-40-4-B) were also installed in March 2023. Each probe had four Pt sensors (2 mm width) placed at 10 cm, 20 cm, 30 cm, and 40 cm below the ground surface. Redox data were referenced to an external Ag0/AgCl (3M KCl) reference probe placed in saturated ground and recorded on CR1000X dataloggers (Campbell Scientific) powered by solar panels. A second reference probe was positioned near the primary reference probe for backup and data correction. The tops of the soil redox probes and soil moisture probes were flush with the soil surface so that sensors are reported at their indicated depths below ground surface. Here, we report data collected between 15 March 2023 to 15 March 2024 for all sensors, with some differences due to exact dates of sensor placement or data gaps associated with sensor malfunction. For example, water depth at OT4 was not recorded between March to November 2023. Data flags indicate whether a value is valid (1) or was excluded from data analysis in the associated manuscript (-1).

EARTH SCIENCE > LAND SURFACE > SOILS↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗