Search NASA⌕ Search

SEARCH · Search NASA

Results for “million core hours”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The ePIC Simulation Campaign Workflow on the Open Science Grid

The ePIC collaboration is realizing the first experiment of the future Electron-Ion Collider (EIC) at the Brookhaven National Laboratory that will allow for a precision study of the nucleons and the nucleus at the scale of sea quarks and gluons through the study of electron-proton/ion collisions. This paper will discuss the current workflow for running centralized simulation campaigns for ePIC on the Open Science Grid (OSG) infrastructure. This involves monthly releases of ePIC software and container deployments to CVMFS, generation of input datasets in HepMC format according to collaboration-defined policy, using Snakemake in CI/CD for validation and benchmarking, and submitting jobs to the OSG condor scheduler for opportunistic running on available resources. File transfers utilize XrootD, and Rucio is used for data management. The workflow is continuously refined to improve daily throughput (currently 50-100k core hours per day) and minimize job failures. Since May 2023, monthly simulation campaigns employing the workflow have cumulatively used over 20 million core hours on the OSG and produced over 350 TB of simulation data. The campaigns incorporate simulations for the broad science program of the EIC and are actively used for the detector and physics studies in preparation of the Technical Design Report (TDR).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

HEPCloud Operations at Fermilab—The First Five Years

The HEPCloud Facility at Fermilab has now been in production operation for five years. This facility is a unified provisioning gateway to US high performance computing centers, including NERSC, OLCF, and ALCF, other large supercomputers run by the NSF, and commercial clouds. HEPCloud delivers hundreds of millions of core-hours yearly for CMS. HEPCloud also serves other Fermilab experiments including DUNE, Mu2e, Muon g-2, and NOvA. In this paper we present the practical considerations of operating a distributed facility such as HEPCloud. We also mention some of the interesting research and development that HEPCloud has been used for including GPU-based machine learning inference servers, and tests of Quantum Computing.

Timm, Steven [Fermilab]↗

Accelerating multiscale electronic stopping power predictions with time-dependent density functional theory and machine learning

Knowing the rate at which particle radiation releases energy in a material, the “stopping power,” is key to designing nuclear reactors, medical treatments, semiconductor and quantum materials, and many other technologies. While the nuclear contribution to stopping power, i.e., elastic scattering between atoms, is well understood in the literature, the route for gathering data on the electronic contribution has for decades remained costly and reliant on many simplifying assumptions, including that materials are isotropic. We establish a method that combines time-dependent density functional theory (TDDFT) and machine learning to reduce the time to assess new materials to hours on a supercomputer and provide valuable data on how atomic details influence electronic stopping. Our approach uses TDDFT to compute the electronic stopping from first principles in several directions and then machine learning to interpolate to other directions at a cost of 10 million times fewer core-hours. We demonstrate the combined approach in a study of proton irradiation in aluminum and employ it to predict how the depth of maximum energy deposition, the “Bragg Peak,” varies depending on the incident angle—a quantity otherwise inaccessible to modelers and far outside the scales of quantum mechanical simulations. The lack of any experimental information requirement makes our method applicable to most materials, and its speed makes it a prime candidate for enabling quantum-to-continuum models of radiation damage. The prospect of reusing valuable TDDFT data for training the model makes our approach appealing for applications in the age of materials data science.

36 MATERIALS SCIENCE↗

Dragonfly Rotor Optimization using Machine Learning Applied to an OVERFLOW Generated Airfoil Database

NASA’s 4th New Frontiers Mission is the Titan Dragonfly relocatable lander. This coaxial quadrotor vehicle will be launched on a rocket to Titan in 2028. Following a gravity assisted Earth flyby and an approximate 6-year transit, Dragonfly will enter the Titan atmosphere around 2034 with the goal of exploring Titan’s pre-biotic chemistry and habitability. The multirotor design for this unique application has continually evolved since 2016 with constraints such as Titan’s cryogenic atmosphere at 95 Kelvin (-288 F), gravity 14% that of Earth’s, atmospheric density 440% of standard sea-level air, and the inability to test the entire system together under all these conditions until the first flight on Titan. This paper focuses on rotor design aspects of the Dragonfly lander and introduces a novel framework for multirotor design optimization considering multiple flight conditions. The methodology leverages machine learning methods and is demonstrated in the context of Dragonfly. A new OVERFLOW Machine Learning Airfoil Performance (PALMO) database is first presented. PALMO is then wrapped inside a Bayesian optimization framework and applied to a 4-rotor system (one side of the Dragonfly lander). Training data is generated on each iteration of the optimization using the CAMRAD-II comprehensive analysis software to evaluate successive rotor designs in multiple relevant flight conditions. An optimal design for the 4-rotor system was found with approximately 900 rotor designs analyzed in CAMRAD-II, which required 9 million queries of the PALMO surrogate models. This demonstration case evaluated 10,000,000 potential candidate rotor designs in 5.5 hours on 114 CPU cores using uniform inflow, and in 27.8 hours using the prescribed wake model. This work thus enables mid-fidelity rotor design optimization without requiring access to high-performance computing.

Dragonfly↗

The Global Ecosystem Dynamics Investigation (GEDI) Lidar Laser Transmitter

The Global Ecosystems Dynamics Investigation (GEDI) Lidar, is an Earth Science remote sensing instrument aboard the International Space Station (ISS) and the Japanese Experiment Module (JEM). Its core mission is to measure the global carbon balance of Earth's forests by using a set of three solid state laser transmitters in a multibeam waveform capture lidar technique. GEDI's laser transmitters and precision optical system transmits over 3.4 million laser pulses to the Earth every hour, each pulse producing an individual 3-D biomass column measurement. To enable a successful two-year mission, the lasers had to be reliable, highly repeatable in performance with each measurement power cycle, and designed with minimal part count for reduced manufacture complexity and cost. These transmitters are in-house products; developed, constructed, qualified, and fully integrated into the GEDI instrument at NASA's Goddard Space Flight Center. We will present the lasers' path from initial design to flight operation, with emphasis on the major milestones, critical issues, and lessons learned. Full credit goes to the excellent team effort that led to the successful commissioning and initiation of full-time science operations in March 2019.

Coyle, D. Barry↗

A New Approach to using a Cloud-Resolving Model to Study the Interactions between Clouds, Precipitation and Aerosols

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 kilometers, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-kilometer scales are resolved in horizontal domains as large as 10,000 kilometers in two dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically-based parameterizations for climate models and numerical weather prediction models. A review of developments and applications of cloud models in the past, present and future will be presented in this talk. In particular, a new approach to using cloud-resolving models to study the interactions between clouds, precipitation and aerosols will be presented.

Tao, Wei-Kuo↗

A New Approach to Using a Cloud-resolving Model to Study the Interactions Between Clouds, Precipitation and Aerosols

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 km, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-km scales are resolved in horizontal domains as large as l0,OOO km in two-dimensions, and 1,OOO x 1,OOO km2 in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically-based parameterizations for climate models and numerical weather prediction models. A review of developments and applications of cloud models in the past, present and future will be presented in this talk. In particular, a new approach to using cloud-resolving models to study the interactions between clouds, precipitation and aerosols will be presented.

Tao, Wei-Kuo↗

A New Approach to using a Cloud-Resolving Model to Study the Interactions between Clouds, Precipitation and Aerosols

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 km, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-km scales are resolved in horizontal domains as large as 10,000 km in two-dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically-based parameterizations for climate models and numerical weather prediction models. A review of developments and applications of cloud models in the past, present and future will be presented in this talk. In particular, a new approach to using cloud-resolving models to study the interactions between clouds, precipitation and aerosols will be presented.

Tao, Wei-Kuo↗

Using Multi-scale Modeling System to Study the Interactions between Clouds, Precipitation, Aerosols, Radiation and Land Surface

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 kilometers, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-kilometer scales are resolved in horizontal domains as large as 10,000 kilometers in two-dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically based parameterizations for climate models and numerical weather prediction models. It is also expected that NWP and mesoscale model can be run in grid size similar to cloud resolving model through nesting technique. A review of developments, improvements and applications of cloud models (GCE and WRF) at Goddard will be presented in this talk. In particular, a new approach to using multi-scale modeling system to study the interactions between clouds, precipitation, aerosols and land will be presented.

Tao, Wei-Kuo↗

Using Multi-scale Modeling System to Study the Interactions between Clouds, Precipitation, Aerosols, Radiation and Land Surface

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 kilometers, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial-lkm scales are resolved in horizontal domains as large as 10,000 kilometers in two-dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically based parameterizations for climate models and numerical weather prediction models. It is also expected that NWP and mesoscale model can be run in grid size similar to cloud resolving model through nesting technique. A review of developments, improvements and applications of cloud models (GCE and WRF) at Goddard wlll be is presented in this talk. In particular, a new approach to using multi-scale modeling system to study the interactions between clouds, precipitation, aerosols and land will be presented.

Tao, Wei-Kuo↗

Using Multi-scale Modeling System to Study the Interactions between Clouds, Precipitation, Aerosols, Radiation and Land Surface

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 km, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-km scales are resolved in horizontal domains as large as 10,000 km in two-dimensions, and 1,000 x 1,000 sq km in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically based parameterizations for climate models and numerical weather prediction models. It is also expected that NWP and mesoscale model can be run in grid size similar to cloud resolving model through nesting technique. A review of developments, improvements and applications of cloud models (GCE and WRF) at Goddard will be presented in this talk. In particular, a new approach to using multi-scale modeling system to study the interactions between clouds, precipitation, aerosols and land will be presented.

Tao, Wei-Kuo↗

Optical Multi-Gas Monitor Technology Demonstration on the International Space Station

There are a variety of both portable and fixed gas monitors onboard the International Space Station (ISS). Devices range from rack-mounted mass spectrometers to hand-held electrochemical sensors. An optical Multi-Gas Monitor has been developed as an ISS Technology Demonstration to evaluate long-term continuous measurement of 4 gases. Based on tunable diode laser spectroscopy, this technology offers unprecedented selectivity, concentration range, precision, and calibration stability. The monitor utilizes the combination of high performance laser absorption spectroscopy with a rugged optical path length enhancement cell that is nearly impossible to misalign. The enhancement cell serves simultaneously as the measurement sampling cell for multiple laser channels operating within a common measurement volume. Four laser diode based detection channels allow quantitative determination of ISS cabin concentrations of water vapor (humidity), carbon dioxide, ammonia and oxygen. Each channel utilizes a separate vertical cavity surface emitting laser (VCSEL) at a different wavelength. In addition to measuring major air constituents in their relevant ranges, the multiple gas monitor provides real time quantitative gaseous ammonia measurements between 5 and 20,000 parts-per-million (ppm). A small ventilation fan draws air with no pumps or valves into the enclosure in which analysis occurs. Power draw is only about 3 W from USB sources when installed in Nanoracks or when connected to 28V source from any EXPRESS rack interface. Internal battery power can run the sensor for over 20 hours during portable operation. The sensor is controlled digitally with an FPGA/microcontroller architecture that stores data internally while displaying running average measurements on an LCD screen and interfacing with the rack or laptop via USB. Design, construction and certification of the Multi-Gas Monitor were a joint effort between Vista Photonics, Nanoracks and NASA-Johnson Space Center (JSC). Vista Photonics developed the core technology and built the sensor. Nanoracks designed, constructed the enclosure, interfaces, and battery power management circuitry, integrated all subsystems into the enclosure, and then managed the certification tests, documentation and manifesting. The unit was calibrated in the JSC Toxicology Laboratory. The Multi-Gas Monitor is manifested to fly as a technology demonstration to the ISS in November 2013 and will operate for at least 6 months with data sent to the ground for evaluation. The primary goal is to demonstrate long term interference free operation in the real spacecraft environment.

Pilgrim, Jeffrey S.↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Development of Machine Learning Algorithms to Segment and Study Images of Astromaterial Samples

Introduction: Micrometer-scale chemical analyses of chondritic meteorites and mission-returned asteroid samples can reveal details of the physical and chemical processes operating in the early solar system, including processes that gave rise to planets, moons, and minor bodies. These primitive astromaterials are comprised of chondrules, calcium- and aluminum-rich inclusions (CAI), and many other silicates, oxides, metals, sulfides, and fine-grained materials. The chemical and mineralogical complexity of these samples, vast populations of different components, and heterogeneity across mm to km scales, all limit our understanding of the origin and evolution of these materials. Here, we describe recent efforts to use machine learning techniques to automate the segmentation of chemical maps of chondritic meteorites, designed to aid studies of asteroid samples returned by spacecraft. By automating the task of segmentation it will become possible to rapidly analyze and interpret the sizes, shapes, mineralogy, chemistry, and other properties of every chondrule, calcium- and aluminum-rich inclusion (CAI) and other clast within and between asteroid samples. Sample return missions significantly accelerate and heighten the need to develop such new data analysis techniques, and associated data repositories. Techniques: Neural networks require abundant training data, i.e. images which have been segmented by a human user. We have manually segmented data available from previous petrologic and chemical work at NASA Johnson Space Center and the American Museum of Natural History [1-4]. These data were derived from energy- and wavelength-dispersive X-ray spectroscopy (EDS, WDS) mapping of samples from many chondrite groups. The Deeplabv3+ [5] neural network architecture was trained on human-labeled masks and used to create machine-labeled masks. Several different algorithms were investigated, with inputs ranging from common RGB image formats through to hyperspectral datasets, with raw data comprising greyscale maps of Mg, Ca, and Al, with or without Si, Fe, Ti for both EDS and WDS data, and extending to other elements in EDS only. Each greyscale image was paired with a binary mask for each labelled particle type. Results: The trained algorithms can segment (Fig 1), classify, and measure the dimensions of thousands of particles in chemical maps of a standard 1-inch round petrographic section in seconds to minutes, rather than many hours needed by a human. Accuracy of the algorithms varied from chondrite to chondrite and across particle types. Further results and details of the algorithms will be presented at the workshop. Future directions: Machine learning has the potential to revolutionize our understanding of complex particle populations contained within primitive astromaterial, with segmentation being a critical first step. Example applications include better understanding of particle transport, nebular reservoirs, parent body accretion, and a deeper understanding of the relationships between particle populations and bulk rock elemental and isotopic compositions. In addition to benefits that machine learning can bring to individual researchers, building a community data repository of thousands to millions of particles across hundreds of samples will open up many other possibilities. For example, with a large enough dataset it will be possible to search for exceptionally closely matching particles across disparate samples. Such a capability would enable a single CAI from OSIRISREx or Hayabusa/II samples to be matched to chondritic CAIs that exhibit near-identical size, texture, and mineralogy, down to the level of similar core phenocrysts, zonation, and rim sequences. Such comparative analyses will help to disentangle precursor chemistry, chronology, gas/dust reservoirs during heating, and accretion. Such an endeavor would be impossible without machine learning and a large community data repository of astromaterial chemical/mineralogic maps.

Machine Learning↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗