Search NASASearch

SEARCH · Search NASA

Results for “Training Principles”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Next Gen One Portal Usability Evaluation

Each exercise device on the International Space Station (ISS) has a unique, customized software system interface with unique layouts / hierarchy, and operational principles that require significant crew training. Furthermore, the software programs are not adaptable and provide no real-time feedback or motivation to enhance the exercise experience and/or prevent injuries. Additionally, the graphical user interfaces (GUI) of these systems present information through multiple layers resulting in difficulty navigating to the desired screens and functions. These limitations of current exercise device GUI's lead to increased crew time spent on initiating, loading, performing exercises, logging data and exiting the system. To address these limitations a Next Generation One Portal (NextGen One Portal) Crew Countermeasure System (CMS) was developed, which utilizes the latest industry guidelines in GUI designs to provide an intuitive ease of use approach (i.e., 80% of the functionality gained within 5-10 minutes of initial use without/limited formal training required). This is accomplished by providing a consistent interface using common software to reduce crew training, increase efficiency & user satisfaction while also reducing development & maintenance costs. Results from the usability evaluations showed the NextGen One Portal UI having greater efficiency, learnability, memorability, usability and overall user experience than the current Advanced Resistive Exercise Device (ARED) UI used by astronauts on ISS. Specifically, the design of the One-Portal UI as an app interface similar to those found on the Apple and Google's App Store, assisted many of the participants in grasping the concepts of the interface with minimum training. Although the NextGen One-Portal UI was shown to be an overall better interface, observations by the test facilitators noted specific exercise tasks appeared to have a significant impact on the NextGen One-Portal UI efficiency. Future updates to the NextGen One Portal UI will address these inefficiencies.

Cross, E. V., III

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE

Mission X: Train Like an Astronaut Pilot Study

Mission X: Train Like an Astronaut is an international educational challenge focusing on fitness and nutrition as we encourage students to "train like an astronaut." Teams of students (aged 8-12) learn principles of healthy eating and exercise, compete for points by finishing training modules, and get excited about their future as "fit explorers." The 18 core exercises (targeting strength, endurance, coordination, balance, spatial awareness, and more) involve the same types of skills that astronauts learn in their training and use in spaceflight. This first-of-its-kind cooperative outreach program has allowed 14 space agencies and various partner institutions to work together to address quality health/fitness education, challenge students to be more physically active, increase awareness of the importance of lifelong health and fitness, teach students how fitness plays a vital role in human performance for exploration, and inspire and motivate students to pursue careers in STEM fields. The project was initiated in 2009 in response to a request by the International Space Life Sciences Working Group. USA, Netherlands, Italy, France, Germany, Austria, Colombia, Spain, and United Kingdom hosted teams for the pilot this past spring, and Japan held a modified version of the challenge. Several more agencies provided input into the preparations. Competing on 131 teams, more than 3700 students from 40 cities worldwide participated in the first round of Mission X. OUTCOMES AND BEST PRACTICES Members of the Mission X core team will highlight the outcomes of this international educational outreach pilot project, show video highlights of the challenge, provide the working group s initial assessment of the project and discuss the future potential of the effort. The team will also discuss ideas and best practices for international partnership in education outreach efforts from various agency perspectives and experiences

Lloyd, Charles W.

Helicopter Human Factors

Even under optimal conditions, helicopter flight is a most demanding form of human-machine interaction, imposing continuous manual, visual, communications, and mental demands on pilots. It is made even more challenging by small margins for error created by the close proximity of terrain in NOE flight and missions flown at night and in low visibility. Although technology advances have satisfied some current and proposed requirements, hardware solutions alone are not sufficient to ensure acceptable system performance and pilot workload. However, human factors data needed to improve the design and use of helicopters lag behind advances in sensor, display, and control technology. Thus, it is difficult for designers to consider human capabilities and limitations when making design decisions. This results in costly accidents, design mistakes, unrealistic mission requirements, excessive training costs, and challenge human adaptability. NASA, in collaboration with DOD, industry, and academia, has initiated a program of research to develop scientific data bases and design principles to improve the pilot/vehicle interface, optimize training time and cost, and maintain pilot workload and system performance at an acceptable level. Work performed at Ames, and by other research laboratories, will be reviewed to summarize the most critical helicopter human factors problems and the results of research that has been performed to: (1) Quantify/model pilots use of visual cues for vehicle control; (2) Improve pilots' performance with helmet displays of thermal imagery and night vision goggles for situation awareness and vehicle control; (3) Model the processes by which pilots encode maps and compare them to the visual scene to develop perceptually and cognitively compatible electronic map formats; (4) Evaluate the use of spatially localized auditory displays for geographical orientation, target localization, radio frequency separation; (5) Develop and flight test control/display concepts; (6) Quantify, model, predict, and improve pilots, workload-management strategies; and (7) Design computer-game trainers to reduce training time and cost.

Hart, Sandra G.

Portable Presentation And Instruction Unit

Proposed electronic display unit reminiscent of kiosk serves as portable, interactive, multimedia information terminal. Used as traveling science exhibit, aid for teaching science in schools, or training and skill-refresher device for space flight crews. Provides interactive video and audio displays, including three-dimensional-appearing video simulations. Speeds learning and improves retention by applying principles of scientific visualization. Also helps previously trained but recently unpracticed personnel relearn special skills and procedures quickly.

Christman, L.

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science

An Educational Program on Concentrated Solar Power and Heliostats for Power Generation and Industrial Processes

The objective of this project was to design and implement a comprehensive educational and applied research program in Concentrated Solar Thermal Power (CSTP) and heliostat technologies at Northeastern University. In alignment with the U.S. Department of Energy's Heliostat Consortium (HelioCon) goals, the project aimed to expand student and public understanding of CSTP systems while simultaneously contributing to workforce development and the broader decarbonization strategy. A particular emphasis was placed on integrating hands-on student design projects and publicly disseminating educational content relevant to CSTP systems. The project addressed a critical gap in renewable energy education: CSTP and heliostats, despite their importance in utility-scale solar energy, are rarely included in standard mechanical engineering programs. This project established new pathways for students to engage with the topic through the creation of a 4-credit graduate/senior elective course, development of five industry-facing short courses, and the inclusion of CSTP-based capstone design projects. Over two academic years, 36 students across six senior design teams developed and tested technologies such as deformable heliostats, beacon-based tracking systems, and solar-powered pyrolizers for biomass-to-biochar conversion. Concurrently, 30 undergraduate and graduate students were enrolled in the new academic course centered around CSTP principles. To ensure the relevance and accessibility of the short course content, the project team engaged with industry professionals, technical policy stakeholders, and potential course participants through structured surveys and informal consultations. Feedback from 28 respondents guided the structure, length, and delivery format of the courses - resulting in a modular design broken into five workshops. The feedback emphasized the need for flexible, asynchronous delivery and practical case studies, particularly in areas such as heliostat control, thermal storage, and solar fuel production. This engagement helped align the courses with the evolving knowledge demands of the renewable energy workforce and ensured that participants from both technical and policy backgrounds could meaningfully benefit from the material. The research and educational activities advanced the understanding of heliostat control systems, optical performance under misalignment, and thermal system integration in solar-driven pyrolysis applications. Methods and designs explored in this project proved to be both technically effective and economically feasible at the lab scale. Prototypes were constructed using commercially available components and custom-fabricated elements, demonstrating that meaningful performance improvements can be achieved with modest material and fabrication costs, supporting the feasibility of student-led research in this field. The public benefit of this project is twofold. First, it cultivates a pipeline of engineers trained to be familiar with CSTP principles, an essential workforce need identified by the Department of Energy for achieving its 2030 cost and deployment targets. Second, it contributes openly accessible educational materials, course content, and experimental frameworks to the broader community, enabling other institutions to adopt or adapt similar programming. Through outreach activities, curriculum integration, and technical exposure, this project contributes to a more informed and capable renewable energy workforce while supporting innovation in heliostat and CSTP system design. A new technical report is being prepared to document the development of the course and its outcomes, with plans to publish it in the ASME Open Access Journal of Engineering to ensure global accessibility, free of cost.

14 SOLAR ENERGY

Gradient-based optimization of complex nanoparticle heterostructures enabled by deep learning on heterogeneous graphs

Applications of deep learning (DL) to design nanomaterials are hampered by a lack of suitable data representations and training data. Here, in this study, we report efforts to overcome these limitations and leverage DL to optimize the nonlinear optical properties of core–shell upconverting nanoparticles (UCNPs). UCNPs, which have applications in fields such as biosensing, super-resolution microscopy and three-dimensional printing, can emit visible and ultraviolet light from near-infrared excitations. We report a large-scale dataset of UCNP emission spectra based on accurate but expensive kinetic Monte Carlo simulations (N > 6,000) and use these data to train a heterogeneous graph neural network using a physically motivated representation of UCNP nanostructure. Applying gradient-based optimization on the trained graph neural network, we identify structures with 6.5× higher predicted emission under 800-nm illumination than any UCNP in our training set. Our work reveals design principles for UCNP heterostructures and presents a roadmap for DL-based inverse design of nanomaterials.

Sivonxay, Eric [Lawrence Berkeley National Laborat

Crew/Automation Interaction in Space Transportation Systems: Lessons Learned from the Glass Cockpit

The progressive integration of automation technologies in commercial transport aircraft flight decks - the 'glass cockpit' - has had a major, and generally positive, impact on flight crew operations. Flight deck automation has provided significant benefits, such as economic efficiency, increased precision and safety, and enhanced functionality within the crew interface. These enhancements, however, may have been accrued at a price, such as complexity added to crew/automation interaction that has been implicated in a number of aircraft incidents and accidents. This report briefly describes 'glass cockpit' evolution. Some relevant aircraft accidents and incidents are described, followed by a more detailed description of human/automation issues and problems (e.g., crew error, monitoring, modes, command authority, crew coordination, workload, and training). This paper concludes with example principles and guidelines for considering 'glass cockpit' human/automation integration within space transportation systems.

Rudisill, Marianne

Preparing for Lunar Exploration: Geology and Field Training for Astronauts

NASA astronauts will soon return to the Moon, this time to the lunar south pole. Astronauts will perform traverses, make geologic observations, and collect samples to return to Earth. Lessons from Apollo show that science returns were optimized because crews were well-trained in both spacewalk operations and field geology. In that spirit, our team of geologists introduced a revised geologic training program for in-coming NASA astronauts that includes classroom activities and fieldwork and is split over 2 years. Year 1 focuses on an introduction to geologic concepts capped by a field exercise. Year two focuses on climate change, the Moon, other planets, and additional fieldwork. Our geology curriculum supports astronaut observations from the ISS and is the foundation for future Artemis geology training, which will include lunar science and increasingly complex field training in planetary-relevant locations. A final piece of our training program provides geology and field mapping experiences for NASA engineers, flight controllers, and managers, to help them understand the principles of fieldwork and relevance to lunar exploration. The Year 1 geology training must be an effective introduction for astronauts who have little or no background in geology. We focus the training around a narrowly defined field problem and use a week of classroom training to prepare the astronauts with the specific skills and background they need to carry out a field exercise on volcanic features, structures, and landforms of the Taos Plateau, New Mexico. Classroom modules are hands-on, using satellite imagery, maps, analog models, and samples, and provide the astronauts with specific and relevant experience in how to make observations and describe what they see, recognize, and interpret patterns, infer processes from products, and analyze relationships to build a story based on evidence from a variety of data sources. At the end of each classroom day, astronauts work in teams to build a preliminary geologic map of the field area, adding new observations and interpretations as they learn about topics in the classroom. Each team’s bucket list of target areas to visit helps shape their investigations during their week in the field. On the final day in the field, each astronaut team presents a geologic map, cross section, and geologic interpretation

Geology

Behavioral and biological interactions with small groups in confined microsocieties

Research on small group performance in confined microsocieties was focused upon the development of principles and procedures relevant to the selection and training of space mission personnel, upon the investigation of behavioral programming, preventive monitoring and corrective procedures to enhance space mission performance effectiveness, and upon the evaluation of behavioral and physiological countermeasures to the potentially disruptive effects of unfamiliar and stressful environments. An experimental microsociety environment was designed and developed for continuous residence of human volunteers over extended time periods. Studies were then undertaken to analyze experimentally: (1) conditions that sustain group cohesion and productivity and that prevent social fragmentation and performance deterioration, (2) motivational effects performance requirements, and (3) behavioral and physiological effects resulting from changes in group size and composition. The results show that both individual and group productivity can be enhanced under such conditions by the direct application of contingency management principles to designated high-value tasks. Similarly, group cohesiveness can be promoted and individual social isolation and/or alienation prevented by the application of contingency management principles to social interaction segments of the program.

Brady, Joseph V.

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields in the last two decades, in part thanks to an increasing culture of open data sharing and reuse. Due to its capability for identifying complex relationships and patterns, AI/ML methodology is particularly well suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are many key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Even with the positive culture of Open Science and data sharing, inexperienced researchers working quickly without proper checks can produce models that perform poorly outside of the immediate training dataset. Lessons learned from biological AI/ML research indicate that Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Andrew Casaletto

Uncertainty quantification of a physics-informed model based on sparse identification of a Thermal Energy Distribution System

Integrated energy systems (IES)s are crucial for enhancing the economy and efficiency of power generation sources (e.g., nuclear energy) necessary to unleash American energy dominance. These systems can be integrated with thermal energy storage (TES) and intermittent renewable energies to optimize overall energy use, peak-load regulation, and demand-side responses. However, the stabilization of energy generation, transport, and utilization introduces operational complexities that exceed the challenges of managing each sub-component individually. Currently, though IESs rely on human operators for efficiency and stability, reducing human error risk and enhancing performance through automation is highly desirable. Recent advances at Idaho National Laboratory have demonstrated successful control of the Thermal Energy Distributed System (TEDS). However, the automatic control system depends on a deterministic Sparse Identification of Nonlinear Dynamics with Control (SINDyC) model, which are trained based on simulation data from physics-based simulations. Because of uncertainties in physics-based simulation, SINDyC model results in large discrepancies against experimental data and cannot be reliably used in automatic control. In this paper, we present an innovative approach to address these discrepancies by quantifying uncertainties and developing a more robust model. We first generated trajectories by using first-principles physics codes to encapsulate the experiment. Next, we trained thousands of models by randomly sampling these trajectories. We then collapsed all those models into one probabilistic SINDyC by fitting a multivariate Gaussian distribution onto the resulting coefficient’s distribution. Despite its simplicity, our approach successfully produced 95% confidence intervals that captured the experimental trajectories. It even did so with a higher probability and better U-pooling score across six of the seven relevant quantities of interest (QoIs), as compared to other classical approaches. In conclusion, ongoing research is focusing on generating new experimental trajectories to validate this approach, and on employing Bayesian calibration to refine parametric uncertainties and guide future model development efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Machine learning approach for vibronically renormalized electronic band structures

Here, we present a machine learning (ML) method for efficient computation of vibrational thermal expectation values of physical properties from first principles. Our approach is based on the nonperturbative frozen phonon formulation in which stochastic Monte Carlo algorithm is employed to sample configurations of nuclei in a supercell at finite temperatures based on a first-principles phonon model. A deep-learning neural network is trained to accurately predict physical properties associated with sampled phonon configurations, thus bypassing the time-consuming ab initio calculations. To incorporate the point-group symmetry of the electronic system into the ML model, group-theoretical methods are used to develop a symmetry-invariant descriptor for phonon configurations in the supercell. We apply our ML approach to compute the temperature dependent electronic energy gap of silicon based on density functional theory (DFT). We show that, with less than a hundred DFT calculations for training the neural network model, an order of magnitude larger number of sampling can be achieved for the computation of the vibrational thermal expectation values. Our work highlights the promising potential of ML techniques for finite temperature first-principles electronic structure methods.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING

Anchoring

This software provides methods and functions for training deep image classification models based on the principle of anchoring. It features a user-friendly PyTorch wrapper that facilitates the easy conversion of any model into an anchored model. The software supports various standard datasets and includes scripts for conducting evaluations. Developed with PyTorch, it is compatible with common neural network architectures used for image data. Additionally, it offers tools for computing evaluation metrics to assess model performance.

Narayanaswamy, Vivek Sivaraman