Search NASA⌕ Search

SEARCH · Search NASA

Results for “fair machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

50 records · Page 3

Overcoming sparse datasets with multi-task learning as applied to high entropy alloys

Abstract The design of novel High Entropy Alloys for use in high-temperature applications is an area of active interest due to their potential to provide exceptional properties compared to conventional alloys. Since the increased popularity of machine learning, an important cog in the design process has been training surrogate models on alloy properties. However, these Single-Task models are trained on individual mechanical properties and do not take advantage of the relatedness between properties. Multi-Task models can capture the interdependencies between tasks, leading to potentially more accurate predictions for all tasks. In this paper, we investigate if Multi-Task models can show improvement over Single-Task models when used for predicting the mechanical properties of these alloys. To ensure fair evaluation between the models, we apply L 0 regularization and skip connections to the models, which allows them to adjust the number of model parameters and depth for optimal performance. We find that the Multi-Task models can leverage task relationships to perform better than Single-Task models, especially for high amounts of missing data in the tasks. Furthermore, adding simple auxiliary targets can boost Multi-Task performance even further despite not being effective as input descriptors to single-task models themselves. We anticipate that the proposed strategies can achieve more accurate predictions and consequently enable better design capabilities for such data-constrained domains without incurring much additional computational cost.

Debnath, Arindam (ORCID:0000000194274499)↗

Using Artificial Intelligence and Machine Learning to Enhance Mission Design and Operations of the Habitable Worlds Observatory (HWO)

One key aspect in the development of HWO is the early deployment of artificial intelligence (AI) and machine learning (ML) to enhance mission science and operations. Our subtask group is part of the HWO AI/ML working group and focuses on AI and ML for mission operations. Our task group seeks to educate other HWO working groups about AI and ML capabilities for mission operations, investigate how to bridge technology gaps, and enable new capabilities particularly in the areas of observational scheduling, instrument health monitoring, and downlink operations. We focus on mission tasking / scheduling both for mission analysis in development and operations. AI and ML for mission scheduling includes: tools to support proposal calls and review, ensuring fairness in calls for proposals, community peer reviews and ease workloads, as well as in-flight and ground software development (e.g., using natural language processing (NLP) to support process automation from requirements). AI and ML for the mission’s development and operations include 1) anomaly detection and prediction (from onboard and ground based tools) to monitor the spacecraft’s health, 2) ground-based automated scheduling for mission operations including long-term and short-term planning and maintenance, and 3) flight system flexible execution (as flight proven for Spitzer and JWST) to enable robust execution despite execution variations, and 4) data analysis for prioritization (e.g., real-time data evaluation leading to autonomous actions and adjustments, high-priority identification, onboard data compression, etc.). Incorporation of ML and AI will enable HWO to address the major science questions related to exoplanet characterization, general astrophysics, and solar system exploration and also extend the boundaries of space mission technologies.

Mark Moussa↗

Data Readiness for AI: A 360-Degree Survey

Artificial Intelligence (AI) applications critically depend on data. Poor-quality data produces inaccurate and ineffective AI models that may lead to incorrect or unsafe use. Evaluation of data readiness is a crucial step in improving the quality and appropriateness of data usage for AI. R&D efforts have been spent on improving data quality. However, standardized metrics for evaluating data readiness for use in AI training are still evolving. In this study, we perform a comprehensive survey of metrics used to verify data readiness for AI training. This survey examines more than 140 papers published by ACM Digital Library, IEEE Xplore, journals such as Nature, Springer, and Science Direct, and online articles published by prominent AI experts. This survey aims to propose a taxonomy of data readiness for AI (DRAI) metrics for structured and unstructured datasets. We anticipate that this taxonomy will lead to new standards for DRAI metrics that would be used for enhancing the quality, accuracy, and fairness of AI training and inference.

97 MATHEMATICS AND COMPUTING↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

Discriminative versus generative approaches to simulation-based inference

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

high energy physics↗

First Estimation of Model Parameters for Neutrino-Induced Nucleon Knockout Using Simulation-Based Inference

To enable an accurate determination of oscillation parameters, accelerator-based neutrino experiments require detailed simulations of nuclear interaction physics in the GeV regime. While substantial effort from both theory and experiment is currently being invested to improve the fidelity of these simulations, their present deficiencies typically oblige experimental collaborations to resort to empirical tuning of simulation model parameters. As the precision requirements of the field continue to become more stringent, machine learning techniques may provide a powerful means of handling corresponding growth in the complexity of future neutrino interaction model tuning exercises. To study the suitability of simulation-based inference (SBI) for this physics application, in this paper we revisit a tuned configuration of the GENIE neutrino event generator that was originally developed by the MicroBooNE collaboration. Despite closely reproducing the adopted values of four physics parameters when confronted with the tuned cross-section predictions as input, we find that our trained SBI algorithm prefers modestly different values (within MicroBooNE's assigned uncertainties) and achieves slightly better goodness-of-fit when inference is run on the experimental data set originally used by MicroBooNE. We also find that our trained algorithm can create a fair approximation of an alternative neutrino scattering simulation, NuWro, that shares only a subset of its physics model parameters with GENIE.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

Supporting ARPA-E Power Grid Optimization (Final Report)

Pacific Northwest National Laboratory (PNNL), Arizona State University (ASU), Georgia Institute of Technology (Georgia Tech), Los Alamos National Laboratory (LANL), National Renewable Energy Laboratory (NREL), Texas A&M University (TAMU), The University of Texas at Austin (UT), and the University of Wisconsin-Madison (UW-M) supported the ARPA-E Grid Optimization (GO) Competition by providing a common problem formulation, data format, datasets, evaluation mechanism, scoring, rules, and results that resulted in the awarding of $\$9.24$ million dollars to teams from academia, industry, and national labs for solving three sets of increasingly difficult non-linear, security- constrained AC Optimal Powerflow (AC-OPF) optimization problems in order to increase the efficiency of the US Electric Grid. It is estimated that a 1% increase in efficiency can save $\$1$ billion. Current industry practices typically use a linear DC model (DC-OPF) in order solve the OPF problem within the time constraints of the operation schedule. The GO Competition challenges the best power engineers, mathematicians, and computer scientists to make possible operational decisions based on accurate physical models. To accomplish this, the GO Competition created a series of Challenges and funded teams to produce the best solver. Challenge 1 was to solve the security constrained Alternating Current Optimal Power Flow (ACOPF) problem. Challenge 2 extended that to by adding adjustable transformer tap ratios, phase shifting transformers, switchable shunts, price-responsive demand, ramp rate constrained generators and loads, and fast-start unit commitment (UC). Furthermore, Challenge 2 was a maximization problem while Challenge 1 was a minimization problem. While Challenge 3 was being developed, the entrants were invited to find better solutions to the Challenge 2 synthetic datasets with no restrictions on time, hardware, or algorithms. The Challenge 2 solutions turned out to be very good. Challenge 3 expanded the Challenge 2 problem further by using multiperiod dynamic markets, including advisory models for extreme weather events, day-ahead markets, and the real-time markets with an extended look-ahead. These problems included active bid-in demand and topology optimization. Together the Challenges used nearly 30 million CPU hours. Since each team was working on the same problem, using the same data, and running on the same hardware, fair comparisons could be drawn as to the best solver. The datasets were varied enough, however, that the best solver for one dataset was not necessarily the best at another, so cumulative scores were used. The process was managed by the PNNL maintained website https://GOCompetition.energy.gov, where Entrants could find information about the problem, the data, the rules, submit their solver for evaluation, and see the scores of all the competing teams on a Leaderboard. Interest was world-wide but only American teams were eligible for prizes. The Competition has produced 34 journal articles 115 papers and been cited over 500 times in the literature, including 12 dissertations (4 from foreign countries; Columbia (2), Germany, and Italy) and 3 from the DOE ExaScale project. Software developed by Pearl Street Technologies for Challenges 1 and 2 is now deployed by Southwest Power Pool (SPP) and Midcontinent Independent Service Operator (MISO). Other teams have received inquiries from venture capitalists. Google DeepMind has thanked the Competition for making the datasets developed for the Competition public. They are using it to train machine learning models. The larger datasets have billions of unknowns to be solved for, but only a small percent matter in the final solution. Knowing what unknowns are important can dramatically speedup the solution.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unsupervised Detection of SOC Spoofing in OCPP 2.0.1 EV Charging Communication Protocol Using One-Class SVM

The electric vehicles (EVs) market keeps growing globally; thus, it is critical to secure the EV charging communication protocols in order to guarantee reliable and fair charging operations among the customers. The Open Charge Point Protocol (OCPP) 2.0.1 supports the communication between the Electric Vehicle Supply Equipment (EVSE) and Charging Station Management Systems (CSMSs); therefore, it becomes vulnerable to several types of attacks, which aim to jeopardize smart charging, billing, and energy management. Specifically, OCPP 2.0.1 allows the self-reporting of the State of Charge (SOC) values, which makes it vulnerable to spoofing-based cyberattacks, which target manipulating the scheduling priorities, distorting the load forecasts, and extending the charging sessions in an unfair manner. In this paper, we try to address this type of attack by providing a comprehensive analysis of the SOC spoofing attacks and introducing a novel unsupervised detection framework based on the One-Class Support Vector Machine (OCSVM) algorithm. Specifically, two types of attack scenarios are analyzed (i.e., priority manipulation and session extension) by deriving engineered features that capture the nonlinear relationships under normal charging behavior. Detailed simulation-based results are derived by utilizing the DESL-EPFL Level 3 EV charging dataset. Our results demonstrate high F1-score and recall in identifying spoofed SOC values and that the proposed OCSVM model demonstrates superior performance compared to alternative clustering and deep-learning based detectors.

EV charging↗

A deep learning approach to fast analysis of collective Thomson scattering spectra

Fast analysis of collective Thomson scattering ion acoustic wave features using a deep convolutional neural network model is presented. The network was trained from spectra to predict the plasma parameters, including ion velocities, population fractions, and ion and electron temperatures. A fully kinetic particle-in-cell simulation was used to model a laboratory astrophysics experiment and simulate a diagnostic image of the ion acoustic wave feature. Network predictions were compared with Bayesian inference of the plasma model parameters for both the simulated and experimentally measured images. Both approaches were fairly accurate predicting the simulated image and the network predictions matched a good portion of the Bayesian results for the experimentally measured image. The Bayesian approach is more robust to noise and motivates future work to train deep learning models with realistic noise. The advantage of the deep learning model is making thousands of predictions in a few hundred milliseconds, compared to a few seconds to minutes per prediction for the optimization and Bayesian approaches presented here. The results demonstrate promising capabilities of deep learning models to analyze Thomson data orders of magnitude faster than conventional methods when using the neural network for standalone analysis. If more rigorous analysis is needed, neural network predictions can be used to quickly initialize other optimization methods and increase chances of success. This is especially useful when the dataset becomes very large or highly dimensional and manually refining initial conditions for the entire dataset are no longer tractable.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

An Ethics-Based Review of Generative Artificial Intelligence: Assuring Responsible Use (Version 1.0)

The rapid expansion of generative artificial intelligence (GenAI) has generated excitement regarding its potential benefits and concern over its ethical implications. Governments, corporations, and standards organizations have described ethical principles to direct GenAI's development and use; however, practical guidance for implementing these principles is limited. Addressing this gap is critical, especially considering the array of risks associated with GenAI, such as legal liabilities, privacy concerns, security threats, and potential misuse. Robust policies and procedures are critical to support responsible deployment of GenAI. This report examines Pacific Northwest National Laboratory (PNNL)’s approach to promoting responsible GenAI use. Proposed initiatives include developing policies based on ethical principles, creating a governance process to review projects relative to those principles, and implementing onboarding processes for training staff. The governance framework described in this report adapts the structure and principles of Institutional Review Boards (IRBs), traditionally used in human subjects research, for GenAI ethical review, providing oversight. Ethical principles guiding responsible GenAI usage include transparency and accountability, privacy, fairness, safety, security, and validity and reliability. To operationalize these principles, we propose forming a GenAI Assurance Council (GAC) that mirrors the IRB's structure. The GAC will evaluate GenAI projects across privacy, accountability, transparency, safety, security, fairness, and validity dimensions. Complementing policy and governance is AI literacy training to support staff understanding of GenAI's ethical implications. An initial training effort for AI Incubator Chat—a GenAI tool deployed at PNNL—showed promising results, underscoring the importance of clear guidelines and user accountability. Collaborative efforts and the dissemination of best practices are also discussed. The proposed GAC model and AI literacy training provide a blueprint for establishing ethical GenAI use and governance, offering practical tools to bridge the gap between ethical principles and real-world applications. The responsible integration of GenAI at PNNL entails a multifaceted approach involving policy development, ethical governance, and AI literacy training. The positive initial feedback and collaborative opportunities position PNNL to lead by example in GenAI's responsible use, reflecting a proactive stance in addressing the ethical, legal, and societal challenges associated with this emerging technology. PNNL's systematic and ethical approach to GenAI offers a model for other institutions to emulate, promoting safe and responsible technological advancements in the AI domain.

97 MATHEMATICS AND COMPUTING↗

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods.

knowledge↗