Search NASA⌕ Search

SEARCH · Search NASA

Results for “FAIR Digital Object”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Challenges for Implementing FAIR Digital Objects with High Performance Workflows

New types of workflows are being used in science that couple traditional distributed and high-performance computing (HPC) with data-intensive approaches, and orchestrate ensembles of numerical simulations and artificial intelligence (AI) models. Such workflows may use AI models to supplement computation where numerical simulations may be too computationally expensive, to automate trivial yet time consuming operations, to perform preliminary selections among intractable numbers of combinations in domains as diverse as protein binding, fine-grid climate simulations, and drug discovery.

97 MATHEMATICS AND COMPUTING↗

Making digital objects FAIR in high energy physics: An implementation for Universal FeynRules Output (UFO) models

Research in the data-intensive discipline of high energy physics (HEP) often relies on domain-specific digital contents. Reproducibility of research relies on proper preservation of these digital objects. This paper reflects on the interpretation of principles of Findability, Accessibility, Interoperability, and Reusability (FAIR) in such context and demonstrates its implementation by describing the development of an end-to-end support infrastructure for preserving and accessing Universal FeynRules Output (UFO) models guided by the FAIR principles. UFO models are custom-made python libraries used by the HEP community for Monte Carlo simulation of collider physics events. Our framework provides simple but robust tools to preserve and access the UFO models and corresponding metadata in accordance with the FAIR principles.

Neubauer, Mark S.↗

F*** workflows: when parts of FAIR are missing

The FAIR principles for scientific data (Findable, Accessible, Interoperable, Reusable) are also relevant to other digital objects such as research software and scientific workflows that operate on scientific data. The FAIR principles can be applied to the data being handled by a scientific workflow as well as the processes, software, and other infrastructure which are necessary to specify and execute a workflow. The FAIR principles were designed as guidelines, rather than rules, that would allow for differences in standards for different communities and for different degrees of compliance. There are many practical considerations which impact the level of FAIR-ness that can actually be achieved, including policies, traditions, and technologies. Because of these considerations, obstacles are often encountered during the workflow lifecycle that trace directly to shortcomings in the implementation of the FAIR principles. Here, we detail some cases, without naming names, in which data and workflows were Findable but otherwise lacking in areas commonly needed and expected by modern FAIR methods, tools, and users. We describe how some of these problems, all of which were overcome successfully, have motivated us to push on systems and approaches for fully FAIR workflows.

Wilkinson, Sean↗

Applying the FAIR Principles to computational workflows

Recent trends within computational and data sciences show an increasing recognition and adoption of computational workflows as tools for productivity and reproducibility that also democratize access to platforms and processing know-how. As digital objects to be shared, discovered, and reused, computational workflows benefit from the FAIR principles, which stand for Findable, Accessible, Interoperable, and Reusable. The Workflows Community Initiative’s FAIR Workflows Working Group (WCI-FW), a global and open community of researchers and developers working with computational workflows across disciplines and domains, has systematically addressed the application of both FAIR data and software principles to computational workflows. We present recommendations with commentary that reflects our discussions and justifies our choices and adaptations. These are offered to workflow users and authors, workflow management system developers, and providers of workflow services as guidelines for adoption and fodder for discussion. The FAIR recommendations for workflows that we propose in this paper will maximize their value as research assets and facilitate their adoption by the wider community.

97 MATHEMATICS AND COMPUTING↗

Codebase release 2.0 for UFOManager

Research in the data-intensive discipline of high energy physics (HEP) often relies on domain-specific digital contents. Reproducibility of research relies on proper preservation of these digital objects. This paper reflects on the interpretation of principles of Findability, Accessibility, Interoperability, and Reusability (FAIR) in such context and demonstrates its implementation by describing the development of an end-to-end support infrastructure for preserving and accessing Universal FeynRules Output (UFO) models guided by the FAIR principles. UFO models are custom-made python libraries used by the HEP community for Monte Carlo simulation of collider physics events. Our framework provides simple but robust tools to preserve and access the UFO models and corresponding metadata in accordance with the FAIR principles.

Neubauer, Mark S.↗

KBase Credit Metadata Schema

As part of KBase’s commitment to promote open science, we offer users the ability to obtain a DOI (Digital Object Identifier) for their work, which can then be cited in an associated science publication. To further support the community-wide shift towards FAIR (Findable, Accessible, Interoperable, Reusable) data, KBase is expanding our data descriptors so that KBase DOIs have comprehensive citations for datasets, in addition to referencing publications or software used in the workflow. This helps encourage a culture of giving attribution for all research inputs and outputs; standard practice for literature, but still relatively new for software products or datasets. It also promotes open science by building trust that contributors get credit for their work, and accelerates knowledge discovery by supporting and incentivizing the release of data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Guidelines for Publicly Archiving Terrestrial Model Data to Enhance Usability, Intercomparison, and Synthesis

Scientific communities are increasingly publishing data to evaluate, accredit, and build on published research. However, guidelines for curating data for publication are sparse for model-related research, limiting the usability of archived simulation data. In particular, there are no established guidelines for archiving data related to terrestrial models that simulate land processes and their coupled interactions with climate. Terrestrial modelers have a unique set of challenges when publishing data due to the diversity of scientific domains, research questions, and the types and scales of simulations. Researchers in the U.S. Department of Energy’s (DOE) projects use a variety of multiscale models to advance robust predictions of terrestrial and subsurface ecosystem processes. Here, we synthesize archiving needs for data associated with different DOE models, and provide guidelines for publishing terrestrial model data components following FAIR (Findable, Accessible, Interoperable, Reusable) principles. The guidelines recommend archiving model inputs and testing data used in final simulation runs along with associated codes, workflow scripts, and metadata in public repositories. Researchers should consider archiving model outputs if they are within the storage limits of the repository. We also provide considerations for how to bundle files into different data publications with citable digital object identifiers. Finally, we identify repository features and tools that would enable storage and reuse of model data. Given the diversity of DOE terrestrial models, these guidelines are transferable to other model types and will enable efficient reuse of simulation data for purposes such as model intercomparisons, initialization, benchmarking, synthesis, and comparisons with field observations.

58 GEOSCIENCES↗

A review of privacy in energy applications

As the distribution system continues to experience an increase in distributed energy resource (DER) and electric vehicle (EV) penetration, so does the need for new solutions that can help grid operators manage and leverage their capabilities. This will undoubtedly lead to new operational schemes and business opportunities that will transform the traditional consumer into a prosumer who will be more actively engaged in grid operations. Although the field is still under active development, many of the potential use cases presented in literature or industry are built upon edge computing, two-way communications, and other innovative computational constructs to attain their goals. However, at their core, many use cases assume a great level of data access to aid with the decision-making process, an assumption that may need to be revised to ensure fair and equitable operational processes are maintained. This may be particularly true as edge resources are predicted to participate in retail-side, many-to-many, or peer-to-peer markets and thus may lead to financial impacts if data access considerations are ignored. The need to revise data access mechanisms can be further justified by the introduction of new participants into the operational process, who do not have the same level of trust, nor the incentives to focus on energy delivery as their primary objective. At the same time, more consumers are becoming aware of their own data, and the potential impacts of its abuse. To help solution developers better understand these risks, this report has been developed to offer an initial introduction to the topic of privacy. This is achieved by 1) Highlighting the need for privacy-aware solutions; 2) Encouraging system designers to be inquisitive about the status quo; 3) Documenting the existing threat space; 4) Presenting and evaluating tools that may be helpful towards enabling better privacy postures; and 5) Making recommendations to encourage the adoption of better practices. From a technical perspective, the report focuses on evaluating two potential techniques by applying them to the Transactive Energy Space. Based on the obtained results, it can be established that differential privacy (DP) methods may have limited applicability when highly correlated, time-series data records need to be protected. However, DP may be a powerful tool when it is used to aggregate and analyze mid-size and large-size data sets in a more traditional statistical environment. The second tool under evaluation is threshold cryptography, which can guarantee complete secrecy (and thus privacy) but requires the establishment of key management procedures and dedicated communication channels for key coordination. Therefore, due to its increased computational overhead, the use of threshold cryptography must be weighted using a cost/benefit analysis on a per-application basis.

24 POWER TRANSMISSION AND DISTRIBUTION↗