Search NASA⌕ Search

SEARCH · Search NASA

Results for “open source software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

800 records · Page 45

IM3 Projected US Data Center Locations

IM3 Projected US Data Center Locations This dataset contains model projections of new data center facilities in the contiguous United States (CONUS) through 2035 using the CERF – Data Centers model. Data center locations are modeled across four data center electricity demand growth scenarios (low, moderate, high, higher) and five market gravity scenarios (0%, 25%, 50%, 75%, 100%). Projected locations are intended to be regional representations of feasible siting locations in the future to assess potential grid and water stress impacts. The data center load growth scenarios correspond with the rates outlined in EPRI (2024) and include 3.71%, 5%, 10%, and 15% annual growth of electricity demand for data centers from 2023 values in 37 states across the CONUS. Market gravity scenarios correspond to the relative importance of proximity to data center markets or high population areas compared to locational cost in the siting algorithm. 0% market gravity means that siting decisions were entirely determined by the locational cost in each feasible location. 100% market gravity means that only market proximity was considered when siting. Other scenarios have weight placed on both components where total weight always equals 100%. Locational cost is dependent on facility cooling type and corresponding electricity cost, taxes, and other factors. Facility cooling type is spatially determined where high water stress and/or areas with high summer wet bulb temperatures are assumed to operate with mechanical cooling for a higher fraction of the year rather than evaporative cooling. Feasible data center siting areas are based on geospatial suitability raster data developed with open-source information. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands Because we use open-source information, proprietary information that can influence siting decisions such as individual tax agreements with cities, detailed fiber line connectivity, electric grid power capacity agreements, and others, are not currently accounted for in the modeling process. Using specific building locations and footprints in the dataset for local planning purposes is not advised. Technical Information Geospatial data is provided in geojson format using the Albers Equal Area Conic (ESRI:102003) coordinate reference system. The datasets contain the following parameters: id - unique identification number within given scenario file growth_scenario – data center demand growth scenario market_gravity_weight – market gravity weight scenario (%) region – name of region (i.e., US State) total_cost_million_usd – locational siting cost ($million) campus_size_square_ft – total land acquired for data center facility (square ft) data_center_it_power_mw – IT power of data center facility (MW) mechanical_cooling_frac – fraction of year when data center uses mechanical cooling system water_cooling_frac– fraction of year when data center uses evaporative cooling system cooling_energy_demand_mwh – total annual facility energy demand for cooling (MWh) cooling_water_demand_mgy – total annual facility water demand for cooling (MG) cooling_water_consumption_mgy – total annual facility water consumed (MG) normalized_locational_cost – normalized total locational cost score for location normalized_gravity_score – normalized market gravity score for location weighted_siting_score – total weighted siting score of locational cost and gravity score geometry – polygon geometry of facility Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall (ORCID:0000000328077088)↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Catalight─An Open-Source Automated Photocatalytic Reactor Package Illustrated through Plasmonic Acetylene Hydrogenation

An open-source and modular Python package, Catalight, is developed and demonstrated to automate (photo)catalysis measurements. (Photo)catalysis experiments require studying several parameters to evaluate performance, including the temperature, gas flow rate and composition, illumination power, and spectral profile. Catalight orchestrates measurements over this complicated parameter space and systematically stores, analyzes, and visualizes the results. To showcase the capabilities of Catalight, we perform an automated apparent activation barrier measurement of acetylene hydrogenation over a plasmonic AuPd catalyst on an Al 2 O 3 support, simultaneously varying laser power, wavelength, and temperature in a multiday experiment controlled by a simple Python script. Our chemical results unexpectedly show an increased activation barrier upon light excitation, contrary to previous findings for other plasmonic reactions and catalysts. We show that the reaction rate order with respect to both acetylene and hydrogen remains unchanged upon illumination, suggesting that molecular surface coverage is not changed by light. By analyzing the inhomogeneity of the laser-induced heating, we attribute these results to a partial photothermal effect combined with a photochemical/hot electron-driven mechanism. In conclusion, our findings highlight the capabilities of a new experiment automation tool; explore the photocatalytic mechanism for an industrially relevant reaction; and identify systematic sources of error in canonical photocatalysis experimental procedures.

Catalysts↗

Marine Algae Industrialization Consortium (MAGIC): Combining biofuel and high-value bioproducts to meet the RFS

The Marine Algae Industrialization Consortium (MAGIC) was formed to address pressing challenges in the commercialization of microalgae as a source of biofuel. The “Marine Algae Industrialization Consortium (MAGIC): Combining biofuel and high-value bioproducts to meet the RFS” project formally addressed two US Department of Energy Bioenergy Technologies Office (BETO) goals: (1) Model the sustainable supply of 1 million metric tonnes ash free dry weight (AFDW) cultivated algal biomass and (2) Demonstrate valuable co-products produced along with biofuel intermediates to increase value of algal biomass by 30%. To achieve these goals, the project demonstrated and validated high-value co-products to drive down the cost of biofuel by increasing the value of algae “co-products” towards increasing the selling price of total algae biomass as one of the key drivers of economics and adoption. This was accomplished through five core, interdependent tasks including: (1) strain selection to identify and deliver strains for mass culture, (2) mass culture using a hybrid cultivation system and following key operating parameters for downstream applications to provide algae feedstock, (3) recovery and conversion to evaluate two alternative methods to separate dry algae biomass into oil and residuals for downstream testing, (4) product assessment to determine biofuel, aquafeed or poultry feed product efficacy using algae biomass fractions as well as to provide critical performance data for valuation and (5) commercialization to use technoeconomic and life cycle assessments (TEA/LCA) as iterative design and assessment tools including consideration of target markets, competitors, and distribution channels to guide product assessment, development and valuation. A total of 46 peer-review publications, many open-access, provide detail of much of the work carried out and the results of the tasks. Additional reports and presentations provide other technical and public engagement material. At a high level, using a variety of approaches, more than 1000 marine microalgae strains were evaluated to ultimately identify the seven winners that were down-selected to be grown in mass culture. Strain selection demonstrated that there were no ‘super strains’ and that each candidate had strengths and limitations for specific products, growth conditions or operational considerations. Mass culture growth of these seven strains at >5000 L / 29 m 2 scale found that four them were suitable for product assessment. More than 250 kg of biomass was produced across hundreds of pond runs along with thousands of cultivation entries on the growth and biomass characteristics as well as environmental parameters. In the process, dozens of standard operating procedures were generated as was custom software to process and analyze cultivation data. Recovery and conversion of algae biomass demonstrated that a hexane solvent based extraction protocol was most effective at recovering oil (biocrude) from algae and four strains were processed to produce oil and lipid extracted algae (residuals) for downstream testing. Membrane-based oil separation was less successful, but may still be applicable to other commercial applications in the future. Product testing demonstrated that algae biocrude is of high quality and hydrotreating generated numerous fractions of high quality composition for fuel and lubricate based applications. Aquafeed studies performed at a variety of scales showed that both whole and defatted (lipid extracted algae) microalgae were suitable as a feed ingredient, but that the specifics of the fed animal and biochemical composition of the algae are critical factors when determining formulation. Similarly, poultry studies on whole and defatted microalgae generally showed positive outcomes on animal growth and health, with some microalgae providing enhanced nutritional composition of the animal product. Economic and life cycle assessments covered a wide range of possible commercialization and sustainability scenarios. Replacement value, improved product value added, consumer values marketing added valuation and improved animal health were considered as alternatives for microalgae valuation. Using the open pond system, algae productivity was identified as the key driver of commercialization economics, but combination of co-products (e.g. animal feed) with biofuel production substantially increased the total selling price of algae. Modeled microalgae selling price exceeded $\$$1500/tonne and could generate competitive biofuel selling prices below $\$$5 gallon gas equivalents using realistic algal productivities. Short (process scale) and longer (decadal trends) sustainability assessments show that marine microalgae can enhance the sustainability of energy production and lead to other realized benefits in water, fertilizer and land use for other sectors (e.g. agriculture). This project successfully demonstrated all of the components of an end-to-end process from mass microalgae cultivation and dewatering, to recovery and conversion of algae biomass components, to final product demonstration and process valuation; the combined results provide a framework for future commercialization of algae based biofuels.

09 BIOMASS FUELS↗

Open Innovation for a NASA Architecture Library

NASA’s Center of Excellence for Collaborative Innovation (CoECI) uses open innovation, or “crowdsourcing”, to access the global public to find ideas, concepts, designs, or solutions that meet a previously unmet need possibly resulting in significant advances in performance. The Center of Excellence for Collaborative Innovation was launched at the request of the White House Office of Science and Technology Policy. This is both a non-traditional method of innovation and a non-traditional method of outreach to the public to involve them in space technologies and programs. It has been used often for software development and new hardware technology. In this case we applied it to innovate with systems engineering tools for creating space architectures. The challenge was sponsored by NASA Engineering and Safety Center Systems Engineering Technical Fellow as part of a program for NASA’s adoption of MBSE. It was a trial to see if there would be as much participation or quality submissions with this more specialized topic and skill. The challenge sought space architecture representations and decompositions to create a library of modeled parts in a system modeling language (SysML). Mission architects mostly start from scratch to build model elements representing the functional and physical architecture of a system in SysML. There are a few beginning libraries, but these are also local to a program or group. A common library will save system engineers a large amount of time, will allow project stakeholders to recognize common graphics and quickly understand the architecture options. The challenge was promoted internationally, especially through professional organizations and universities with a systems engineering focus. It was open for 4 months, purposefully over the winter holiday break time to allow participants extra time outside of work or school. The challenge was designed so that expertise in space hardware was not necessary but getting to play with models of space architecture could provide motivation to participate. We did not receive as many entries as other broader outreach challenges, but the ones we received were extremely thorough and high quality. Solutions came from individuals and teams, students and professional consultants from the United States and Europe. We learned a few lessons about how to engage with the public and what characteristics of a problem result in good crowdsourcing results. The outreach challenge produced several useful ideas and modeled space elements, and the group will be engaging the winners to learn more about their new approaches.

innovation↗

Traffic Control via Connected and Automated Vehicles (CAVs): An Open-Road Field Experiment with 100 CAVs

The CIRCLES project aims to reduce instabilities in traffic flow, which are naturally occurring phenomena due to human driving behavior. Also called “phantom jams” or “stop-and-go waves,” these instabilities are a significant source of wasted energy. Toward this goal, the CIRCLES project designed a control system, referred to as the MegaController by the CIRCLES team, that could be deployed in real traffic. Our field experiment, the MegaVanderTest (MVT), leveraged a heterogeneous fleet of 100 longitudinally controlled vehicles as Lagrangian traffic actuators, each of which ran a controller with the architecture described in this article. The MegaController is a hierarchical control architecture that consists of two main layers. The upper layer is called the Speed Planner and is a centralized optimal control algorithm. It assigns speed targets to the vehicles, conveyed through the LTE cellular network. The lower layer is a control layer, running on each vehicle. It performs local actuation by overriding the stock adaptive cruise controller, using the stock onboard sensors. The Speed Planner ingests live data feeds provided by third parties as well as data from our own control vehicles and uses both to perform the speed assignment. The architecture of the Speed Planner allows for the modular use of standard control techniques, such as optimal control, model predictive control (MPC), kernel methods, and others. The architecture of the local controller allows for the flexible implementation of local controllers. Corresponding techniques include deep reinforcement learning (RL), MPC, and explicit controllers. Depending on the vehicle architecture, all onboard sensing data can be accessed by the local controllers or only some. Likewise, control inputs vary across different automakers, with inputs ranging from torque or acceleration requests for some cars to electronic selection of adaptive cruise control (ACC) setpoints in others. The proposed architecture technically allows for the combination of all possible settings proposed previously, that is {Speed Planner algorithms} × {local Vehicle Controller algorithms} × {full or partial sensing} × {torque or speed control}. As a result, most configurations were tested throughout the ramp up to the MegaVandertest (MVT).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Thermodynamic constraints on the textural evolution of eucrite, EET 90020

Basaltic eucrites, which are possibly parts of the Vestan crust, formed as either lava flows or intrusions, and can offer insights into early crust formation. However most eucrites have experienced thermal metamorphism, which exacerbates the challenges of understanding these samples. A variety of petrogenetic models such as, partial melting of a primitive source, fractional crystallization of magmas emplaced in a primitive crust, and/or partial melting of a eucritic source coupled with melt mixing and assimilation have been proposed in order to explain observed petrologic and chemical characteristics of eucrites. EET 90020 is an unbrecciated eucrite that has experienced significant thermal metamorphism, with temperatures of metamorphic equilibration calculated to be ~840-1040°C. Its petrologic history remains contentious, in part due to differences in trace element analyses [10,17, this work], which has led to multiple metamorphic interpretations [8,10-11]. Here, we combined petrologic observations, chemical analyses, and thermodynamic modeling, to interpret micro-domain textures identified in EET 90020 and better constrain its petrologic history. Ultimately, the development of such textures is a direct result of the geologic processes operating on a young Vesta or similar asteroid. Sample description: We identified three textural domains in EET 90020 (Figure 1); (I) A coarse grain domain with granoblastic plagioclase and pyroxene, (II) a medium grain domain with curved grain boundaries between plagioclase and pyroxene, and (III) a fine grain domain dominated by tridymite that fills in interstitial space between spherical plagioclase. Major element chemistry in phases is consistent across domains. Pyroxene have distinctive high Ca lamellae (Wo40.92En28.1Fs31.0) forming from the low-Ca pigeonite host (Wo3.5En33.7Fs62.8). There are small amounts of fayalitic olivine along pyroxene/oxide grain boundaries, where pyroxenes have reacted with ilmenite/chromite clasts. Plagioclase is anorthitic (Avg ~An88), with little variation (An86-92.5). Trace element bulk rock analyses were collected via ICP-MS and individual mineral analyses with LA-ICP-MS (Figure 2). The bulk sample has an observable enrichment in light rare earth elements, with a notable depletion in Eu. Plagioclase is enriched in LREEs, depleted in HREEs, and has a positive Eu anomaly while pyroxene is enriched in HREEs with a negative Eu anomaly. Variations in trace element analyses is likely influenced by the presence of phosphates because are highly enriched in REEs and have a dramatic effect on the bulk domain compositions. Thermodynamic modeling: The software package, Perple_X, was used to calculate mineral phase equilibria over a range of conditions using a Gibbs free energy minimization approach [12]. Models used all major/minor elements except P2O5. Thermodynamic properties from [13] were used to constrain endmember phase stabilities. Activity models were used to describe mixing in phases with solid solution. Isochemical P-T phase diagrams were constructed for the bulk thin section, and the coarse, medium and fine grain domain compositions. Domain compositions were determined using microprobe analyses of phases identified in each domain and observed modal mineral abundancies. Thermodynamically stable pyroxene compositions were extracted from model results and compared to pyroxene compositions collected via EMPA, allowing us to calculate temperatures for metamorphic equilibrium. Results: Results that use the coarse domain composition (Fig. 3a) demonstrate overlap at T ~ 1020°C indicating high T equilibration. There was no overlap for the medium and fine grain domains, or the bulk composition (e.g., Fig. 3b) indicating that these are not equilibrium assemblages. These results were replicated using data from [8,20]. Discussion: The disequilibrium results from the bulk rock, fine grain and medium grain models imply that EET 90020 was modified during and/or after peak metamorphism (open system behavior). However, this is in contrast to the coarse domain where it appears as if equilibrium was maintained. We suggest that this conflict is explained by highly localized equilibrium occurring at millimeter length-scales, which has been observed in terrestrial metamorphism where aqueous fluid infiltrate [14] or melt loss occurs [15]. It is unlikely that aqueous alteration was significant in modifying the bulk composition of EET 90020, because it lacks hydrated minerals. However, the presence of melt is supported by the presence of an abundance of curved grain boundaries and spherical grains in the medium and fine grain domains This contrasts the coarse domain, where grain boundaries are angular. Given the evidence for disequilibrium in the medium and fine grain domains, it would be inappropriate to draw additional conclusions about textural evolution using thermodynamic models of those domains. However, modeling results of the coarse grain domain, where equilibrium was maintained, can provide useful insights into the geologic evolution of EET 90020. We approximate maximum metamorphic temperatures, using the Ca component in pyroxene, at ~1020°C (Fig. 4), which is consistent with previous two-pyroxene thermometry [11]. At this temperature, up to 10 vol % melt can be produced in the coarse grain domain, which is at the boundary of minimum melt needed in order to segregate and form a melt network [16]. Tridymite and plagioclase would be the first minerals to melt out during heating starting ~990°C. This is consistent with the observed melt textures formed by tridymite and plagioclase in the fine grain domain. We suggest that a partial melting model best explains the textural development of EET 90020. Upon heating, melt generated was either trapped, resulting in no net change in the localized rock composition, or migrated, resulting in disequilibrium between pyroxene and the surrounding matrix. A positive Eu anomaly in plagioclase indicates that all of the plagioclase was melted previously. This is consistent with previous studies that have suggested that EET 90020 represents a residual eucrite that experienced partial melting and subsequent melt loss [8]. We speculate that partial melting could have occurred when EET 90020 was heated in the asteroid’s lower crust or adjacent to an intrusion [7]. Future work will study other, texturally heterogeneous samples to see if EET 90020 presents a unique circumstance or whether crustal partial melting was ubiquitous on the eucrite asteroid.

Eucrites↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗