Search NASA⌕ Search

SEARCH · Search NASA

Results for “entry”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

WHOC (Wind Hybrid Open Controller) [SWR-25-54]

The Wind Hybrid Open Controller (WHOC) is a python-based tool for real-time plant-level wind farm control and wind-based hybrid plant control. WHOC is primarily run in simulation, although we intend that it could be used for physical plants in future. WHOC provides simple farm-level (and hybrid plant-level) controls such as wake steering control, spatial filtering/consensus, active power control, and coordinated control of hybrid power plant assets; and creates an entry point for the development of more advanced controllers.

Sinner, Michael (Misha) [National Renewable Energy↗

RxnRover/CyRxnOpt

CyRxnOpt aims to provide a single software interface to various optimization algorithms, mainly designed for chemical process optimization applications. CyRxnOpt generalizes the optimization process into four high-level “phases”: Installation, Configuration, Training, and Prediction. This allows developers to program to a general interface for each phase of the optimization, simplifying the development of user-friendly tools to lower the barrier of entry into chemical process optimization, especially for automated laboratory workflows which can greatly benefit from access to various optimization techniques. It is also designed so researchers can easily add new or existing algorithms into existing workflows in a user-friendly manner.

Kulathunga, Dulitha Prasanna [Iowa State Universit↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Analyzing and Exploring Training Recipes for Large-Scale Transformer-Based Weather Prediction

Abstract The rapid rise of deep learning (DL) in numerical weather prediction (NWP) has led to a proliferation of models which forecast atmospheric variables with comparable or superior skill than traditional physics-based NWP. However, among these leading DL models, there is a wide variance in both the training settings and architecture used. Further, the lack of thorough ablation studies makes it hard to discern which components are most critical to success. In this work, we show that it is possible to attain high forecast skill even with relatively off-the-shelf architectures, simple training procedures, and moderate compute budgets. Specifically, we train a minimally modified Swin Transformer V2 (SwinV2) on ERA5 data and find that it attains superior skill in terms of mean-square errors of deterministic forecasts when compared against the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS). Almost all DL–NWP systems share a core set of hyperparameters and design decisions. To aid and expedite future DL–NWP research, we present an in-depth, systematic exploration of different loss functions, model sizes and depths, patch sizes, and multistep training objectives. We also examine the model performance with metrics beyond the typical accuracy (ACC) and RMSE and investigate how the performance scales with model size. Through our open-source code, scoring pipelines, and models, we share our findings on key aspects of the training pipeline. These ablations reduce the necessity for expensive hyperparameter tuning and lower the barrier to entry for future DL–NWP research. Significance Statement This study investigates the potential of using large-scale transformer-based models for weather prediction, showing that it is possible to achieve high forecast accuracy with simpler, off-the-shelf architectures. By training a minimally modified SwinV2 transformer on ERA5 data, we show that the model achieves competitive forecast skill in terms of mean-square error for key variables, outperforming the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS) at all lead times. Our findings suggest that effective training strategies, such as multistep fine-tuning and channel-weighted losses, significantly enhance the model’s performance. However, we also highlight that these improvements come with trade-offs in other areas, such as ensemble spread and high-frequency spatial detail. This work highlights the promise of deep learning in improving weather forecasts, which could lead to better preparedness and response to weather events, ultimately benefiting society by providing more reliable weather predictions.

Willard, Jared D. [Lawrence Berkeley National Labo↗

Sempervirens: A Fast Reconstruction Algorithm for Noisy and Incomplete Binary Matrix Representations of Trees

Applications such as reconstructing cell lineage trees (represented as phylogenetic trees) from single-cell sequencing data require reconstructing a {0,1}-matrix that has many errors and missing entries. We introduce Sempervirens, a very fast matrix reconstruction algorithm for noisy and incomplete matrix representations of phylogenetic trees. Sempervirens uses an iterative maximum-likelihood approach to determine the topology tree represented by the corrupted data. We show that Sempervirens is at least three orders of magnitude faster than other methods on thousand by thousand matrices, with the speed gap widening with larger matrices. We also show that Sempervirens matches state-of-the-art methods in reconstruction accuracy. The speed of Sempervirens enables it to be tractably applied to reconstructing much larger matrices than those that other methods can reconstruct. In addition to experimental results, we justify the algorithm with a mathematical treatment of its subprocedures.

algorithms↗

Chromosome duplication causes premature aging via defects in ribosome quality control

Down syndrome, caused by an extra copy of Chromosome 21, causes lifelong problems. One of the most common phenotypes among people with Down syndrome is premature aging, including early tissue decline, neurodegeneration, and shortened life span. Yet the reasons for premature systemic aging are a mystery and difficult to study in humans. Here we show that chromosome amplification in wild yeast also produces premature aging and shortens life span. Chromosome duplication disrupts nutrient-induced cell-cycle arrest, entry into quiescence, and cellular health during chronological aging, across genetic background and independent of which chromosome is amplified. Using a genomic screen, we discovered that these defects are due in part to aneuploidy-induced dysfunction in Ribosome Quality Control (RQC). We show that aneuploids entering quiescence display aberrant ribosome profiles, accumulate RQC intermediates, and harbor an increased load of protein aggregates compared to euploid cells. Although they maintain proteasome activity, aneuploids also show signs of ubiquitin dysregulation and sequestration into foci. Remarkably, inducing ribosome stalling in euploids produces similar aging phenotypes, while up-regulating limiting RQC subunits or poly-ubiquitin alleviates many of the aneuploid defects. We propose that the increased translational load caused by having too many mRNAs accelerates a decline in translational fidelity, contributing to premature aging.

Aneuploidy↗

Chickpea NCR13 disulfide cross-linking variants exhibit profound differences in antifungal activity and modes of action

Small cysteine-rich antifungal peptides with multi-site modes of action (MoA) have potential for development as biofungicides. In particular, legumes of the inverted repeat-lacking clade express a large family of nodule-specific cysteine-rich (NCR) peptides that orchestrate differentiation of nitrogen-fixing bacteria into bacteroids. These NCRs can form two or three intramolecular disulfide bonds and a subset of these peptides with high cationicity exhibits antifungal activity. However, the importance of intramolecular disulfide pairing and MoA against fungal pathogens for most of these plant peptides remains to be elucidated. Our study focused on a highly cationic chickpea NCR13, which has a net charge of +8 and contains six cysteines capable of forming three disulfide bonds. NCR13 expression in Pichia pastoris resulted in formation of two peptide folding variants, NCR13_PFV1 and NCR13_PFV2, that differed in the pairing of two out of three disulfide bonds despite having an identical amino acid sequence. The NMR structure of each PFV revealed a unique three-dimensional fold with the PFV1 structure being more compact but less dynamic. Surprisingly, PFV1 and PFV2 differed profoundly in the potency of antifungal activity against several fungal plant pathogens and their multi-faceted MoA. PFV1 showed significantly faster fungal cell-permeabilizing and cell entry capabilities as well as greater stability once inside the fungal cells. Additionally, PFV1 was more effective in binding fungal ribosomal RNA and inhibiting protein translation in vitro. Furthermore, when sprayed on pepper and tomato plants, PFV1 was more effective in reducing disease symptoms caused by Botrytis cinerea, causal agent of gray mold disease in fruits, vegetables, and flowers. In conclusion, our work highlights the significant impact of disulfide pairing on the antifungal activity and MoA of NCR13 and provides a structural framework for design of novel, potent antifungal peptides for agricultural use.

59 BASIC BIOLOGICAL SCIENCES↗

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

The Global Spectra-Trait Initiative: A database of paired leaf spectroscopy and functional traits associated with leaf photosynthetic capacity (v1.0.0)

The Global Spectra-Trait Initiative (GSTI) aims to generate generalizable spectra trait models using reflectance data to predict leaf traits associated with the photosynthesis capacity of leaves. It comprises a synthesized dataset of leaf trait data, input datasets and code. Leaf traits include the maximum carboxylation rate of rubisco (Vcmax), the maximum electron transport rate (Jmax), the dark respiration, as well as the prediction of leaf nitrogen, leaf mass per area (LMA), and leaf water content (LWC). The dataset comprises >7500 paired observations from around 400 species from a broad range of biomes. This dataset comprises a zip file of the GSTI GitHub repository (https://github.com/plantphys/gsti), the synthesized database (.csv) and database metadata files. This dataset was updated on 2025-12-12 with minor edits to mirror the accepted manuscript version and GitHub release (Version 1.0.0 (ESSD accepted version)). Edits included minor changes to the project documentation on GitHub and removal of 12 duplicate entries from the database.

54 ENVIRONMENTAL SCIENCES↗

Sap Velocity Data for East River Watershed Sites (2023-2025)

This dataset includes sap velocity measurements for aspen (populus tremuloides), fir (abies lasiocarpa), spruce (Engelmann spruce) and lodgepole pine (pinus contorta) trees at nine sites in the East River Watershed near Gothic, CO. This dataset was generated following a similar method as a previous dataset (Dataset. doi:10.15485/1647654) but was conducted at different sites in the area and now includes lodgepole pine. Site selection was done to explicitly improve understanding of topographic controls of tree water use. The data collection began in June 2023 and we provide data until December 2025 - though data collection is ongoing. The sap flux data were collected using ICT SFM1 sensors and are presented in both units of cm h^-1 and as kg h^-1 by multiplying the sap flux by the sapwood area of the tree. All sap flow data has been been corrected using estimates of wounding diameter, water content of wood and sap wood depth. We also provide a normalized sap velocity estimate by subtracting each measurement from that trees' annual minimum and dividing by that trees' annual maximum. This provides data for each tree and year on a 0-1 scale. This data entry contains one CSV file that includes all available sap flow data. The timestamps are provided in local time as year, day of year and hour and each measurement contains an associated latitude, longitude and site number which can be used to identify trees in a given stand. Each species is given a numeric value as: 1= populus tremuloides, 2=Engelmann spruce, 3=abies lasiocarpa and 4=pinus contorta.

abies↗

Detection of a Space Capsule Entering Earth’s Atmosphere with Distributed Acoustic Sensing (DAS)

On 24 September 2023, the Origins, Spectral Interpretation, Resource Identification, and Security-Regolith Explorer sample return capsule (SRC) entered Earth’s atmosphere after successfully collecting samples from an asteroid. The known trajectory and timing of this return provided a rare opportunity to strategically instrument sites to record geophysical signals produced by the capsule because it travelled at hypersonic speeds through the atmosphere. For this work, we deployed two distributed acoustic sensing (DAS) interrogators to sample over 12 km of surface-draped fiber-optic cables, as well as six collocated seismometer-infrasound sensor pairs, spread across two sites near Eureka, Nevada. This campaign-style rapid deployment is the first reported recording of an SRC re-entry with any distributed fiber-optic sensing technology. The DAS interrogators recorded an impulsive arrival with an extended coda which had features that were similar to recordings from both the seismometers and infrasound sensors. While the signal-to-noise ratio of the DAS data was lower than the seismic-infrasound data, the extremely dense spacing of fiber-optic sensors allowed for more phases to be clearly distinguished and the visualization of the continuous transformation of the wavefront as it impacted the ground. Unexpectedly, the DAS recordings contain less low-frequency content than is present in both the seismic and infrasound data. The deployment conditions strongly affected the recorded DAS data; in particular, we observed that fiber selection and placement exert strong controls on data quality.

58 GEOSCIENCES↗

Twinac: initiation of a community-driven accelerator digital twin framework

We present the initiation of a community-driven framework for the integration of accelerator digital twins into control systems: Twinac. Few facilities have fully integrated accelerator digital twins like at Cornell’s CHESS. Many facilities have active research to employ surrogate models to aid in operational decisions like at Argonne’s ALS, MSU’s FRIB, SLAC’s LCLS-II, and Fermilab’s FAST/IOTA, PIP-II, and main complex. To lower the barrier to entry for all accelerator facilities to build and benefit from a digital twin of their own accelerators, we propose the following software framework. Twinac will provide the capability to compose one’s own digital twin using reusable components engineered at other facilities. With this model in place, Twinac will also support tools for (1) predictive maintenance systems; (2) discovery of correlated but uncontrolled environmental factors, like seasonal temperature variations causing performance changes on power supplies, magnets, etc.; and (3) prototyping and updating sophisticated optimization and controls algorithms. The Twinac framework will enable sharing and simplified deployment of modeled components and control algorithms at all facilities. With an inter-facility team to build and support the Twinac framework, it will be easy to publish and try out the latest advancements at one’s own facility.

Miceli, Tia [Fermilab]↗

ICAT: The Interactive Corpus Analysis Tool

The Interactive Corpus Analysis Tool (ICAT) is a Python library for creating dashboards to explore textual datasets and build simple binary classification models to help filter through them and focus on entries of interest. This tool uses a form of interactive machine learning (IML), a paradigm of “machine teaching” (Simard et al., 2017) that sits at the intersection of the fields of human computer interaction (HCI), visual analytics, and machine learning. The intent of ICAT is to allow subject matter experts (SME) with limited to no experience in machine learning to benefit from an iterative human-in-the-loop (HITL) approach to building their own model without needing to understand the details of the underlying algorithm. This interactivity is achieved by allowing the user to create features, label data points, and visually manipulate a representation of the features to manually cluster and investigate data, while a model is trained on the fly based on these actions. ICAT is built on top of the Panel (Holoviz, 2018) library, using a combination of Vega, a custom IPyWidget using D3, and ipyvuetify, and is intended to be used inside of a Jupyter environment.

Martindale, Nathan [Oak Ridge National Laboratory ↗

A microscopic realization of dS$_3$

We propose a precise duality between pure de Sitter quantum gravity in 2+1 2 + 1 dimensions and a double-scaled matrix integral. This duality unfolds in two distinct aspects. First, by carefully quantizing the gravitational phase space, we arrive at a novel proposal for the quantum state of the universe at future infinity. We compute cosmological correlators of massive particles in the universe specified by this wavefunction. Integrating these correlators over the metric at future infinity yields gauge-invariant observables, which are identified with the string amplitudes of the complex Liouville string [S. Collier et al., arXiv: 2409.17246]. This establishes a direct connection between integrated cosmological correlators and the resolvents of the matrix integral dual to the complex Liouville string, thereby demonstrating one aspect of the dS _3 3 /matrix integral duality. The second aspect concerns the cosmological horizon of the dS static patch and the Gibbons-Hawking entropy it is conjectured to encode. We show that this entropy can be reproduced exactly by counting the entries of the matrix.

Collier, Scott (ORCID:0000000286476653)↗

ARM Data-Oriented Metrics and Diagnostics Package for Climate Model Evaluation

A Python-based metrics and diagnostics package is currently being developed by the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Infrastructure Team at Lawrence Livermore National Laboratory (LLNL) to facilitate the use of long-term, high-frequency measurements from the ARM Facility in evaluating the regional climate simulation of clouds, radiation, and precipitation. This metrics and diagnostics package computes climatological means of targeted climate model simulation and generates tables and plots for comparing the model simulation with ARM observational data. The Coupled Model Intercomparison Project (CMIP) model data sets are also included in the package to enable model intercomparison as demonstrated in Zhang et al. (2017). The mean of the CMIP model can serve as a reference for individual models. Basic performance metrics are computed to measure the accuracy of mean state and variability of climate models. The evaluated physical quantities include cloud fraction, temperature, relative humidity, cloud liquid water path, total column water vapor, precipitation, sensible and latent heat fluxes, and radiative fluxes, with plan to extend to more fields, such as aerosol and microphysics properties. Process-oriented diagnostics focusing on individual cloud- and precipitation-related phenomena are also being developed for the evaluation and development of specific model physical parameterizations. The version 1.0 package is designed based on data collected at ARM’s Southern Great Plains (SGP) Research Facility, with the plan to extend to other ARM sites. The metrics and diagnostics package is currently built upon standard Python libraries and additional Python packages developed by DOE (such as CDMS and CDAT). The ARM metrics and diagnostic package is available publicly with the hope that it can serve as an easy entry point for climate modelers to compare their models with ARM data. In this report, we first present the input data, which constitutes the core content of the metrics and diagnostics package in section 2, and a user's guide documenting the workflow/structure of the version 1.0 codes, and including step-by-step instruction for running the package in section 3.

54 ENVIRONMENTAL SCIENCES↗

Marine Algae Industrialization Consortium (MAGIC): Combining biofuel and high-value bioproducts to meet the RFS

The Marine Algae Industrialization Consortium (MAGIC) was formed to address pressing challenges in the commercialization of microalgae as a source of biofuel. The “Marine Algae Industrialization Consortium (MAGIC): Combining biofuel and high-value bioproducts to meet the RFS” project formally addressed two US Department of Energy Bioenergy Technologies Office (BETO) goals: (1) Model the sustainable supply of 1 million metric tonnes ash free dry weight (AFDW) cultivated algal biomass and (2) Demonstrate valuable co-products produced along with biofuel intermediates to increase value of algal biomass by 30%. To achieve these goals, the project demonstrated and validated high-value co-products to drive down the cost of biofuel by increasing the value of algae “co-products” towards increasing the selling price of total algae biomass as one of the key drivers of economics and adoption. This was accomplished through five core, interdependent tasks including: (1) strain selection to identify and deliver strains for mass culture, (2) mass culture using a hybrid cultivation system and following key operating parameters for downstream applications to provide algae feedstock, (3) recovery and conversion to evaluate two alternative methods to separate dry algae biomass into oil and residuals for downstream testing, (4) product assessment to determine biofuel, aquafeed or poultry feed product efficacy using algae biomass fractions as well as to provide critical performance data for valuation and (5) commercialization to use technoeconomic and life cycle assessments (TEA/LCA) as iterative design and assessment tools including consideration of target markets, competitors, and distribution channels to guide product assessment, development and valuation. A total of 46 peer-review publications, many open-access, provide detail of much of the work carried out and the results of the tasks. Additional reports and presentations provide other technical and public engagement material. At a high level, using a variety of approaches, more than 1000 marine microalgae strains were evaluated to ultimately identify the seven winners that were down-selected to be grown in mass culture. Strain selection demonstrated that there were no ‘super strains’ and that each candidate had strengths and limitations for specific products, growth conditions or operational considerations. Mass culture growth of these seven strains at >5000 L / 29 m 2 scale found that four them were suitable for product assessment. More than 250 kg of biomass was produced across hundreds of pond runs along with thousands of cultivation entries on the growth and biomass characteristics as well as environmental parameters. In the process, dozens of standard operating procedures were generated as was custom software to process and analyze cultivation data. Recovery and conversion of algae biomass demonstrated that a hexane solvent based extraction protocol was most effective at recovering oil (biocrude) from algae and four strains were processed to produce oil and lipid extracted algae (residuals) for downstream testing. Membrane-based oil separation was less successful, but may still be applicable to other commercial applications in the future. Product testing demonstrated that algae biocrude is of high quality and hydrotreating generated numerous fractions of high quality composition for fuel and lubricate based applications. Aquafeed studies performed at a variety of scales showed that both whole and defatted (lipid extracted algae) microalgae were suitable as a feed ingredient, but that the specifics of the fed animal and biochemical composition of the algae are critical factors when determining formulation. Similarly, poultry studies on whole and defatted microalgae generally showed positive outcomes on animal growth and health, with some microalgae providing enhanced nutritional composition of the animal product. Economic and life cycle assessments covered a wide range of possible commercialization and sustainability scenarios. Replacement value, improved product value added, consumer values marketing added valuation and improved animal health were considered as alternatives for microalgae valuation. Using the open pond system, algae productivity was identified as the key driver of commercialization economics, but combination of co-products (e.g. animal feed) with biofuel production substantially increased the total selling price of algae. Modeled microalgae selling price exceeded $\$$1500/tonne and could generate competitive biofuel selling prices below $\$$5 gallon gas equivalents using realistic algal productivities. Short (process scale) and longer (decadal trends) sustainability assessments show that marine microalgae can enhance the sustainability of energy production and lead to other realized benefits in water, fertilizer and land use for other sectors (e.g. agriculture). This project successfully demonstrated all of the components of an end-to-end process from mass microalgae cultivation and dewatering, to recovery and conversion of algae biomass components, to final product demonstration and process valuation; the combined results provide a framework for future commercialization of algae based biofuels.

09 BIOMASS FUELS↗

Towards Unlocking Insights from Logbooks Using AI

Electronic logbooks contain valuable information about activities and events concerning their associated particle accelerator facilities. However, the highly technical nature of logbook entries can hinder their usability and automation. As natural language processing (NLP) continues advancing, it offers opportunities to address various challenges that logbooks present. This work explores jointly testing a tailored Retrieval Augmented Generation (RAG) model for enhancing the usability of particle accelerator logbooks at institutes like DESY, BESSY, Fermilab, BNL, SLAC, LBNL and CERN. The RAG model uses a corpus built on logbook contributions and aims to unlock insights from these logbooks by leveraging retrieval over facility datasets, including discussion about potential multimodal sources. Our goals are to increase the FAIR-ness (findability, accessibility, interoperability, and reusability) of logbooks by exploiting their information content to streamline everyday use, to enable macro-analysis for root cause analysis, and to facilitate problem-solving automation.

43 PARTICLE ACCELERATORS↗