Search NASASearch

SEARCH · Search NASA

Results for “informatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science

Spaceflight Environmental-Telemetry Data for Biological Science

There is a critical need for better access and visualization of spaceflight environmental telemetry and mission hardware data from sensors including relative humidity, carbon dioxide, oxygen, radiation, airflow, temperature, acceleration, and acoustics. Under the stewardship of the Ames Life Sciences Data Archive (ALSDA) and GeneLab, an effort is underway to consolidate, normalize and provide accessibility of archived mission environmental data and hardware information, with the purpose of providing important context to biological data. This effort is necessary to provide scientific context of its impact upon biological and biomedical data from spaceflight missions and experiments (genomic, metagenomic, gene expression, proteomic, metabolomic, physiological, phenomics, behavioral; tabular, imaging, video). Environmental spaceflight data is derived from dozens of sources, with various formats, and in the past year a pipeline is in development to collect, curate and present this data efficiently. In the upcoming year, a new Data Visualization Portal will utilize the standardized pipeline data to provide easy user access to compare parameters and environmental conditions between missions, locations, subjects, and durations. Environmental and hardware data enables broad accessibility and analytics, without the need for advanced data informatic expertise. Familiarity with the capabilities and limitations of a variety of existing hardware/tools is a strength that could be applied to creation of improved hardware for future ecosystems on the Moon and Mars. The intention is to make biological and environmental telemetry data maximally open-access and FAIR (findable, accessible, interoperable, reusable) for data mining-informatic approaches to support knowledge discovery necessary for low Earth orbit, cis-Lunar, Mars transit, and Mars surface missions.

Danielle K. Lopez

Assessments of Physiology and Cognition in Hybrid-Reality Environments (APACHE)

NASA is planning to return to the Moon in the mid-2020s as a stepping stone to Mars missions in the 2030s. Spacewalks, or extravehicular activities (EVAs), performed on the Moon and Mars will differ in a variety of ways from those that have been performed in decades past. NASA has identified multiple risks to human health and performance associated with a crewed mission to Mars, especially those associated with exploration EVAs which are expected to be a primary mission activity. Crew may be expected to conduct up to 24 hours of EVA per person per week, where the likelihood of injury and/or mental mistakes are increased compared to ground-based training or current microgravity EVAs and the consequences of which can be catastrophic. Current test environments for exploration EVA research and technology development are large, costly facilities that are limited in their availability or capabilities. Spacesuit testing in a reduced gravity environment such as NASA’s Neutral Buoyancy Laboratory, while a good representation of the crew’s physical workload during exploration EVAs, typically has small datasets and is difficult to integrate physiological sensors or other types of crew performance measures. Meanwhile, scientific field-based testing such as NASA’s Desert Research and Technology Studies offers an operationally relevant environment for exploration EVAs, particularly for cognitive workload, but is also limited by small datasets, lack of a pressurized spacesuit, and obtrusive measures. The limitations of current analogs for exploration EVAs identify a need for a new test environment that can approximate both the physical and cognitive demands associated with exploration EVAs to enable rapid, controlled, and repeatable evaluations of human health and performance risks of exploration missions. In response, the Human Physiology, Performance, Protection, and Operations Laboratory (H-3PO) at NASA Johnson Space Center has developed a hybrid reality exploration EVA analog named the Assessments of Physiology And Cognition in Hybrid-reality Environments (APACHE) to address these limitations using a combination of virtual, physical, and hybrid reality techniques. The APACHE facility resides at NASA Johnson Space Center and serves as a large “sandbox” for EVA research and simulation. At its center is a roughly 15x20ft space surrounded by a 14” tall sandbox partially filled with lunar regolith simulant to emulate the physical feeling of walking on a planetary surface and to allow for simulated geology operations. Nearby, a curved passive treadmill (Skillmill Connect, Technogym, Fairfield, NJ) and an omnidirectional treadmill (Infinadeck, Infinadeck, Rocklin, CA) are included to enable exploration of these large virtual environments while also imposing the physical demands, representative timelines, and cognitive burdens required to navigate and traverse these distances during exploration EVA. A 6DOF motion platform is used to simulate rover operations and supports various human performance evaluations and associated risks. Lastly, APACHE can support two extravehicular (EV) crewmembers working in tandem. A computer workstation is located nearby and also supports an intravehicular (IV) crewmember as part of a full mission simulation. The IV crewmember has direct video and audio communication with the EV crew in VR to provide operational and procedural support. The software used in APACHE was created by the JSC Engineering Directorate, in partnership with Buendea, powered by a custom Unreal Engine 5 (UE5.3, Epic Games) project. APACHE currently utilizes the HTC Vive Pro Eye in a wireless configuration for VR simulations. There are two virtual environments that subjects can explore within APACHE, a Lunar and Martian surface. The virtual Lunar surface was created from LIDAR data of the Lunar South Pole to create roughly 16 sq km of explorable terrain. The virtual Martian surface contains roughly 400 sq km of explorable terrain derived from Mars Reconnaissance Orbiter LIDAR data of the Jezero Crater. The immersion and related cognitive burdens of conducting a planetary EVA is simulated through a series of EVA-relevant tasks performed in the VR environment, using these high-fidelity visual representations. Additionally, APACHE includes biosensor driven informatics, such as real-time heart rate monitoring and/or derived values from model simulations, for active monitoring by the EV crew and added cognitive demand. A “Wizard of Oz” control panel enables test operators to activate contingency events such as simulated spacesuit malfunctions, loss of communications, and/or limited visibility. Embedded performance measures such as accuracy, completeness, and execution time have been developed for various exploration tasks to objectively quantify crew performance during an EVA and compare impacts to performance when different environmental stressors, both physical and cognitive, are added to or removed from the simulation. Additionally, validated cognitive and operational performance measures such as the Digit Symbol Substitution Task have been recreated and embedded in VR for direct and relatively unobtrusive measurement of motor perception. The APACHE environment currently supports multiple research studies at NASA. Examples include the CHAPEA project, a series of simulated year-long missions on Mars by a 4-person crew; and the CO2 Contingency Walk Back Study, an investigation of elevated CO2 exposure on crew performance during a contingency EVA scenario. APACHE also provides a test environment to support the development of the Crew State and Risk Model, which is a collection of individualized, mathematical models of crew physical and cognitive state; and the Personalized EVA Informatics and Decision Support system, an operational tool for flight controllers, and eventually a self-reliant Martian crew, to make biomedically-informed decisions in real-time to optimize the EVA planning and execution with respect to crew health and performance. Some technical challenges associated with developing the APACHE environment, as well as current limitations, include VR limitless natural walking with a hybrid spacesuit simulator, optimizing performance for wireless PC VR streaming while maintaining a high degree of visual fidelity, and the integration of various physiological (metabolic masks) and psychometric (eye tracking) sensors with the VR headset.

Human Performance

The Role and Evolution of NASA's Earth Science Data Systems

One of the three strategic goals of NASA is to Advance understanding of Earth and develop technologies to improve the quality of life on our home planet (NASA strategic plan 2014). NASA's Earth Science Data System (ESDS) Program directly supports this goal. NASA has been launching satellites for civilian Earth observations for over 40 years, and collecting data from various types of instruments. Especially since 1990, with the start of the Earth Observing System (EOS) Program, which was a part of the Mission to Planet Earth, the observations have been significantly more extensive in their volumes, variety and velocity. Frequent, global observations are made in support of Earth system science. An open data policy has been in effect since 1990, with no period of exclusive access and non-discriminatory access to data, free of charge. NASA currently holds nearly 10 petabytes of Earth science data including satellite, air-borne, and ground-based measurements and derived geophysical parameter products in digital form. Millions of users around the world are using NASA data for Earth science research and applications. In 2014, over a billion data files were downloaded by users from NASAs EOS Data and Information System (EOSDIS), a system with 12 Distributed Active Archive Centers (DAACs) across the U. S. As a core component of the ESDS Program, EOSDIS has been operating since 1994, and has been evolving continuously with advances in information technology. The ESDS Program influences as well as benefits from advances in Earth Science Informatics. The presentation will provide an overview of the role and evolution of NASAs ESDS Program.

Remote Sensing

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods. REFERENCES [1] Open science in space. Nature Medicine, 2021. 27(9): p. 1485-1485. [2] Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. [3] Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5. [4] Whetzel, P.L., et al., BioPortal: enhanced functionality via new Web services from the National Center for Biomedical Ontology to access and use ontologies in software applications. Nucleic Acids Res, 2011. 39(Web Server issue): p. W541-5.

informatics

Computer Human Interface Challenges in Space Exploration

NASA’s plans to return humans to the Lunar surface require overcoming a variety of challenging technical and operational obstacles. In 2022, NASA formed the Extravehicular Activity (EVA) and Human Surface Mobility (HSM) Program (EHP) at the Johnson Space Center with responsibilities including development of space suits and surface mobility systems for Lunar missions. This program includes a Technology Development and Partnerships office chartered to identify high priority gaps in capabilities for Lunar surface mobility and to coordinate resources to close those gaps. This presentation details the EHP technology roadmap for “Informatics and Decision Support,” a subset of spacecraft avionics focused on effective and autonomous crew interaction with spaceflight systems. The gaps, grouped into displays, audio systems, and information technology infrastructure, are largely driven by the unique interaction requirements for human spacecraft and the severe radiation environments beyond low earth orbit. The roadmap identifies ongoing activities and paths to technology infusion into Lunar spacecraft. NASA is seeking input on the content and ideas for alternative paths to gap closure. Closing these gaps is important to successful human operations on the Lunar surface and vital to NASA’s long-term goal of human missions to Mars.

Spacecraft Displays

High-Throughput Strategies that Encompass Experiments and Machine Learning to Predict the Mechanical Properties of Additive Manufactured Aerospace Alloys

Small Punch Test (SPT) uses a thin disk of material to predict mechanical properties. While SPT has existed for decades, it has been used largely as a qualitative evaluator of mechanical properties. Recent advances in computational modeling have enabled the extraction of uniaxial stress-strain response from the measured SPT load-displacement data. Due to small sample volumes and unidirectional testing, SPT is conducive to high-throughput automation and ideally suited to extract properties from high-cost materials. Aerospace alloys have been of recent interest to the Additive Manufacturing (AM) community due to AM’s unique ability to fabricate complex designs not possible, or extremely arduous, with conventional manufacturing. In this research, SPT, coupled with Materials Informatics and computational modeling, is used to develop relevant Process-Structure-Property relationships to decrease the cost and time of process optimization for AM aerospace alloys, namely Inconel 718, Inconel 625, and Niobium C103.

High-throughput Testing

Combining computational modeling and experimental library screening to affinity-mature VEEV-neutralizing antibody F5

Engineered monoclonal antibodies have proven to be highly effective therapeutics in recent viral outbreaks. However, despite technical advancements, an ability to rapidly adapt or increase antibody affinity and by extension, therapeutic efficacy, has yet to be fully realized. We endeavored to stand-up such a pipeline using molecular modeling combined with experimental library screening to increase the affinity of F5, a monoclonal antibody with potent neutralizing activity against Venezuelan Equine Encephalitis Virus (VEEV), to recombinant VEEV (IAB) E1E2 antigen. We modeled the F5/E1E2 binding interface and generated predictions for mutations to improve binding using a Rosetta-based approach and dTERMen, an informatics approach. The modeling was complicated by the fact that a high-resolution structure of F5 is not available and the H3 loop of F5 exceeds the length for which current modeling approaches can determine a unique structure. A subset of the predicted mutations from both methods were incorporated into a phage display library of scFvs. This library and a library generated by error-prone PCR were screened for binding affinity to the recombinant antigen. Results from the screens identified favorable mutations which were incorporated into 12 human-IgG1 variants. The best variant, containing eight mutations, improved KD from 0.63 nM (parental) to 0.01 nM. While this did not improve neutralization or therapeutic potency of F5 against IAB, it did increase cross-reactivity to other closely related VEEV epizootic and enzootic strains, demonstrating the potential of this method to rapidly adapt existing therapeutics to emerging viral strains.

affinity-maturation

FREDA: A Web Application for the Processing, Analysis, and Visualization of Fourier‐Transform Mass Spectrometry Data

The high-resolution measurement capability of Fourier-transform mass spectrometry (FT-MS) has made it a necessity for exploring the molecular composition of complex organic mixtures, like soil, plant, aquatic, and petroleum samples. This demand has driven a need for informatics tools to explore and analyze FT-MS data in a robust and reproducible manner. FREDA is an interactive web application developed to enable spectrometrists to format, process, and explore their FT-MS data without the need for statistical programming expertise. FREDA was built to explore outputs from a molecular identification tool, like CoreMS, and provide a suite of methods to filter data, compute chemical properties of peaks, statistically compare samples and groups of samples, conduct exploratory data analysis, and download the results with a report detailing all steps conducted. To demonstrate the utility of FREDA, an example analysis was conducted using FT-MS data from a soil microbiology study of samples collected in two different soil depths at the Sphagnum bog forest north of Grand Rapids, Minnesota. Differences between the two depths are observed using Kendrick, Gibbs free energy, and van Krevelen plots. G-tests are used to quantify a significant difference between the groups. All analyses and plotting are conducted using only the FREDA application. FREDA is an open-source and readily available web application that allows users to explore and make statistically valid conclusions about their FT-MS data. The application is available online (https://map.emsl.pnnl.gov/app/freda) with a tutorial web series (https://youtu.be/k5HLE2kNSBY?si=yB6sGoyvzxrFf5MP) and freely accessible code on Github (https://github.com/EMSL-Computing/FREDA).

47 OTHER INSTRUMENTATION

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection

A machine learning approach to quantify degradation of nuclear fuels and the effects of fission products

Nuclear fuel performance is critically dependent on understanding the evolution of fuel properties under operational conditions, a complex challenge driven by chemical changes and substantial radiation damage during fission. Traditionally, property evolution has been determined via empirical data collected following irradiation. However, these empirical correlations are limited in their applicability beyond the specific conditions in which they were obtained. This study explores a novel approach to address this challenge by applying materials informatics to develop a machine learning random forest (ML-RF) model that captures the effects of fission products on fuel compounds. The model predicts formation enthalpy (ΔH f ) by leveraging extensive quantum materials property data and correlating it with material descriptors such as composition, atomic and site features, and crystal lattice properties. This ML-RF model enables rapid interpolation across the compositional and structural spaces covered by the training data, thus supporting high-throughput screening and energetic ranking of candidate phases. The model demonstrates the ability to predict ΔH f with a mean absolute error (MAE) of approximately 0.1 to 0.2 eV/atom across a wide range of compounds, including key nuclear fuel systems (U-O, U-N, U-C, U-Si, and U-Mo). For example, it was used to assess shifts in stoichiometry for UO 2 (O/M) and UN (N/M) fuels, revealing their distinct tendencies in chemical potential variation and enabling preliminary convex hull analyses. Furthermore, the model provides insights into how individual fission products affect fuel properties. Results indicate that larger fission products (e.g., Nd, Pu, Ce) have a more pronounced impact on UO 2 , while lighter ones (e.g., Zr) strongly influence UN. Here, the model developed in this work can be used to support the Accelerated Fuel Qualification approach by facilitating preliminary evaluations prior to extensive materials modeling and experimentation. To this end, the trained model has been made available to the fuel community to support ongoing fuel development efforts.

Accelerated fuel qualification

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance

Autonomous Synthesis and Inverse Design of Electrochromic Polymers with High Efficiency and Accuracy

Here, the design and synthesis of functional polymers, aimed at targeted properties through specific structures, have long been challenged by their complex and often nonlinear structure–property relationships. Key processes, including knowledge accumulation for predictive design and experimental refinement and validation, are traditionally labor-insensitive and time-consuming, making it difficult to balance accuracy and efficiency. Here, we introduce an accelerated, autonomous system for the on-demand synthesis of electronic polymers that achieves the desired electrochromic functionality with high accuracy and efficiency. Our approach leverages large language model-assisted data mining, a physics-informed copolymer machine learning model, and an AI-driven autonomous robotic workflow in the Polybot lab. Within 72 h, Polybot autonomously synthesized electrochromic polymers (ECPs) with targeted, previously-unreported color values, including green polymers with specific absorption profiles, precisely fine-tuning copolymer structures with a 5% step size in comonomer composition within a three-monomer system. A publicly accessible ECP informatics database has also been created to foster knowledge exchange.

AI-driven Robotic Lab

Modeling the impact of structure and coverage on the reactivity of realistic heterogeneous catalysts

Adsorbates often cover the surfaces of catalysts densely as they carry out reactions, dynamically altering their structure and reactivity. Understanding adsorbate-induced phenomena and harnessing them in our broader quest for improved catalysts is a substantial challenge that is only beginning to be addressed. Here, in this work, we chart a path toward a deeper understanding of such phenomena by focusing on emerging in silico modeling methodologies, which will increasingly incorporate machine learning techniques. We first examine how adsorption on catalyst surfaces can lead to local and even global structural changes spanning entire nanoparticles, and how this affects their reactivity. We then evaluate current efforts and the remaining challenges in developing robust and predictive simulations for modeling such behavior. Last, we provide our perspectives in four critical areas—integration of artificial intelligence, building robust catalysis informatics infrastructure, synergism with experimental characterization, and adaptive modeling frameworks—that we believe can help surmount the remaining challenges in rationally designing catalysts in light of these complex phenomena.

catalytic mechanisms

A high-throughput and data-driven computational framework for novel quantum materials

Two-dimensional layered materials, such as transition metal dichalcogenides (TMDs), possess an intrinsic van der Waals gap at the layer interface, allowing for remarkable tunability of the optoelectronic features via external intercalation of foreign guests such as atoms, ions, or molecules. Herein, we introduce a high-throughput, data-driven computational framework for the design of novel quantum materials derived from intercalating planar conjugated organic molecules into bilayer transition metal dichalcogenides and dioxides. By combining first-principles methods, material informatics, and machine learning, we characterize the energetic and mechanical stability of this new class of materials and identify the fifty (50) most stable hybrid materials from a vast configurational space comprising ∼105 materials, employing intercalation energy as the screening criterion.

Kastuar, Srihari M. (ORCID:0000000279001561)

Quasiparticle spectroscopy in technologically relevant niobium using London penetration depth measurements: experiment and theory

Abstract The London penetration depth, λ ( T ) , was measured in various forms of niobium, including foils, thin films, single crystals, and samples from superconducting radio-frequency (SRF) cavities. We observed a significant difference in λ ( T ) at low temperatures, T < T c / 3 , due to low-energy quasiparticles. In particular, an unusual downturn of λ ( T ) on cooling in the SRF cavity samples required to take into account deep in-gap bound states. Theoretical modeling using the generalized Dynes density of states shows that such in-gap states lead to a downturn or a peak in λ ( T ) upon cooling. Combined, experimental and theoretical findings provide a method for detecting two-level systems or states related to magnetic impurities in the bulk of niobium. This result is particularly relevant for the quantum informatics sciences technologies used in qubits and circuit quantum electrodynamics architecture based on SRF cavities.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo

Dynamic Retrieval Augmented Generation of Ontologies using Artificial Intelligence (DRAGON-AI)

Ontologies are fundamental components of informatics infrastructure in domains such as biomedical, environmental, and food sciences, representing consensus knowledge in an accurate and computable form. However, their construction and maintenance demand substantial resources and necessitate substantial collaboration between domain experts, curators, and ontology experts. We present Dynamic Retrieval Augmented Generation of Ontologies using AI (DRAGON-AI), an ontology generation method employing Large Language Models (LLMs) and Retrieval Augmented Generation (RAG). DRAGON-AI can generate textual and logical ontology components, drawing from existing knowledge in multiple ontologies and unstructured text sources.We assessed performance of DRAGON-AI on de novo term construction across ten diverse ontologies, making use of extensive manual evaluation of results. Our method has high precision for relationship generation, but has slightly lower precision than from logic-based reasoning. Our method is also able to generate definitions deemed acceptable by expert evaluators, but these scored worse than human-authored definitions. Notably, evaluators with the highest level of confidence in a domain were better able to discern flaws in AI-generated definitions. We also demonstrated the ability of DRAGON-AI to incorporate natural language instructions in the form of GitHub issues.These findings suggest DRAGON-AI's potential to substantially aid the manual ontology construction process. However, our results also underscore the importance of having expert curators and ontology editors drive the ontology generation process.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION