Search NASA⌕ Search

SEARCH · Search NASA

Results for “knowledge extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Single cell RNA sequencing reveals shifts in cell maturity and function of endogenous and infiltrating cell types in response to acute intervertebral disc injury

Intervertebral disc (IVD) degeneration contributes to disabling back pain. Degeneration can be initiated by injury and progressively leads to an irreversible loss of cells and function. IVD function restoration through cell replacement therapies have had limited success due to knowledge gaps in the critical cell populations important for repair. Here, in this study, we used single cell RNA sequencing to identify the transcriptional changes of IVD resident and infiltrating cell populations from Control and Injured coccygeal IVDs extracted from 12-week-old female C57BL/6J mice 7 days post injury. Clustering, gene ontology, and pseudotime trajectory analyses determined transcriptomic divergences with injury, flow cytometry identified they types of infiltrating immune cells, and immunofluorescence was utilized to define mesenchymal stem cell (MSC) localization. We identified 11 distinct clusters that included IVD, immune, vascular cells, and MSCs. Differential gene expression analysis determined that Outer Annulus Fibrosus, Neutrophils, Saa2-High MSCs, Macrophages, and Krt18 + Nucleus Pulposus (NP) cells were the major drivers of transcriptomic differences between Control and Injured cells. Gene ontology revealed that the most upregulated biological pathways were angiogenesis and T cell-related while wound healing and ECM regulation were downregulated. Pseudotime trajectory analyses revealed that IVD injury directed cells towards increased differentiation in all clusters, except for Krt18 + NP cells which remained in a less mature cell state. Saa2-High and Grem1-High MSCs populations shifted towards more differentiated IVD cells profiles with injury and localized distinctly within the IVD. This study revealed novel MSC populations with the potential to be leveraged for future IVD repair studies.

Cartilage↗

Biological response of eelgrass epifauna, Taylor's Sea hare ( Phyllaplysia taylori ) and eelgrass isopod ( Idotea resecata ), to elevated ocean alkalinity

Abstract. Marine carbon dioxide removal (mCDR) approaches are under development to mitigate the effects of climate change by sequestering carbon in stable reservoirs, with the potential co-benefit of local reductions in coastal acidification impacts. One such method is ocean alkalinity enhancement (OAE). A specific OAE method is the generation of aqueous alkalinity via electrochemistry to enhance the alkalinity of the receiving water by the extraction of acid from seawater, thereby avoiding the issues of solid dissolution kinetics and the release of impurities into the ocean from alkaline minerals. While electrochemical acid extraction is a promising method for increasing the carbon dioxide sequestration potential of the ocean, the biological effects of increasing seawater alkalinity and pH within an OAE project site are relatively unknown. This study aims to address this knowledge gap by testing the effects of increased pH and alkalinity, delivered in the form of aqueous NaOH, on two eelgrass epifauna in the US Pacific Northwest, Taylor's sea hare (Phyllaplysia taylori) and eelgrass isopod (Idotea resecata), chosen for their ecological importance as salmon prey and for their role in eelgrass ecosystems. Four-day experiments were conducted in closed bottles to allow measurements of the evolution of carbonate species throughout the experiment, with water refreshed twice daily to maintain elevated pH, across pHNBS (NBS standard scale) treatments ranging from 7.8 to 9.3. Sea hares experienced mortality in all pH treatments, ranging from 37 % mortality at pHNBS 7.8 to 100 % mortality at pHNBS 9.3. Isopods experienced lower mortality rates in all treatment groups, ranging from 13 % at pHNBS 7.8 to 21 % at pHNBS 9.3, which did not significantly increase with higher pH treatments. These experiments represent an extreme of constant exposure to elevated pH and alkalinity, which should be considered in the context of both the natural variation and the dilution of alkalinity experienced by marine communities across an OAE project site. Different invertebrate species will likely have different responses to increased pH and alkalinity, depending on their physiological vulnerabilities. Investigation of the potential vulnerabilities of local marine species will help inform the decision-making process regarding mCDR planning and permitting.

marine carbon dioxide removal↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Weak-charge form-factor determination at the electron-ion collider

Determining the weak charge form factor, 𝐹 𝑊 ⁡(𝑄 2 ), of nuclei over a continuous range of momentum transfers, 0 ≲ 𝑄 2 ≲ 0.1 GeV 2 , is essential for mapping out the distribution of neutrons in nuclei. The neutron density distribution has significant implications for a broad range of areas, including studies of nuclear structure, neutron stars, and physics beyond the Standard Model. Currently, our knowledge of 𝐹 𝑊 ⁡(𝑄 2 ) comes primarily from fixed target experiments that measure the parity-violating asymmetry in coherent elastic electron-ion scattering. Fixed target experiments, such as CREX and PREX-1,2, have provided high-precision weak charge form factor extractions for the 48 Ca and 208 Pb nuclei, respectively. However, a major limitation of fixed target experiments is that they each provide data only at a single value of 𝑄 2 . With the proposed electron-ion collider (EIC) on the horizon, we explore its potential to impact the determination of the weak charge form factor. While it cannot compete with the precision of fixed target experiments, it can provide data over a wide and continuous range of 𝑄 2 values, and for a wide variety of nuclei. We show that with data corresponding to an integrated luminosity of ℒ ∼ 500/𝐴 fb −1 , where 𝐴 is the nucleus atomic weight, the EIC can significantly impact constraints by lifting degeneracies in theoretical models of the neutron density distribution. Ensuring EIC detector coverage at low 𝑄 2 and large negative pseudorapidities will be essential for such 𝐹 𝑊 ⁡(𝑄 2 ) measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multi-Tiered Estimation for Correlation Spectroscopy in 3D (MTECS3D) v0.1

This innovative software estimates rotational diffusion coefficients of particles from X-ray photon correlation spectroscopy (XPCS) data of monodispersed particle systems. It is the first method capable of extracting rotational diffusion information from three-dimensional particle systems using XPCS. Using the angular-temporal cross-correlation of the XPCS images, this software is able to estimate the rotational diffusion coefficients with only a few percent relative errors while requiring minimal prior knowledge of particle structures. This software enhances XPCS analysis capabilities, allowing researchers to study translational and rotational Brownian dynamics of particles in suspension across various temporal and spatial scales.

Hu, Zixi↗

Microwave-assisted catalytic conversion of waste biomass and plastic feedstocks via thermochemical routes

Microwave-assisted catalytic conversion of waste feedstocks to fuels and value-added chemicals shows incredible promise as an efficient pathway to support the U.S. Department of Energy’s vision toward strengthening the nation’s energy independence. Microwave-heated systems have the potential to outperform conventional technologies through energy-efficient heating and improved product selectivity. This chapter emphasizes microwave-assisted catalytic approaches for waste conversion, allowing maximum energy recovery and extraction of valuable chemicals from waste feedstock such as biomass and plastics while reducing undesired byproducts. A gap remains in understanding how microwaves interact with materials to enable rapid and selective heating, which is crucial for improving catalytic efficiency. This chapter attempts to address this knowledge gap by proposing mechanisms that explain the microwave-catalytic interactions for efficient conversion of biomass-plastic wastes. In addition, comparisons with conventional catalytic technologies as well as the potential for scale ups and future commercialization of microwave-catalytic waste conversion technologies are also discussed.

microwave-assisted catalytic conversion↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Spectroscopic Online Monitoring: Using a Multi-Track Visible Spectrometer to Facilitate a Mass Balance Study in a Simulated TALSPEAK Process

Nuclear energy is a promising low-carbon energy candidate to meet the increased demand for green energy, where the integration of fuel recycling can have significant benefits for material usage and waste reduction. Utilizing in situ monitoring tools can provide ample opportunities to better control and safeguard nuclear material recycle processes while also offering knowledge and insight into real-time solution properties. The simultaneous measurement of analytical targets in multiple process locations can enable real-time mass balance and material accountancy calculations. This is demonstrated here with a mass balance study of Nd 3+ on countercurrent aqueous/organic metal extraction within a single centrifugal contactor. The Nd 3+ concentration was simultaneously monitored at the inlets and outlets of both aqueous and organic phases using a visible absorbance detector that allowed for the simultaneous measurement of up to six locations. The Nd 3+ concentration was calculated by using chemical data science algorithms, where model training sets were collected on a single track of the detector. The discussion includes addressing the challenges of using a model collected on a single track and applying it as a model across the other tracks on the detector. Each track of the detector corresponds to one measurement location on the contactor. The difference in the integrated moles of Nd 3+ between the inlet and outlet at the end of the experiment was near zero, indicating that the mass balance of this experiment was maintained. Overall, the online spectroscopic monitoring was able to follow changing solution conditions and accurately measure the concentration of Nd 3+ in different locations within the contactor system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Generative AI to implement the discrepancy checker for a Nearly Autonomous Management and Control System for Advanced Reactors

Developments related to generative artificial intelligence (AI) have brought a major breakthrough in AI. These developments are rapidly accelerating developments in different science and engineering applications. Nearly Autonomous Management and Control (NAMAC) system provides recommendations to the operator for maintaining the safety and performance of the reactor. The discrepancy checker (DC) is an important component of the NAMAC) system, whose goal is to determine if the plant is moving towards the expected system state after the control actions are injected. In this work, we explore generative AI methods, particularly, a generative pretrained transformer (GPT) for implementing the DC function in NAMAC. The GPT-based DC aims to alert the operator in situations outside NAMAC’s scope and act as a chatbot the operator can use to retrieve relevant information. This study involves two versions of GPT developed by OpenAI: GPT-3.5 and GPT-4. These GPTs are trained on huge amounts of undisclosed general domain datasets. We explored two methods to adapt GPTs for DC implementation in NAMAC: fine-tuning and retrieval augmented generation. A small knowledge base (information file) that encompasses rules for DC implementation and some general information related to NAMAC has been created to support DC implementation using GPT. In this work, the GPT-based DC implementations have been tested for their reasoning abilities, comprehension, information retrieval, and extraction abilities. It should be noted that this paper only presents a preliminary study to test the feasibility of DC implementation using generative AI technology. Given the potential risks and severe consequences associated with nuclear reactor applications, combined with the black-box nature of AI, extensive offline and online testing and reliability analyses of GPT-based DCs are needed for further developing such capabilities.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Investigating the Impacts of Direct Dissolution Conditions on the Radiolytic Longevity of Butyramide Extractants

Removing the nitric acid (HNO3) dissolution step in used nuclear fuel (UNF) reprocessing would reduce the volume of radioactive waste streams generated, thereby, improving process efficiency. A promising strategy for this is the direct dissolution of UNF that has been pretreated by voloxidation into an organic solvent composed of specialized extractants and diluent. However, removal of the aqueous HNO3 phase from the envisioned reprocessing system has the potential to drastically change the suite of radiation-induced processes occurring, and thus, alter the longevity of proposed reagents. Furthermore, the impacts of fission product and transuranic metal ion complexation on the aforementioned radiation-induced processes is poorly understood, and yet can cause significant changes in radiolytic longevity. To bridge these knowledge gaps and support the continued development of direct dissolution strategies, we present an investigation into the impacts of direct dissolution conditions on the gamma radiation-induced degradation of N,N-di-(2-ethylhexyl) butyramide (DEHBA) and N,N-di-(2-ethylhexyl)isobutyramide (DEHiBA) ligands—candidate replacements for tributyl phosphate—in pre-equilibrated n-dodecane solvent in the presence and absence of envisioned loading amounts of uranium.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

From Machine Learning to Machine Reasoning: A Model-based Approach to Analyze Equipment Reliability Data

In current nuclear power plants (NPPs) a large amount of condition-based data which can be used to assess and monitor component health and performance. Assessing component health from such data can be performed with a large variety of methods. While the analysis of numeric data can be performed with several methods, the extraction of information from textual data remains a challenge. Currently employed natural language processing (NLP) methods do not really provide quantitative information that might be contained in IRs. In addition, the integration of numeric and textual data to identify possible causal relationships between data elements is still an unresolved challenge. This paper presents an approach to extract information from textual (e.g., incident or maintenance reports) and numeric data that relies on model based system engineer (MBSE) models. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence while semantic analysis is designed to analyze the logic structure of a sentence. An innovative element of our approach is that semantic analysis uses MBSE models to identify links between textual elements. Similarly, numeric data is directly linked to elements of the MBSE models in order to map which functions are being monitored.

97 - MATHEMATICS AND COMPUTING↗

Williston Basin CORE-CM Initiative Final Report

The University of North Dakota Energy & Environmental Research Center (EERC) is leading the Williston Basin Carbon Ore, Rare Earth, and Critical Minerals (CORE-CM) Initiative to drive the expansion and transformation of coal and coal-based resource usage within the Williston Basin to produce rare-earth elements (REEs), CMs, and nonfuel carbon-based products (CBPs). This project is the first phase in a long-term program and set the stage for future work by assessing resource, market, technology, and infrastructure knowledge; identifying knowledge gaps; developing a series of plans to be carried out in future work; and initiating stakeholder engagement. Composed of several tasks, the project sought to identify, characterize, and assess several necessary aspects vital to make this future work a reality. The project’s fundamental task was to characterize the Williston Basin CORE-CM resources. Over 2500 samples from multiple sources were utilized to begin the assessment. Several locations were identified in western North Dakota where sample analysis identified the total REE (TREE) concentration as being over 500 parts per million (ppm), which is at a concentration level that would be suitable to consider for mining and extraction. Current operating coal mines have sufficient concentrations of TREEs for consideration. However, the current data across the basin are still not adequate to fully characterize REE and CM content nor give reliable estimates of the total resource potential. Waste stream reuse was also considered, and several streams were identified which ranged from potential energy sources to chemicals to material wastes. This includes streams that result from oil and gas production. These streams are not fully characterized, and further data are needed before they can be accurately assessed. Infrastructure within the Williston Basin is suitable for expansion of a new industry to mine, extract, and concentrate REEs and CMs. The development of this industry will not only preserve many existing jobs in the coal-mining industry but produce many new jobs. The supply chain for REEs and CMs is currently controlled outside of the United States in nations such as China, but the potential to develop the supply chain within the basin is considered possible. Processing of the mined materials for REEs and CMs needs further research. The technology and knowhow exist outside of the United States, and within the country much of the knowledge has been lost and must be regained. To develop the supply chain and regain lost processing technology, the creation of technology innovation centers (TICs) is crucial. The Williston Basin contains several similar centers and entrepreneurial assistance for other industries that can be applied in the development of REE and CM innovation centers. Education to develop the new skill sets required is also needed. Outreach is important for the development of the REE and CM industry within the basin. Understanding throughout federal and state governments, state agencies, industry, and resource end users is vital for the industry to form and grow. Through this project these groups have been contacted through bulletins, presentations, webinars, and annual symposiums. The report is a summary of the work conducted and throughout refers to a series of appendixes which contain more thorough and specific information about each section.

01 COAL, LIGNITE, AND PEAT↗

Enhancing f -Element Separations with ADAAM-EH: The Impact of Phase Modifiers and a DGA Aqueous Complexant

Recent investigations have used a 2-ethylhexyl diamide amine (ADAAM-EH) for Am/Cm separations in combination with N,N,N ',N '-tetraethyldiglycolamide as an aqueous complexant to achieve an unprecedented separation factor of 41. The aim of this research effort is to understand the speciation of trivalent lanthanide (Ln) and actinide (An) ions in the organic phase of an ADAAM-EH extraction system, both with and without phase modifiers (PM) (1-octanol and tri-n-butyl phosphate (TBP)). Leveraging spectroscopic techniques in combination with distribution ratio measurements provides an understanding of organic phase f-element ligand complexation. In the absence of PM, Ln is extracted in a stoichiometric 1:1 [M(ADAAM-EH) 1 (NO 3 ) x (H 2 O) 1 ](NO 3 ) 3-x complex. The addition of 1-octanol at 20 vol % results in multiple species present. One of the species is the same as the no PM case, and the other species results in an increased -OH coordination to the inner sphere, potentially displacing some NO 3 . In the case of TBP, increasing concentration results in additional red-shifted bands in the UV-visible spectra, suggesting the complexation of additional ligands of either ADAAM-EH or TBP. Finally, the new system knowledge obtained by and spectroscopic experiments will provide benchmarking information for computational studies of the inner- and outer-sphere coordination environments of f-element cations and insights into ADAAM-EH adduct formation with PM, like 1-octanol and TBP.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Role of Fluid and Temperature in Fracture Mechanics and Coupled THMC Processes for Enhanced Geothermal Systems

The ability to sustain high energy extraction efficiency from Enhanced Geothermal Systems (EGS) is affected by the ability to measure, characterize, and predict the effect of different perturbations that arise from fluid injection rates, fluid temperature, and shut-in conditions on existing and stimulated fracture systems throughout a subsurface reservoir’s lifecycle. The difficulty arises from limited knowledge of the interaction of fluid and temperature driven fractures with frictional interfaces that govern local deformation and frictional behavior in the surrounding rock. Thermo-poro-mechanical coupled processes tend to dominate this interaction and strongly affect the permeability and local stress distributions that impact efficiency and prevent optimal stimulation conditions over multi-decadal time frames.

Pyrak-Nolte, Laura [Purdue Univ., West Lafayette, ↗

Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high-throughput identification and lineage assignment in Pakistan

Bacillus anthracis, the causative agent of anthrax, is a highly virulent zoonotic pathogen primarily affecting domesticated and wild herbivores. Human exposure to B. anthracis is primarily through contact with infected animals or contaminated animal products. In Pakistan, where livestock vaccines are largely unavailable and infected carcasses are often disposed of improperly, the risk to humans, wildlife and livestock is significant. Currently, the diagnosis of anthrax infections and outbreak tracing necessitates the isolation and culturing of B. anthracis, a process that requires BSL-3 facilities. In this study, we show that positive identification, genome reconstruction and lineage assignment can be accomplished using bioinformatic analysis of DNA extracted directly from environmental samples that would otherwise provide the starting material for isolation and culturing. This approach does not require laboratory target enrichment as is necessary for other pathogens, due in part to the extremely high bacterial load in the bloodstream in the deceased animals. Using these methods, we greatly expand the knowledge of endemic B. anthracis in Pakistan. We provide the first reference B. anthracis genomes from Pakistan since the 1970s and identify A.Br.014 Aust94 as a minor circulating sublineage alongside the dominant A.Br.047 Vollum. Future work will focus on the limits of detection and will determine if this bioinformatic method can be expanded more broadly for B. anthracis or other pathogens to replace typical culture-based methods.

A.Br.047 Vollum↗

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING↗