Search NASA⌕ Search

SEARCH · Search NASA

Results for “Level of Automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Optimizing Cryo-Focused Pyrolysis GC/MS for Tracing Soil Organic Matter Across Diverse Ecosystems

The cycling of organic matter in terrestrial soils and sediments is central to a range of biogeochemical processes that regulate nutrient cycling, crop productivity, trace gas emissions, and contaminant transport. Pyrolysis-gas chromatography/mass spectrometry (py-GC/MS) is a powerful tool for characterizing bulk soil organic matter (SOM) at the molecular level. In this study, we used a cryo-focused py-GC/MS system to analyze soil samples from seven diverse ecosystems: vernal pool, prairie pothole, temperate forest, tropical forest, tundra, wildfire-affected boreal forest, and grassland. We addressed a key bottleneck in molecular-level SOM characterization by developing an automated data analysis pipeline to optimize py-GC/MS and complementary evolved gas analysis/mass spectrometry (EGA/MS) methods, incorporating advanced tools for peak deconvolution, developing a custom compound class library, and implementing fragmentation spectrum-based molecular networking for the first time. This improved workflow was applied to soil samples from all seven ecosystems, including multiple depths and density fractions. Our findings demonstrate that ecosystem type plays a dominant role in shaping compositional differences in SOM. We also identified trends in the source of SOM compounds (e.g., microbial vs plantderived) across soil depth and density fractions, which are critical for understanding persistence and turnover of SOM. Our molecular networking analysis indicated that although many compounds are widespread across ecosystems, others are restricted to specific environments, such as wetlands. This underscores the utility of molecular-level data in elucidating the complexity of SOM composition and the environmental drivers that shape it. Such molecular-level insights can deepen our knowledge of biogeochemical SOM cycles.

54 ENVIRONMENTAL SCIENCES↗

Deformable phrase level attention: A flexible approach for improving AI based medical coding

Objective: Improving the AI-driven automated medical encoding of clinical text plays a vital role in gathering information on the occurrence of diseases to improve population-level health. This work presents a novel attention mechanism designed to enhance text classification models and ensure appropriate classification of medical concepts in unstructured electronic health records. Materials and Methods: We developed a deformable, phrase-level attention mechanism to identify important lexical word-level and contextual phrase-level information from clinical text documents. We evaluated conventional and transformer-based deep learning models that we extended with our attention mechanism on the extraction of critical cancer information (e.g., site, subsite, laterality, histology, behavior) from 629,908 electronic pathology reports and on the automated medical encoding of 52,722 hospital discharge summaries. Results: Transformer-based models with the deformable, phrase-level attention mechanism achieved the best performance on the extraction of critical cancer information from pathology reports. Conventional- and transformer-based models show similar or better performance than their baseline counterparts on the automated medical encoding of clinical documents. Discussion: The addition of phrase-level information allowed models extended with our proposed method to outperform standard word-level attention. Our method showed favorable properties for the real-world application in terms of model robustness and phenotyping. These results indicate that our method is promising for automated data harmonization for common data models. Conclusion: This work proposes a novel deformable, phrase-level attention mechanism that enhances text classification models in the extraction of medical concepts from clinical text documents. We demonstrate strong performances on two clinical text datasets and showcase real-world deployability of our method.

Automated medical encoding↗

Accuracy Enhancement of Nuclear Power Plant Simulators Utilizing High Accuracy Simulation Predictions

More recently, reactor core simulators for core designs associated with commercial nuclear power plants that utilize what is believed to be higher fidelity models have been developed. Features such as neutronics models that utilize transport equation solvers with fine spatial meshes and many energy-groups, thermal-hydraulic models that utilize sub-channel solvers with fine spatial mesh and capable of treating a wide range of fluid conditions, and fuel-coolant chemistry interaction models capable of treating CRUD deposition are to be found in these higher fidelity core simulators. These reactor core simulators require access to higher performance computers, characterized by many processors, cores and large memory. So associated with utilization of these simulators is access to high performance computers and ability to accommodate in one’s workflow longer execution times. By contrast, currently used core simulators by the nuclear industry can execute on engineering workstations and have execution times of seconds to minutes. The desirability for having short execution times is not only desired for support of time critical tasks but supports the mental process of decision making by engineers. The goal of the work reported upon here has the objective of retaining the fidelity of higher fidelity models while retaining the ability to utilize engineering workstations. Beyond the core simulator goal, additional goals of this work include incorporating the just described core simulator capability into a Nuclear Steam Supply System (NSSS) simulator, and to incorporate the resulting capability into an environment supportive of design and operational decision making associated with nuclear power stations. The model selected for the core neutronics model is the NESTLE code, for the core thermal-hydraulic model is the CTF code utilizing coarse mesh, and for the NSSS model is the RELAP5-3D code. WSC’s proprietary 3KEYMASTERTM platform is being used to provide software coupling, user interface, visualization, and reporting. The NESTLE core neutronics simulator was first integrated with the CTF core thermal-hydraulic simulator using CTF developed communication commands which are also used for CTF to communicate with RELAP5-3D under WSC’s proprietary 3KEYMASTERTM platform. To assure NESTLE prediction consistency with higher fidelity core neutronic simulators, buffer codes have been created to automatically generate from output files written by the VERA core simulator the NESTLE nodal neutronic parameter’ library, geometry, and pin-power reconstruction input files, thereby avoiding a number of challenges associated with utilizing lattice physics codes and providing consistency with VERA predictions. To treat absorber rod effects a multi-set library is utilized, where a set refers to a specific absorber rod fully inserted pattern. A coarse spatial mesh CTF model was developed with features added that support using CTF as envisioned in the engineering quality simulator. A hybrid meshing approach was implemented to allow for automated construction of models with mixed levels of refinement. Specifically, a core model could resolve some assemblies at a nodal level (4 subchannels per assembly) and others at a pin-resolution (one subchannel per coolant subchannel in the assembly). The intention is that this will allow for better resolution of limiting conditions such as DNBR and PCT, which are based on local rod and subchannel conditions. Further development was done of features that enhance the capabilities for the envisioned engineering quality simulator that has been developed, but now for RELAP-3D. The RELAP5-3D code development includes ability to model more than 999 components and the addition of the cross-channels turbulence mixing model and the void drift model that are implemented in CTF, aiming to achieve closer prediction agreement of the two codes for transient simulations, specifically, more accurate matches of the overall mass, momentum, and energy exchanges of both the liquid and gas phases between the neighboring core assemblies. Graphics were also developed for the Instructor Station for this project under WSC’s proprietary 3KEYMASTERTM platform to facilitate design and operational decision making.

42 ENGINEERING↗

Assembly of CMS Endcap MIP Timing Detector Module at FNAL

The High-Luminosity LHC (HL-LHC) will enable a more detailed exploration of new phenomena thanks to an anticipated increase in collisions where pileup is expected to reach approximately 200 simultaneous interactions. Many CMS systems will be significantly upgraded to prepare for this new era, including the MIP Timing Detector (MTD) project. The MTD is designed to mitigate the effect of pileup and is set to provide a timestamp accurate to 30 ~ 40 picoseconds for every event, ensuring sustained detector performance at HL-LHC. The MTD is divided into two sections, Barrel Timing Layer (BTL) and Endcap Timing Layer (ETL) which utilize different sensor and ASIC technologies due to the difference in active surfaces, irradiation conditions, and installation schedules. The ETL, composed of two double-sided disks, employs the Low Gain Avalanche Detector (LGAD) sensor and the Endcap Timing Readout Chip (ETROC). More than 8,000 modules, each consisting of four LGAD sensors and ETROCs are required for the ETL detector. These modules will be assembled using an automated robotic gantry that guarantees precise placement at a level of 10 micrometers. In addition, the full assembly of ETL modules includes film application with the jig, wire-bonding, encapsulation with the automated dispensing robot for protecting the wire-bonding, and film curing with a vacuum oven. This talk reports on the successfully completed throughput test with mockup components using the gantry and the successful assembly of real functional modules for beam tests at CERN and FNAL, including the first official ETL module.

Apresyan, Artur↗

A Sensitivity-driven Wide Area Protection (SWAP) Coordination Tool for High Penetration of Inverter-based Resources (IBR)

Traditionally, power system generation sources have been composed of synchronous generators, of which the fault current behavior is understood with minimal differences between generation size and types due to the physics of their construction. Present protection schemes and modeling methods are based upon these understood characteristics. Most renewable generation is composed of inverter-based resources (IBR), in which fault current is determined by switching control software and hardware limitations, each of which can vary between manufacturers and even between models of the same manufacturer. The resulting fault current is low in magnitude, low in negative-sequence current, unpredictable phase angles, and is a challenge to model. These characteristics also result in a challenge to traditional protection schemes and fault simulation software. To address several of these concerns, the project has the following goals: 1. Improve IBR models: Improve IBR models used in short circuit (SC) programs to accurately capture the response of IBRs at the bulk power system (BPS) level for fault and protection studies. 2. Develop automation tool: Develop an automation tool that allows engineers to identify protection coordination and sensitivity issues by performing SC and protection coordination studies in a high IBR-penetrated grid by applying variations to the IBR models, faults, contingencies, etc. 3. Develop schemes: Develop new protection mitigation solution schemes that complement the existing protection systems to ensure safe operation of the BPS with higher IBR penetration levels. The project team did not achieve this final goal, as the Department of Energy (DOE) stopped the project early due to changes in DOE funding priorities. The termination notice came at the beginning of the final project phase, while the team was identifying and beginning to investigate protection issues. It should be noted that the team discussed a 100% penetration scenario. However, this scenario would require the use of grid-forming IBR models that are not presently available. Since developing these models requires additional effort, the 100% penetration scenario was not pursued during this project. In the future, developing the methodology and models for the 100% scenario could benefit the industry.

14 SOLAR ENERGY↗

Towards Generalizable and Efficient Circuit Topology Design: A Graph-Transformer-based Surrogate Model with Curriculum Learning

Unlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. To address the cold start problem when scaling GTN to large-scale circuits, we further introduce a curriculum learning strategy that progressively trains GTN from small-scale to large-scale circuits. This approach enables the model to first learn fundamental physical principles from simpler topologies and gradually adapt to complex configurations, effectively bridging the circuit complexity gap and improving prediction accuracy. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%. With curriculum learning, the enhanced GTN achieves a 51% improvement for performance prediction of large-scale circuits compared to the GTN model without this strategy. These advancements establish GTN as a scalable framework for automated analog circuit design across varying circuit complexity levels.

Lu, Haoshu [New Jersey Institute of Technology (NJ↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

RC-SFA Data Management Templates and Guidance for Standardized, Reusable AI-Ready Data Packages

This data package provides templates and supporting documentation developed by the River Corridor Science Focus Area (RC-SFA; https://www.pnnl.gov/projects/river-corridor) to communicate its approach to managing and publishing AI-ready data. The package is intended to help data users and data producers understand the structures, metadata practices, and quality-control approaches that support consistent, reusable, and machine-actionable data products across RC-SFA studies. Rather than focusing on a single experimental dataset, this package documents the data management framework used to make RC-SFA data easier to find, ingest, navigate, and interpret. The materials in this package reflect RC-SFA practices for standardized data package organization, including the use of a human- and machine-readable README, file-level metadata, data dictionaries, descriptive file naming, method identifiers, and automated and review-based quality assurance procedures. Together, these components illustrate how RC-SFA extends FAIR data principles toward AI-readiness by prioritizing deep metadata, consistency across data packages, and support for informed downstream reuse by both humans and computational tools. This dataset is comprised of (1) readme; (2) presentation slides with an overview of RC-SFA approach and guidance; (3) document of RC-SFA best practices; (4) data dictionary (dd); (5) file level metadata (flmd); and a subfolder containing templates for dd and flmd. All files are .csv and .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

AI-readiness↗

Spatial Proteomics towards cellular Resolution

Introduction: Spatial biology is an emerging interdisciplinary field facilitating biological discoveries through the use of spatial omics technologies. Recent advancements in spatial transcriptomics, spatial genomics (e.g. genetic mutations and epigenetic marks), multiplexed immunofluorescence, and spatial metabolomics/lipidomics have enabled high-resolution spatial profiling of gene expression, genetic variation, protein expression, and metabolites/lipids profiles in tissue. These developments contribute to a deeper understanding of the spatial organization within tissue microenvironments at the molecular level. Areas covered: This report provides an overview of the untargeted, bottom-up mass spectrometry (MS)-based spatial proteomics workflow. It highlights recent progress in tissue dissection, sample processing, bioinformatics, and liquid chromatography (LC)-MS technologies that are advancing spatial proteomics toward cellular resolution. Expert opinion: The field of untargeted MS-based spatial proteomics is rapidly evolving and holds great promise. To fully realize the potential of spatial proteomics, it is critical to advance data analysis and develop automated and intelligent tissue dissection at the cellular or subcellular level, along with high-throughput LC-MS analyses of thousands of samples. In conclusion, achieving these goals will necessitate significant advancements in tissue dissection technologies, LC-MS instrumentation, and computational tools.

59 BASIC BIOLOGICAL SCIENCES↗

Field Validation of a Grid-Interactive Efficient Building Software Solution

The U.S. General Services Administration's (GSA's) Green Proving Ground (GPG) program, in partnership with the National Laboratory of the Rockies (NLR), completed a field study of a Grid-Interactive Efficient Buildings (GEB) software solution. The study focused on a single testbed facility to test the GEB functionality of the software solution, along with other features. The testbed facility - a courthouse - is a common building type in GSA's vast building portfolio, offering potentially impactful findings on a scalable level. The study evaluated Prescriptive Data's technology, Nantum OS, a connected building operating system ("GEB Solution") which aggregates multiple sources of previously siloed building data and combines that data with external sources, such as weather information or utility signals, into a single integrated platform. A GEB Solution is a type of Energy Management Information System (EMIS). EMIS is defined as a system of devices, data services, and software applications that communicates with any building system or third-party data source to aggregate and transform data into new capabilities to aid in the optimization of energy use at the building, campus, or agency level. This specific GEB Solution is an EMIS with ASO, automated system optimization, offering supervisory control of certain aspects of the Building Automation System (BAS). Multiple features were evaluated including, but not limited to, Continuous Demand Management to avoid setting new monthly kilowatt (kW) peaks, energy efficiency for reduction of kilowatt hours (kWh) and natural gas consumption, and automated demand response (ADR) for purposes of lowering demand during a utility called Demand Response (DR) event. The testbed facility was the Foley Federal Building and US Courthouse ("Foley Federal Building") located in Las Vegas, NV. This is a 209,496 sq. ft. building constructed in the 1960s with major renovations in 2004. The facility was a good candidate due to the large prevalence of office and courthouse spaces in the GSA portfolio of buildings. It also has many features which allow integration into and control of the building and a strong facilities team to assist with the study. Quantitative and qualitative performance objectives were developed using GSA's GPG GEB project template along with input from the vendor and building facility staff; these are outlined in Table 1. The quantitative performance objectives focused on continuous demand management, energy efficiency, and automated demand response. The qualitative performance objectives focused on the ease of installation and commissioning as well as the operability of the GEB solution. Other performance metrics that are reported on include carbon reduction, cost effectiveness, and occupant acceptance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The Art of Automation: Translating Electron Microscopy Workflows Into Automated Processes

Acquiring data using a scanning transmission electron microscope (STEM) is a complex, multi-step process. The intricacy of the process depends on the type of sample, composition of the material, desired results of the experiment, resolution requirement and other experimental factors. Each experiment presents unique complications, such as sample drift and contamination, that the microscopist must consider when acquiring data. All these challenges are handled fluidly and expertly by experienced microscopists, but to reach new levels of innovation in material development, including greater reproducibility, throughput, and precision, the automation of these workflows is essential. The initial phase of this work involved translating intuition-based workflows into discrete, programmable steps. Some common key stages in STEM workflows are the initial tuning, scanning the sample for areas of interest, and then acquiring the data. Each stage can be broken further into specific parameter adjustments, such as aberration correction and dwell time optimization, depending on the experiment. When deconstructing various experiments each step was assessed for automation feasibility based on the amount of real time operator decisions. There are steps that lend themselves to automation more readily than others, such as course focusing and sample screening, but there is potential for full automation of all stages with time. As an initial step, an automated montage routine was developed, allowing for the efficient acquisition of large portions of the sample without requiring continuous intervention from the operator. The automation of this small process of the procedure demonstrates the value of this capability. A major challenge in automation arises from discrepancies between commanded, reported and actual stage movements. Using systematic tests, stage movement was quantified. This error can be corrected algorithmically for more accurate workflows in the future. Expanding automation capabilities would result in larger, more efficient data acquisition which allows for more robust statistical analysis. Additionally, this work lays the groundwork for a closed loop system where machine learning algorithms would intake automatically acquired data and make real time decisions. By progressively automating this instrument, this work establishes the foundation for fully automated experimentation in transmission electron microscopy.

97 MATHEMATICS AND COMPUTING↗

Use of Fisher's Ratio assisted multivariate curve resolution- alternating least squares for discovery-based analysis using ultrahigh pressure liquid chromatography-high resolution mass spectrometry

Non-targeted analysis of complex chemical mixtures can be difficult considering the convoluted nature of the matrix and the potential unknown chemical differences between samples or classes of samples. Ultrahigh pressure liquid chromatography coupled to quadrupole time-of-flight mass spectrometry (UHPLC-QTOF) is an ideal technique to probe chemical differences for a wide variety of samples. While UHPLC-QTOF can discover minute chemical differences down to low part per billion (ppb) concentrations with a high degree of confidence, the application of high-resolution mass spectrometry can yield massive amounts of information (∼ 10 gb per sample) that cannot be analyzed manually. Therefore, the application of chemometric techniques is mandatory for the interrogation of complex samples. Fisher's ratio (FR) assisted multivariate curve resolution-alternating least squares (MCR-ALS) was used to the discover and identify the chemical differences between two classes of materials: 1) a pond water matrix and 2) the matrix spiked with a pharmaceutical standard mix containing 17 compounds. Thirteen of the seventeen spiked compounds were discovered using FR analysis, and then five were successfully deconvoluted using MCR-ALS wherein the number of curves chosen were automatically determined using singular value decomposition (SVD). In conclusion, the use of an automated FR assisted MCR-ALS will aid in discovering trace levels of chemical components without the need for the researcher to provide potentially biased input which will aid in non-targeted workflow.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Detecting Important Drivers of Gridded Population Modeling With Machine Learning

High-resolution population datasets have been lever-aged across a broad swath of domains, such as climate change, public policy, humanitarian aid, and rescue operations, among others. Machine learning methods were adopted to generate high-resolution or gridded population estimates by using various geospatial input features such as buildings, roads, and nighttime lights. In this study, we evaluate the importance of population features using Random Forest models across three levels of analysis, utilizing permutation measures. Our research aims to address key questions to enhance our understanding of high-resolution population modeling, such as: Are certain features globally (10 countries collectively) more important than others? Do optimal features vary by country? Within each country, do feature importance differ across administrative units? What similarities exist in feature importance at the global, country, and administrative unit levels? To answer these questions, we leverage the Kneedle algorithm to automate the selection of optimum features. We find that there are patterns displayed by features across spatial boundaries, evidenced by the same feature being the most important indicator of population across 7 of the 10 countries modeled. Our findings indicate that while important features may vary across geographies, certain features consistently hold greater importance than others agnostic of geography.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Architectural Approaches for Integrating ADMS and DERMS: Challenges, Comparisons, and Real-World Use Cases

The electrical distribution landscape is rapidly transforming due to the proliferation of distributed energy resources (DERs) such as solar panels, wind turbines, battery storage systems, combined heat and power units, and electric vehicles, introducing variability and uncontrollability that traditional grid operators are ill-equipped to manage. This transformation is further accelerated by advancements in Information and Communication Technology infrastructure that connects control centers with end devices, demanding automation and a deeper understanding of new technologies by utility personnel. Advanced grid control techniques using system-level optimization, Artificial Intelligence, and Machine Learning at the enterprise level and distributed level are evolving to address these issues. There is also an opportunity to utilize the enormous data created by these new DER technologies in the grid. Advanced Distribution Management Systems (ADMS) and Distributed Energy Resource Management Systems (DERMS) are critical in addressing these challenges by automating grid operations and enhancing reliability. Given the relatively recent development of ADMS and DERMS, and the still relatively low level of ADMS and DERMS deployment in the industry, there is a notable deficiency in the comprehensive understanding of the challenges and benefits associated with these new technologies, especially with their complementary natures and integration architectures. This paper aims to bridge the knowledge gap in ADMS and DERMS integration, presenting three distinct integration architectures currently available, and discussing the challenges and benefits of each architecture to guide utilities, industry professionals, and researchers in optimizing grid management and decision-making processes for a resilient and efficient energy future.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Evaluating the feasibility of LA-ICP-TOF-MS for the analysis of environmental particle collections

Laser ablation-inductively coupled plasma-time-of-flight-mass spectrometry (LA-ICP-TOF-MS) was employed to rapidly analyze environmental particle samples collected using aerosol contaminate extractors (ACE). The ACE particle collectors were placed at various distances (0.5, 1.3, and 4.5 km) from a source that released Ru-bearing particles. Samples for measurement were then generated (as sub-samples) from the ACE collection plates via particle “lift off” with gunshot residue (GSR) tabs. The LA-ICP-TOF-MS method was employed such that 10+ samples could be analyzed in a single unattended analytical session. A 3 × 1 mm area of individual GSR tab samples were analyzed in less than 30 minutes. This provided spatially resolved elemental and isotopic measurements of the particulate content and confirmed the presence of Ru-bearing particles within the complex background environmental particle loading. As anticipated, measurements showed collectors closest to the source had the highest concentration of the released Ru-bearing particles, while all collectors, regardless of distance, contained similar levels of background particles (e.g., Fe and Sr). Sequential scanning electron microscopy – automated particle analysis (SEM-APA) and LA-ICP-TOF-MS analysis was employed for method validation and a demonstration of the multi-modal approach. The same 2-dimensional region was analyzed by both methods and the particles identified via SEM-APA were also detected using LA-ICP-TOF-MS, with 100% accuracy. Overall, LA-ICP-TOF-MS demonstrated its utility for rapid elemental and isotopic particle analysis from environmental air samples.

Manard, Benjamin T. [Oak Ridge National Laboratory↗

TSDC: Transportation Secure Data Center: Real-World Data for Planning, Modeling, and Analysis

The Transportation Secure Data Center is a centralized repository for high-resolution transportation data from hundreds of travel and transit surveys and studies. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. It houses surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Meanwhile, the Livewire Data Platform empowers research, industry, and academic partners to easily and securely preserve, maintain, share, discover, and gain access to transportation and mobility data. Livewire accommodates a range of datasets, including behavioral, experimental, model, analytical, and raw data at the vehicle, traveler, and system levels. Datasets support mobility research and planning spanning urban science, connected and automated vehicles, fueling and charging infrastructure, mobility decision science, multimodal transportation, vehicle efficiency, and more.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Peak2Patch: High-Fidelity Functional Group Identification through Attention-Based Fusion of Infrared and Mass Spectra

Identifying molecular structure based on spectroscopic readings is a key task in a variety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless require expert-level knowledge to decode. Machine learning has emerged as a potential solution for automating structure prediction from chemical spectra; however, current approaches generally focus on single sensor modalities, neglecting to leverage the complementary information contained within differing spectra. In this paper, we introduce Peak2Patch, a novel approach to fusion-enhanced prediction of functional groups from IR and mass spectra. First, we perform a detailed comparison of backbone networks for encoding both sparse mass spectra and dense IR spectra and demonstrate the superior performance of transformer neural networks over current state-of-the-art convolutional neural networks. Second, we evaluate three broad categories of fusion: early (raw feature), middle (deep feature), and late (decision) fusion, demonstrating the potential of a deep feature fusion-based approach. Lastly, we present Peak2Patch, our attention-based fusion scheme, which leverages cross-attention to mix features between encoded tokens of the two modalities. We validate our approach on a publicly available multimodal spectroscopic data set of 790k simulated molecules, demonstrating a large improvement in functional group prediction over both the previous state-of-the-art and our own strong single-modal baselines.

Jacobson, Philip [Sandia National Laboratories (SN↗