Search NASA⌕ Search

SEARCH · Search NASA

Results for “AI Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Abrasive Waterjet Machining

The abrasive waterjet machining process was introduced in the 1980s as a new cutting tool; the process has the ability to cut almost any material. Currently, the AWJ process is used in many world-class factories, producing parts for use in daily life. A description of this process and its influencing parameters are first presented in this paper, along with process models for the AWJ tool itself and also for the jet–material interaction. The AWJ material removal process occurs through the high-velocity impact of abrasive particles, whose tips micromachine the material at the microscopic scale, with no thermal or mechanical adverse effects. The macro-characteristics of the cut surface, such as its taper, trailback, and waviness, are discussed, along with methods of improving the geometrical accuracy of the cut parts using these attributes. For example, dynamic angular compensation is used to correct for the taper and undercut in shape cutting. The surface finish is controlled by the cutting speed, hydraulic, and abrasive parameters using software and process models built into the controllers of CNC machines. In addition to shape cutting, edge trimming is presented, with a focus on the carbon fiber composites used in aircraft and automotive structures, where special AWJ tools and manipulators are used. Examples of the precision cutting of microelectronic and solar cell parts are discussed to describe the special techniques that are used, such as machine vision and vacuum-assist, which have been found to be essential to the integrity and accuracy of cut parts. The use of the AWJ machining process was extended to other applications, such as drilling, boring, milling, turning, and surface modification, which are presented in this paper as actual industrial applications. To demonstrate the versatility of the AWJ machining process, the data in this paper were selected to cover a wide range of materials, such as metal, glass, composites, and ceramics, and also a wide range of thicknesses, from 1 mm to 600 mm. The trends of Industry 4.0 and 5.0, AI, and IoT are also presented.

36 MATERIALS SCIENCE↗

Real-Time Data Display

RT-Display is a MATLAB-based data acquisition environment designed to use a variety of commercial off-the-shelf (COTS) hardware to digitize analog signals to a standard data format usable by other post-acquisition data analysis tools. This software presents the acquired data in real time using a variety of signal-processing algorithms. The acquired data is stored in a standard Operator Interactive Signal Processing Software (OISPS) data-formatted file. RT-Display is primarily configured to use the Agilent VXI (or equivalent) data acquisition boards used in such systems as MIDDAS (Multi-channel Integrated Dynamic Data Acquisition System). The software is generalized and deployable in almost any testing environment, without limitations or proprietary configuration for a specific test program or project. With the Agilent hardware configured and in place, users can start the program and, in one step, immediately begin digitizing multiple channels of data. Once the acquisition is completed, data is converted into a common binary format that also can be translated to specific formats used by external analysis software, such as OISPS and PC-Signal (product of AI Signal Research Inc.). RT-Display at the time of this reporting was certified on Agilent hardware capable of acquisition up to 196,608 samples per second. Data signals are presented to the user on-screen simultaneously for 16 channels. Each channel can be viewed individually, with a maximum capability of 160 signal channels (depending on hardware configuration). Current signal presentations include: time data, fast Fourier transforms (FFT), and power spectral density plots (PSD). Additional processing algorithms can be easily incorporated into this environment.

Pedings, Marc↗

Antiviral discovery using sparse datasets by integrating experiments, molecular simulations, and machine learning

Computational methods have demonstrated success in identifying virucidal agents, effectively contributing to the discovery of novel virucidal molecules. In this study, we developed a machine learning (ML) model, trained on a small dataset, to predict inhibitors of human enterovirus 71 (EV71), a pathological agent that causes severe disease in children and immunocompromised adults. Despite the dataset’s limitation, comprising of only 36 compounds tested, our ML framework demonstrated significant predictive capability. Notably, experimental validation revealed that five out of the eight compounds predicted by our model from the Chinese cosmetic material list exhibited virucidal activity. The inhibitor effects displayed by the main active compounds were further confirmed by molecular dynamics simulation. This underscores the potential of our AI-driven approach to bypass data constraints in identifying active molecules against viral pathogens.

60 APPLIED LIFE SCIENCES↗

MCNP Simulations of Measurement of Insulation Compaction in the Cryogenic Rocket Fuel Tanks at Kennedy Space Center by Fast/Thermal Neutron Techniques

MCNP simulations have been run to evaluate the feasibility of using a combination of fast and thermal neutrons as a nondestructive method to measure of the compaction of the perlite insulation in the liquid hydrogen and oxygen cryogenic storage tanks at John F. Kennedy Space Center (KSC). Perlite is a feldspathic volcanic rock made up of the major elements Si, AI, Na, K and 0 along with some water. When heated it expands from four to twenty times its original volume which makes it very useful for thermal insulation. The cryogenic tanks at Kennedy Space Center are spherical with outer diameters of 69-70 feet and lined with a layer of expanded perlite with thicknesses on the order of 120 cm. There is evidence that some of the perlite has compacted over time since the tanks were built 1965, affecting the thermal properties and possibly also the structural integrity of the tanks. With commercially available portable neutron generators it is possible to produce simultaneously fluxes of neutrons in two energy ranges: fast (14 Me V) and thermal (25 me V). The two energy ranges produce complementary information. Fast neutrons produce gamma rays by inelastic scattering, which is sensitive to Fe and O. Thermal neutrons produce gamma rays by prompt gamma neutron activation (PGNA) and this is sensitive to Si, Al, Na, K and H. The compaction of the perlite can be measured by the change in gamma ray signal strength which is proportional to the atomic number densities of the constituent elements. The MCNP simulations were made to determine the magnitude of this change. The tank wall was approximated by a I-dimensional slab geometry with an 11/16" outer carbon steel wall, an inner stainless wall and 120 cm thick perlite zone. Runs were made for cases with expanded perlite, compacted perlite or with various void fractions. Runs were also made to simulate the effect of adding a moderator. Tallies were made for decay-time analysis from t=0 to 10 ms; total detected gamma-rays; detected gamma-rays from thermal neutron reactions d. detected gamma-rays from non-thermal neutron reactions and total detected gamma-rays as a function of depth into the annulus volume. These indicated a number of possible independent metrics of perlite compaction. For example the count rate for perlite elements increased from 3600 to 8500 cps for an increase in perlite density from 6 lbs/lcf to 16.5 lbs/cf. Thus the MCNP simulations have confirmed the feasibility of using neutron methods to map the compaction of perlite in the walls of the cryogenic tanks.

Livingston, R. A.↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗

Digital Twin Framework for PIP-II Linac: AI-Driven Multi-Scale Modeling from Ion Source to 800 MeV

The PIP-II linac will enable >1.2 MW beam power for DUNE, requiring unprecedented operational reliability across its warm front-end (RFQ, MEBT) and five distinct SRF sections operating at 162.5/325/650 MHz. We present a comprehensive digital twin framework uniquely combining a fully differentiable fast beam transport code with neural network surrogates trained on high-fidelity PIC simulations, capturing space charge and nonlinear dynamics beyond traditional envelope codes while achieving 10⁴ speedup at <1% accuracy. End-to-end differentiability enables gradient-based optimization across 500+ parameters simultaneously previously impossible with conventional tools while the model incorporates static/dynamic errors and serves as a virtual commissioning platform for diverse hardware integration. The framework facilitates reinforcement learning for pulsed/CW mode transitions, predictive maintenance through anomaly detection, and autonomous tuning algorithm development with real-time execution capability. Validation against physics simulations shows excellent agreement for the front-end, with initial results demonstrating potential for 30% commissioning time reduction and proactive fault mitigation, providing a scalable blueprint for operating next-generation high-intensity accelerators.

Pathak, Abhishek [Fermilab] (ORCID:000000021704208↗

Digital Twin Framework for PIP-II Linac: AI-Driven Multi-Scale Modeling from Ion Source to 800 MeV

The PIP-II linac will enable >1.2 MW beam power for DUNE, requiring unprecedented operational reliability across its warm front-end (RFQ, MEBT) and five distinct SRF sections operating at 162.5/325/650 MHz. We present a comprehensive digital twin framework uniquely combining a fully differentiable fast beam transport code with neural network surrogates trained on high-fidelity PIC simulations, capturing space charge and nonlinear dynamics beyond traditional envelope codes while achieving 10⁴× speedup at <1% accuracy. End-to-end differentiability enables gradient-based optimization across 500+ parameters simultaneously—previously impossible with conventional tools—while the model incorporates static/dynamic errors and serves as a virtual commissioning platform for diverse hardware integration. The framework facilitates reinforcement learning for pulsed/CW mode transitions, predictive maintenance through anomaly detection, and autonomous tuning algorithm development with real-time execution capability. Validation against physics simulations shows excellent agreement for the front-end, with initial results demonstrating potential for 30% commissioning time reduction and proactive fault mitigation, providing a scalable blueprint for operating next-generation high-intensity accelerators.

Pathak, Abhishek [Fermilab] (ORCID:000000021704208↗

Connected and Learning Based Optimal Freight Management for Efficiency

The management of the future heterogenous fleet is a complex decision-making problem. The heterogenous fleet is emerging as decarbonization technologies are deployed by fleets toward lowering the freight operation emissions in Medium and Heavy-duty vehicles. Traditionally, in fleets characterized by a homogeneous Diesel Internal Combustion Engine (ICE) powertrain, the process of fleet planning and operational optimization unfolds sequentially without the necessity to account for powertrain and vehicle-specific characteristics during dispatch decisions. Fleets with trucks less than 5 years old tend to maintain stable vehicle efficiency with minimal operational reliability risks for fleet managers. However, the landscape changes with the incorporation of emerging powertrain technologies, which lack extensive operational data and service experiences. This includes technologies like hybrid, Electric, Fuel Cell, or alternative fuel ICE. Operational decisions for fleets featuring heterogeneous powertrain technologies and facing limited access to alternative fueling and charging stations become intricate, requiring careful consideration and optimization at each dispatch. The difference in efficiency characteristics of emerging technologies, their range limitations, and the restricted availability of charging/alternative fueling infrastructure, coupled with sensitivity to driving conditions (e.g., EV range reduction in low temperatures) and their impact on component aging (such as batteries), become pivotal factors influencing the reliable and efficient freight transportation. To make the path toward low emission freight transportation efficient and reliable, an AI-assisted fleet management software is developed in this project to help fleet managers in optimizing both adoption of emerging powertrain decarbonization, connected and automated technologies and also operating the fleet after such technologies are deployed as schematically. Freight transportation requirements are different depending on the cargos to be shipped, customer requirements and regions of operations. This further highlights the need for software and digital solutions to tailor deployment and operation of emerging powertrain, connectivity, and automation technologies toward the specific fleet operation requirements. The fleet management optimizer was also integrated with a model of the fleet to simulate the operation of the fleet over 1 year of the baseline fleet operation (250,000+ shipments) indicating the significance of day-to-day variations on emissions and energy consumption of a freight transportation fleet. The results demonstrate ≥20% improvement in freight efficiency in terms of WTW CO2 per ton-mile of cargo shipments while all fleet operation constraints are enforced, and the cost (CapEx and OpEx) is minimized.

33 ADVANCED PROPULSION SYSTEMS↗

Data and Reasoning Fabric (DRF) Phase I & Phase II Report

A Data & Reasoning Fabric (DRF) is envisioned to enable the full potential of advanced air mobility by providing all data and reasoning where they are needed. The DRF marketplace is based on an open foundational ecosystem of data and reasoning exchange between the many systems that must seamlessly interplay to manage the envisioned highly complex and dense airspace operations. DRF activities will identify, test and - as needed - research and develop critical core technologies, and collaboratively test these technologies, open standards and architectures, and the integrated framework with end-users so as to deliver reference designs and development environments that catalyze broad private and public sector buy-in and self-sustaining development of it and associated standards.

UAM↗

Benefits of Ka-band GaN MMIC High Power Amplifiers With Wide Bandwidth and High Spectral/Power Added Efficiencies for Cognitive Radio Platforms

A cognitive radio on a future NASA near-Earth spacecraft will be capable of sensing its environment and dynamically adapting its operating parameters to provide the desired SATCOM service to the mission. A key component that can enable this type of operation is a high-power amplifier (HPA) that resides on the radio platform. In this paper, we present the RF performance characteristics of a Ka-band gallium nitride (GaN) monolithic microwave integrated circuit (MMIC) based HPA for cognitive radio platforms. These characteristics include the output power, gain, power added efficiency (PAE), RMS error vector magnitude (EVM), spectral efficiency, 3rd-order intermodulation distortion (IMD) products, spectrum, spectral regrowth, noise figure (NF), and phase noise. The data presented indicates that the HPA meets NTIA, military, and commercial spectral mask requirements. In addition, we discuss the benefits offered by the above performance characteristics toward the design and implementation of a cognitive radio platform. Furthermore, as examples, we discuss three potential use cases that apply artificial intelligence (AI) and machine learning (ML) techniques and exploit the performance characteristics discussed above to provide a knowledge-based cognitive radio platform design for SATCOM. Thus, cognitive radios with performance flexibility can enable roaming and provide seamless interoperability autonomously in the future between NASA, commercial, and other space networks owned by U.S. government agencies.

Gallium nitride↗

Benefits of Ka-band GaN MMIC High Power Amplifiers With Wide Bandwidth and High Spectral/Power Added Efficiencies for Cognitive Radio Platforms

A cognitive radio on a future NASA near-Earth spacecraft will be capable of sensing its environment and dynamically adapting its operating parameters to provide the desired SATCOM service to the mission. A key component that can enable this type of operation is a high-power amplifier (HPA) that resides on the radio platform. In this report, we present the RF performance characteristics of a Ka-band gallium nitride (GaN) monolithic microwave integrated circuit (MMIC) based HPA for cognitive radio platforms. These characteristics include the output power, gain, power added efficiency (PAE), RMS error vector magnitude (EVM), spectral efficiency, 3rdorder intermodulation distortion (IMD) products, spectrum, spectral regrowth, noise figure (NF), phase noise, and group delay. The data presented indicates that the HPA meets NTIA, military, and commercial spectral mask requirements. In addition, we discuss the benefits offered by the above performance characteristics toward the design and implementation of a cognitive radio platform. Furthermore, as examples, we discuss three potential use cases that apply artificial intelligence (AI) and machine learning (ML) techniques and exploit the performance characteristics discussed above to provide a knowledge-based cognitive radio platform design for SATCOM. Thus, cognitive radios with performance flexibility can enable roaming and provide seamless interoperability autonomously in the future between NASA, commercial, and other space networks owned by U.S. government agencies.

Gallium nitride↗

Expanding SPoRT RGBs and Machine Learning Techniques to Enhance Air Quality Monitoring in Southern Asia

Air pollution poses significant environmental, public health, and societal concerns in the Hindu Kush Himalaya (HKH) region of south-central Asia, notably during the dry monsoon months (~November to May). Key contributors to poor air quality include dust from the Middle East and western India, persistent nocturnal fog/smog, and biomass burning. To address this issue, we established a robust air quality and chemistry observation and modeling product suite utilizing multi-spectral red-green-blue (RGB) composite satellite products from Korea’s GEO-KOMPSAT-2A satellite, the Hybrid Single-Particle Lagrangian Integrated Trajectory (HYSPLIT) model for dust transport forecasts, and the Weather Research and Forecasting coupled with Chemistry (WRF-Chem) model to predict aerosols and chemical species concentrations. Our team employed similar RGB recipes transitioned by the NASA Short-term Prediction Research and Transition (SPoRT) Center for the GOES-R era products over the Western Hemisphere, with significant success in depicting dust and nocturnal fog / low clouds, and to a lesser extent smoke and fire hot spots. We will extend these capabilities for the HKH region by applying an artificial intelligence (AI) model that objectively identifies dust from multi-spectral satellite data over the Southwestern United States. The AI model will be calibrated for automated dust detection over HKH from GEO-KOMPSAT-2A satellite data, with a goal of developing a similar AI model for objectively identifying smoke as well. The ultimate goal of this effort is to enhance dust and smoke predictions in the region by establishing improved emission initializations in HYSPLIT and/or WRF-Chem through the automated AI detection of dust and smoke.

Jonathan L. Case↗

A Thermal Expert System (TEXSYS) development overview - AI-based control of a Space Station prototype thermal bus

A knowledge-based control system for real-time control and fault detection, isolation and recovery (FDIR) of a prototype two-phase Space Station Freedom external thermal control system (TCS) is discussed in this paper. The Thermal Expert System (TEXSYS) has been demonstrated in recent tests to be capable of both fault anticipation and detection and real-time control of the thermal bus. Performance requirements were achieved by using a symbolic control approach, layering model-based expert system software on a conventional numerical data acquisition and control system. The model-based capabilities of TEXSYS were shown to be advantageous during software development and testing. One representative example is given from on-line TCS tests of TEXSYS. The integration and testing of TEXSYS with a live TCS testbed provides some insight on the use of formal software design, development and documentation methodologies to qualify knowledge-based systems for on-line or flight applications.

Glass, B. J.↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

A Data & Reasoning Fabric to Enable Advanced Air Mobility

A Data & Reasoning Fabric (DRF) is envisioned to enable the full potential of advanced air mobility by providing all data and reasoning where they are needed. The DRF marketplace is based on an open foundational ecosystem of data and reasoning exchange between the many systems that must seamlessly interplay to manage the envisioned highly complex and dense airspace operations. DRF activities will identify, test and - as needed - research and develop critical core technologies, and collaboratively test these technologies, open standards and architectures, and the integrated framework with end-users so as to deliver reference designs and development environments that catalyze broad private and public sector buy-in and self-sustaining development of it and associated standards.

Urban Air Mobility↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Data Centers and Digital Assurance Workshop 3 – Mitigations for Digital Assurance Risks

The third session of the TADA (Technical Assistance for Digital Assurance) Data Centers Cohort, held on November 18, 2025, focused on developing mitigation strategies for digital assurance risks identified in previous workshops. Hosted by Idaho National Laboratory (INL) and ScottMadden, the session emphasized the application of Cyber-Informed Engineering (CIE) to data center infrastructure, particularly at the utility–data center interface. Participants revisited and ranked key digital assurance risks, including architecture and interface weaknesses, governance gaps, and AI-enabled threats. The workshop introduced the 12 principles of CIE, advocating for consequence-focused design, engineered controls, and secure information architecture to proactively reduce cyber-physical vulnerabilities. These principles were applied to critical data center systems such as power distribution, UPS, cooling, SCADA/BMS, and grid-forming batteries. The session also addressed governance challenges at the interconnection boundary, highlighting the need for clear roles in telemetry sharing, firmware management, and trip settings. Special attention was given to emerging risks from behind-the-meter (BTM) generation, including reverse-power flow and the integration of small modular reactors (SMRs), which shift data centers from large loads to complex generation nodes. Participants explored how interconnection agreements can serve as enforceable instruments for digital assurance, and reviewed gaps in current standards such as NERC CIP, IEC 62443, and IEEE 1547. The workshop concluded with pathways to standardization, including model agreement language, state-level programs, and expanded NERC guidance. INL also presented tools and frameworks for secure procurement and supplier risk management, reinforcing the need for integrated engineering and policy solutions to secure the evolving data center–grid ecosystem. Session 3 of 3.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗