Large Language Model (LLM) Driven Document Clustering: Improving Real time Security Intelligence Extraction and Threat Analysis
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The File Transfer Service (FTS3) is a distributed data movement service developed at CERN and widely used to transfer data across the Worldwide LHC Computing Grid (WLCG). At Fermilab, FTS3 supports data transfers for multiple experiments, including Intensity Frontier experiments such as DUNE, enabling reliable data movement between WebDAV endpoints in Europe and the Americas. At CHEP 2021, we reported on the initial containerized deployment of FTS3 on OKD, the community Kubernetes distribution of Red Hat OpenShift. In this work, we present the subsequent evolution of this deployment, focusing on new operational capabilities introduced to improve scalability, robustness, and long-term maintainability. We describe the adoption of more secure and reproducible container build workflows, the integration of DevOps-driven operational practices, and enhancements in monitoring and automation. A key new result is the introduction of horizontal scaling and elastic resource management, allowing FTS3 components to dynamically adapt to workload variations while maintaining service reliability. We also discuss improvements in fault tolerance and operational procedures derived from production experience. Finally, we summarize lessons learned from operating FTS3 as a Kubernetes-native service and outline how these developments have improved the resilience and efficiency of data movement operations at Fermilab.
Engineering design in the field of industrial engineering, such as designing automated factories or warehouses, is critical for the effective operation of facilities. Any design flaws introduced early can result in significant capital expenses to correct. However, early-stage engineering design is inherently complex. The systems are not yet built, requiring designers to integrate various aspects, including digital engineering and cybersecurity, to support virtual representations throughout the design process. In this study, we propose an approach to integrate Cyber-Informed Engineering (CIE) principles into model-based systems engineering (MBSE). This approach facilitates the development of a digital thread for engineering systems, ensuring secure digital artifacts in the design of industrial engineering systems.
Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.
A multidisciplinary team at Argonne National Laboratory explores the application of advanced technologies to enhance knowledge transfer and retention within the nuclear safeguards domain. Specifically, it examines the feasibility of leveraging secure large language models (LLMs) to streamline the creation of e-learning modules for the U.S. National Nuclear Security Administration (NNSA) Office of International Nuclear Safeguards (NA-241). The initiative addresses the critical need for preserving institutional memory and accelerating skill development amidst the imminent retirement of senior professionals in the field in addition to supporting good knowledge management practices. The project integrates instructional design theory with cutting-edge AI technologies to transform curated document sets from the Safeguards Knowledge Repository (SKR) into modular online courses. By automating the generation of learning objectives and instructional content, the effort aims to reduce manual effort while maintaining high-quality educational outcomes. A limited measure of human supervision, however, ensures accuracy, relevance, and alignment with NNSA’s strategic priorities. Key findings highlight the potential of AI-assisted course generation to support safeguards professionals by creating structured, interactive learning experiences. The report underscores the importance of SME validation to address limitations in AI-generated content, such as terminology errors and gaps in coverage. Recommendations include adopting a structured workflow combining LLM acceleration with expert oversight to ensure accuracy, usability, and alignment with learner needs. This work demonstrates Argonne’s commitment to advancing national security and scientific excellence through innovative knowledge management solutions.
Join SoCalGas and the National Laboratory of the Rockies (NLR) for an insightful webinar on securing operational technology and energy systems. Los Angeles has the second-largest metro by population in the nation, making it vital to protect the energy infrastructure that powers day-to-day life. However, cybersecurity for energy systems is an immense challenge due to an increasing number of interconnected devices and stakeholders. While there was traditionally a limited need to secure energy infrastructure, the grid is only becoming smarter and more software-defined. Amid aging infrastructure and evolving cyber threats, the need to ensure the security of energy is at an all-time high. Together, SoCalGas and NLR are assessing the current state-of-the-art of the energy ecosystem to understand gaps in current best practices and technologies relevant to powering the city of Los Angeles.
With the conclusion of the Laboratory Directed Research and Development (LDRD) project on Provable Security and Resilience (PSaR) in Critical Infrastructure, we present forward-looking technical concepts and strategies that build on the project’s outcomes and INL’s long-standing expertise in infrastructure protection. The challenge is to protect critical infrastructure and functions much more efficiently at scale than capable adversaries can attack at scale. After summarizing progress and ongoing work we’ll discuss what are the challenges that remain and what are new/emerging technologies, strategies, and processes to meet those challenges. Finally, we’ll layout concepts that integrate with other protection work in the coming year and beyond. For example, building secure function-specific platforms based on the seL4 microkernel, and considering the successes of Cyber-Informed Engineering as a model for engage, collaboration, and adoption. We look forward to your feedback and collaboration as we refine and expand this vision.
We present a mathematical model of the dynamics of Bacillus anthracis bacteria within the lymph nodes and blood of a host, following inhalation of an initial dose of spores. We also incorporate the dynamics of protective antigen, which is the binding component of the anthrax toxin produced by the bacteria. The model offers a mechanistic description of the early infection dynamics of inhalational anthrax, while its stochastic nature allows us to study the probabilities of different outcomes (for example, how likely it is that the infection will be cleared for a given inhaled dose of spores) in order to explain dose-response data for inhalational anthrax. The model is calibrated via a Bayesian approach, using in vivo data from New Zealand white rabbit and guinea pig infection studies, enabling within-host parameters to be estimated. We also leverage incubation-period data from the Sverdlovsk 1979 anthrax outbreak to show that the model can accurately describe human time-to-symptoms data under reasonable parameter regimes. Finally, we derive a simple approximate formula for the probability of symptom onset before time t, assuming that the number of inhaled spores has a Poisson distribution.
In September 2011, the federal government announced a “call to action” to create a mechanism to enable utility consumers to download their energy usage history from their utilities’ secure websites in a standardized electronic format. That effort was christened the Green Button initiative, drawing on the success of the Blue Button initiative delivered by the US Department of Defense in 2009 as a secure, online access mechanism to download patient health information through the TriCare online portal.
In 2018, the Advanced Research Project Agency – Energy (ARPA-E) launched the Grid Optimization Competition (GOC) [1], a series of competitive challenges intended to accelerate innovation in decision support software used to schedule power grid operations, making them as efficient as possible, while respecting operational constraints of power equipment and operational security. This report covers the participation of the LLGoMAX team—a collaboration of the Lawrence Livermore National Laboratory (LLNL) and ECCO International, Inc.—in Challenge 3 of the competition.
This presentation focuses on the challenge of integrating AI-driven data centers with the power grid at scale. It examines the AI data center capacity challenge and the role of new Medium Voltage Direct Current (MVDC) and other grid-enhancing technologies in enabling efficient and reliable power delivery. The session will highlight the National Laboratory of the Rockies' ARIES capabilities and planning tools, along with collaborative examples involving Verrus, Compass, and Schneider through the Agora test bed for grid-friendly data center evaluations, and ON. Energy for UPS evaluation. It will showcase the NLR Stable Grid Platform for studying oscillations caused by large-scale data centers, along with planning tools to assess grid security and reliability. Additionally, the presentation covers reconductoring strategies to increase grid capacity and explores innovative data center architectures, including the Advanced DC Architectures with Power-electronic Transformers (ADAPT) platform, which enables testing of complete DC architectures for data centers.
This report examines the application of artificial intelligence (AI) technologies for insider threat mitigation (ITM) programs in nuclear security facilities. Insider threat detection presents unique challenges due to the subtle and adaptive nature of these threats, the complex signatures involved, and the scarcity of available data for analysis. Traditional human-centered approaches, while essential, face limitations in processing large amounts of data continuously and detecting subtle patterns across multiple systems. AI technologies can potentially address these limitations by providing 24/7 monitoring capabilities, identifying complex patterns that might escape human observation, and offering consistent application of security criteria. However, the deployment of AI in nuclear security contexts introduces significant new risks, including workflow disruption, expanded attack surfaces, potential for misuse, and ethical concerns regarding privacy, fairness, transparency, safety, and security. The high-consequence nature of nuclear security decisions demands careful consideration of these risks and systematic approaches to their mitigation.
We demonstrate a wireless security application to protect the weakest link in phone-to-phone communication, using a terahertz metasurface. To our knowledge, this is the first example of an eavesdropping countermeasure in which the attacker is actively misled.
This report presents the design of defensive cybersecurity architectures (DCSAs) for High Temperature, Gas-Cooled Reactors (HTGRs). A DCSA is a cybersecurity design feature that places systems into security zones in a graded approach according to the importance of the functions performed by the systems. DCSA design efforts for advanced reactors may commence as early as the system-level design phase. This design approach is consistent with the draft regulatory guide for advanced reactor cybersecurity programs (DG-5075) and enables advanced reactor designers to consider the effects of security-by-design (SeBD) features on their DCSAs. Integration of DCSA design and other cybersecurity activities with the traditional design process as part of a SeBD framework may enable advanced reactor designers to improve the security posture of their plants while reducing implementation and operating costs. This report provides a DCSA template for an exemplar HTGR and describes a DCSA design process using event tree analysis so that the template may be optimized for a given HTGR design.
The Broadband Automation for Distributed Grid Efficiency and Resilience (BADGER) project aligns with national strategic priorities for integrating emerging wireless technologies and advancing AI-driven security. As critical infrastructure modernizes toward increasingly software-defined and interconnected systems, the ability to leverage 5G/NextG networks and AI-enabled control becomes essential. This report outlines work at the National Laboratory of the Rockies (NLR) to develop a NextG-native security architecture powered by AI-RAN concepts and evaluate workflows that enable efficient and reliable architectures. Together, these efforts position the laboratory to accelerate innovation while directly supporting national security and resilience objectives.
Reliable, secure access to energy is a major focus for national security efforts. One potential route to such energy is through fusion reactions in inertial confinement fusion (ICF) experiments. Such experiments are carried out at facilities such as the National Ignition Facility (NIF) in Livermore, California, where high powered lasers are used to compress a DT fuel-containing target to the necessary high temperature, high pressure conditions. These experiments are limited in number, which creates a heavy dependence on high fidelity predictive physics simulations and analysis performed “pre shot,” or before the experiment occurs. Many of these simulations in higher dimensions (2D and 3D) are computationally expensive, so finding optimal simulation-based designs presents its own challenges. In this work, we present our multi-fidelity Bayesian optimization with Gaussian processes (GPs) for ICF double shell targets, where a 1D surrogate model is used to help find a 2D surrogate model, enabling us to find optimal targets in the higher fidelity (2D), while saving computational cost.
The rapid expansion of distributed and edge computing platforms—spanning autonomous vehicles, IoT sensors, and healthcare monitors—has heightened concerns about data privacy. Differential Privacy (DP) offers a rigorous mathematical framework to protect sensitive information while retaining analytical utility. This tutorial introduces the foundations of DP for both numerical and categorical datasets and extends the discussion to correlation-aware techniques tailored for structured and high-dimensional data. Hands-on demonstrations will begin with the PETINA (Privacy prEservaTIoN Algorithms) package for numerical data and continue with MIC-DP (Maximum Information Correlated Differential Privacy) for tabular data. Designed for researchers and practitioners in secure systems, embedded architectures, and AI accelerators, the tutorial emphasizes practical and scalable methods for integrating DP into real-world system designs.