Search NASA⌕ Search

SEARCH · Search NASA

Results for “adversarial policies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Automated Adversary-in-the-Loop Cyber-Physical Defense Planning

Security of cyber-physical systems (CPS) continues to pose new challenges due to the tight integration and operational complexity of the cyber and physical components. To address these challenges, this article presents a domain-aware, optimization-based approach to determine an effective defense strategy for CPS in an automated fashion—by emulating a strategic adversary in the loop that exploits system vulnerabilities, interconnection of the CPS, and the dynamics of the physical components. Our approach builds on an adversarial decision-making model based on a Markov Decision Process (MDP) that determines the optimal cyber (discrete) and physical (continuous) attack actions over a CPS attack graph. The defense planning problem is modeled as a non-zero-sum game between the adversary and defender. We use a model-free reinforcement learning method to solve the adversary’s problem as a function of the defense strategy. We then employ Bayesian optimization (BO) to find an approximate best-response for the defender to harden the network against the resulting adversary policy. This process is iterated multiple times to improve the strategy for both players. We demonstrate the effectiveness of our approach on a ransomware-inspired graph with a smart building system as the physical process. Numerical studies show that our method converges to a Nash equilibrium for various defender-specific costs of network hardening.

97 MATHEMATICS AND COMPUTING↗

Reinforcement Learning Approach to Cybersecurity in Space (RELACSS)

Securing satellite groundstations against cyber-attacks is vital to national security missions. However, these cyber threats are constantly evolving. As vulnerabilities are discovered and patched, new vulnerabilities are discovered and exploited. In order to automate the process of discovering existing vulnerabilities and the means to exploit them, a reinforcement learning framework is presented in this report. We demonstrate that this framework can learn to successfully navigate an unknown network and detect nodes of interest despite the presence of a moving target defense. The agent then exfiltrates a file of interest from the node as quickly as possible. This framework also incorporates a defensive software agent that learns to impede the attacking agents progress. This setup allows for the agents to work against each other and improve their abilities. We anticipate that this capability will help uncover unforeseen vulnerabilities and the means to mitigate them. The modular nature of the framework enables users to swap out learning algorithms and modify the reward functions in order to adapt the learning tasks to various use cases and environments. Several algorithms, viz., tabular Q learning, deep Q networks, proximal policy optimization, advantage actor-critic, generative adversarial imitation learning, are explored for the agents and the results highlighted. The agent learns to solve the tasks in a light-weight abstract environment. Once the agent learns to perform sufficiently well, it can be deployed in a minimega virtual machine environment (or a real network) with wrappers that map abstract actions to software commands. The agent also uses a local representation of the actions called a ‘slot-mechanism’. This allows the agent to learn in a certain network and generalize it to different networks. The defensive agent learns to predict the actions taken by an offensive agent and uses that information to anticipate the threat. This information can then either be used to raise an alarm or to take actions to thwart the attack. We believe that with the appropriate reward design, a representative environment, and action set, this framework can be generalized to tackle other cybersecurity tasks. By sufficiently training these agents, we can anticipate vulnerabilities leading to robust future designs. We can also deploy automated defensive agents that can help secure satellite groundstation and their vital national security missions.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

U.S. Nuclear Declaratory Policy 2021: the Renewed Debate about Sole Purpose and No-First-Use. Annotated Bibliography

Declaratory policy and public statements about the potential use of nuclear weapons serve many important roles. They provide an assessment of the security environment, and inform the public debate. These statements also enhance deterrence messages and signals towards adversaries, and reassure allies and partners. On the global level, U.S. declaratory policy has the potential to shape international trends and norms, influence nuclear proliferation, and it may also affect the policy decisions of other nuclear possessors. As the Biden administration reviews the elements of U.S. nuclear declaratory policy, the issue of sole purpose and no-first-use is likely to resurface. Previous administrations have examined these policies in multiple rounds of review, and they decided that the time was not right for such declarations. This literature review was prepared to inform the debate by collecting some of the most prominent articles on the topic that highlight the potential risks and benefits of these policies.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Technology assessment or technology harassment: The attacks on science and technology

The manner in which technology is being assessed by various groups and individuals is discussed. Attacks on science and technology (specifically military uses and funding), and the disillusionment of the public with the lack of relevance of science to the public interest and with the infallible wisdom of scientists are described. The effects of the fear of environmental pollution are emphasized. It is concluded that the important issues of science, technology, and public policy will require a pluralistic definition of the public interest by open adversary proceedings. It is also felt that the danger is not that new technology will receive inadequate assessment for possible deleterious secondary effects, but that harassment by an overemotional political process may prevent its coming to fruition.

Green, L., Jr.↗

American Made Infrastructure: Evolution of Federal Incentives and Requirements

Foreign Entity of Concern (FEOC) restrictions in the One Big Beautiful Bill Act (OBBB) represent the latest evolution of a multi-year legislative trajectory responding to national security concerns about foreign control – and particularly FEOC control – of energy infrastructure. Beginning with Executive Order 14017 (February 2021), which initiated comprehensive federal review of critical supply chain vulnerabilities in semiconductors, battery energy storage systems, and critical minerals, policymakers have progressively expanded restrictions on foreign participation. The National Defense Authorization Act (NDAA) 2019 established precedent for component-level prohibitions on foreign information and communications technology procurement, while NDAA 2024 extended these restrictions to six major People’s Republic of China (PRC) battery manufacturers. Complementary measures such as the Build America, Buy America (BABA) Act and the Infrastructure Investment and Jobs Act (IIJA) introduced domestic content thresholds and FEOC eligibility criteria for federal funding programs. The Inflation Reduction Act (IRA) 2022 further operationalized FEOC restrictions through electric vehicle tax credit requirements, creating a scalable framework for excluding foreign-controlled components. Recent executive actions and state-level policies have reinforced this trajectory, reflecting sustained alignment across federal and state governments. Collectively, these developments demonstrate a bipartisan policy approach that pairs incentives for advanced energy deployment with safeguards designed to prevent subsidizing adversaries or entities that present foreign-sourcing risk.

99 - GENERAL AND MISCELLANEOUS↗

Spatiotemporal Super-Resolution with Generative Machine Learning for Creating Renewable Energy Resource Data Under Climate Change Scenarios

As we plan for a future with higher penetrations of renewables and increasing electrification, it becomes more important to understand how the electricity grid will operate under a variety of weather events. We must also consider that the weather our future grid will experience will be different and possibly more extreme than the historical weather that we have extensive data for. We can use data from global climate models (GCMs) to help understand how our climate may change over the next several decades, but there is often a significant gap between the low-resolution GCM data and the high-resolution weather data required to study power systems under specific weather events. Therefore, our objective in this work is to develop tools that can bridge this gap by using low-resolution GCM data to create realistic high-resolution weather datasets that can be used to study renewable energy generation and electricity demand. To accomplish this objective, we have developed a set of generative machine learning models that can rapidly downscale GCM daily average output data at an approximate grid resolution of 100km to hourly data at an approximate 4 km grid resolution. The models can be used to create high resolution data from nearly any GCM included in the Coupled Model Intercomparison Project (CMIP) Phase 5 or 6. Our methods include all datasets regularly used to study the integration of wind and solar power plants as well as changes in electricity demand due to heating and cooling loads. These models and datasets enable power systems modelers to study climate change-influenced weather events and their impact on the grid. We have downscaled and validated wind, solar, temperature, and humidity data with very promising results. The generative machine learning methods are computationally efficient and produce data that has similar statistical characteristics to current state-of-the-art historical datasets. We have trained initial generative models and produced an initial dataset collectively referred to as Sup3rCC: Super-Resolved Renewable Energy Resource Data with Climate Change Impacts. The data covers a (mostly) historical period from 2015-2025 and a future period from 2050-2059. We have also taken hypothetical high-electrification load data and scaled the heating and cooling loads with respect to the 2050-2059 high-resolution Sup3rCC meteorology. The results show how future levels of renewable energy generation and electrified load may be impacted by climate change, setting the stage for capacity expansion models to consider a dynamic climate through model years.

climate change↗

Model-based Hierarchical Reinforcement Learning for Improved Physical Security Design: A Prototype

Prior work in FY24 developed an adversarial AI agent aid in path analysis of physical protection systems. This agent, trained using a model-based reinforcement learning algorithm, was able to successfully learn the most vulnerable path in facilities. It was able to extend the current state of practice for physical protection design by exhibiting dynamic behavior based on current environmental conditions. Whereas PathTrace largely performs a static, graph-based analysis, the AI agent was able to make decisions based on relative position in the facility, current conditions (was the adversarial agnet discovered?), and proximity to secondary targets. The agent demonstrated some novel capabilities, but had limitations that need to be resolved before it can be used for production purposes. For example, the adversarial agent generalizes poorly and takes a relatively long time to train. Nonetheless, there is still considerable promise for developing the adversarial agent further in order to explore even richer, more dynamic behaviors (e.g., adversary motivations, environmental debris, and more). This work considers a complementary idea; development of a planning agent. The planning agent is envisioned as an auto-complete-like tool that can help accelerate security system design by human experts. The agent would respect existing barriers and sensors placed by a human expert while offering cost-effective suggestions (i.e., implicitly balancing effectiveness with cost) to improve the design. The goal is for this agent to be part of an expert’s toolbox, not to totally upend the current state-of-practice, or to displace human experts. The ultimate goal would be concurrent training of both the adversarial and planning agent together, to learn entirely through self-play. This would represent an entirely new way of performing system deign. We selected a hierarchical, model-based reinforcement learning algorithm to serve as the planning agent. This is an extension of concepts used in the prior FY24 adversarial agent work. There, we had a single agent acting an environment. Here, we have two different sub-agents (policies), working together, to form a complete agent. There is a manager policy, which can select abstract goals on slower time scales, and a worker, which performs primitive actions to reach goals selected by the manager. It is worth noting that this class of algorithm is challenging to work with. From our understanding, our work is one of the first successful uses of model-based reinforcement learning (MBRL) in nuclear energy1 , and likely the first hierarchical model-based reinforcement learning application in nuclear energy. Further, this work is one of the first known attempts to apply AI to perform a design tasks in nuclear energy. Consequently, there were significant implementation challenges and the bulk of the work was focused on successful implementation and algorithm design. The results presented here are very low technology readiness level as a consequence of the lack of related literature, but still represent a significant step forward in the pursuit of applied AI for design.

42 ENGINEERING↗

Essence2.0 Development and Deployment (Final Report)

The objective of the project was to take two core technologies that have been developed under the Recipient’s solid laboratory products and integrate them into a single CyberPhysical awareness platform and complete development on current field-tested prototypes that will extend the integrated capability of the platform. During the final development phase, the Recipient development team and its selected industry partners tested and hardened the platform to ensure resilient and secure operation of the integrated platform. The team also executed substantial field testing and established the framework for defining the organization and/or commercial infrastructure needed to sustain operations and provide readiness for a national scale deployment. The focus of the project was (1) the improvement, refinement, and deployment of technology for the detection of cyber-attacks on utility operational technology (OT) and information technology (IT) networks and assets, including Supervisory Control and Data Acquisition Systems (SCADA) systems; and (2) support for containment and remediation of adversarial threats and actions against those systems and environments.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

Contentious narratives and disinformation about nuclear weapons in strategic deterrence and competition: A SOCOM perspective. Part of Section: United States Special Operations Command (USSOCOM) in Emerging Strategic & Geopolitical Challenges: Operational Implications for US Combatant Commands

Russia’s “special military operation” in Ukraine demonstrates the challenge for strategic deterrence and competition of countering contentious narratives and disinformation about weapons of mass destruction (WMD) during conventional regional wars against a nucleararmed adversary. Moscow uses both tailored, contentious narratives and targeted disinformation about WMD in Ukraine to influence and disrupt local and global perceptions in support of its deterrence and competition objectives vis-à-vis the United States and NATO. Since December 2021, Moscow has made a focal point of chemical, biological, radiological, and nuclear weapons in its efforts to establish a permissive environment for its military buildup on the border with Ukraine and then its military intervention. These information tactics also demonstrate an opportunity for US Special Operations Forces (SOF). They are a case study for considering how SOF can contribute to strategic deterrence and competition objectives, specifically countering adversary gray-zone information efforts to alter regional security orders.1 Such a role is in line with the 2022 Special Operations Forces Vision and Strategy, which provides a framework for the evolution of SOF into “a force capable of creating strategic, asymmetric advantages for the nation as a key contributor of integrated deterrence” (United States Special Operations Command, 2022). This paper briefly examines this strategic challenge and SOF opportunity, focusing narrowly on the distinction between contentious narratives and disinformation about nuclear weapons and the role of SOF in countering these gray-zone information tactics. The nuclear dimension of Moscow’s contentious narratives and disinformation in the “special military operation” is of particular interest because it demonstrates the distinction between strategic efforts to influence and disrupt local and global perceptions in Moscow’s favor. This distinction between influence and disruption is less clear with Russia’s contentious narratives and disinformation about chemical and biological weapons in Ukraine, as disinformation about these two types of WMD appears to overwhelm contentious narratives. We believe this distinction is useful for policymakers and warfighters responsible for countering gray-zone information tactics because it provides a framework for crafting tailored responses to contentious narratives and disinformation about nuclear weapons and other WMD. The chapter concludes with a discussion of efforts that could be undertaken by SOF in cooperation with other relevant stakeholders to address this aspect of adversary gray-zone information tactics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

DTS: Building custom, intelligent schedulers

DTS is a decision-theoretic scheduler, built on top of a flexible toolkit -- this paper focuses on how the toolkit might be reused in future NASA mission schedulers. The toolkit includes a user-customizable scheduling interface, and a 'Just-For-You' optimization engine. The customizable interface is built on two metaphors: objects and dynamic graphs. Objects help to structure problem specifications and related data, while dynamic graphs simplify the specification of graphical schedule editors (such as Gantt charts). The interface can be used with any 'back-end' scheduler, through dynamically-loaded code, interprocess communication, or a shared database. The 'Just-For-You' optimization engine includes user-specific utility functions, automatically compiled heuristic evaluations, and a postprocessing facility for enforcing scheduling policies. The optimization engine is based on BPS, the Bayesian Problem-Solver (1,2), which introduced a similar approach to solving single-agent and adversarial graph search problems.

Hansson, Othar↗

Reconceptualizing Off-Ramps as a Tool of U.S.-China Escalation Management

Great power rivalry, at any point on the competition spectrum, involves assessments of risks and decisions on where to take them. Assessing risk versus reward is also integral to the decision to cooperate in the form of arrangements like arms control. As Donald Brennan wrote more than sixty years ago, the value in arms control lies in whether it “reduce[s] the hazards of the present armament policies by a factor greater than the amount of risk introduced by the control measures themselves.” 2 States can assume in peacetime the risks inherent in cooperating with an adversary, to attempt to mitigate risks likely to arise in crisis or conflict. Or they can choose to forgo cooperation and assume risks later, which will complicate their ability to manage crises, war, and war termination.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Contextual Active Online Model Selection with Expert Advice

How can we collect the most useful labels to learn a model selection policy, when presented with arbitrary heterogeneous data streams? In this paper, we formulate this task as a contextual active model selection problem, where at each round the learner receives an unlabeled data point along with a context. The goal is to output the best model for any given context without obtaining an excessive amount of labels. In particular, we focus on the task of selecting pre-trained classifiers, and propose a contextual active model selection algorithm (CAMS), which relies on a novel uncertainty sampling query criterion defined on a given policy class for adaptive model selection. In comparison to prior art, our algorithm does not assume a globally optimal model. We provide rigorous theoretical analysis for the regret and query complexity under both adversarial and stochastic settings. Our experiments on several benchmark classification datasets demonstrate the algorithm’s effectiveness in terms of both regret and query complexity. Notably, to achieve the same accuracy, CAMS incurs less than 10% of the label cost when compared to the best online model selection baselines on CIFAR10.

Liu, Xuefeng↗

Defending the Homeland: Growing Foreign Challenges to the U.S. Missile Defense Posture

The Office of the Secretary of Defense (OSD) for Policy requested that Lawrence Livermore National Laboratory conduct a Congressionally directed study on homeland missile defense, pursuant to Section 1692 of the fiscal year 2020 National Defense Authorization Act. In accordance with the statutory language, this study considers whether the security benefits obtained by deployment of homeland missile defenses of the United States are undermined or counterbalanced by adverse reactions of potential adversaries, and considers the effectiveness of homeland missile defense efforts of the United States to deter the development of ballistic missiles. Almost a half-century has elapsed since the United States and the Soviet Union signed the Anti-Ballistic Missile (ABM) Treaty, and almost two decades since the former withdrew. Since withdrawing, the United States has sought to develop and deploy a layered missile defense system to defend against regional threats to U.S. and allied interests abroad and to counter limited threats to the U.S. homeland. As such, missile defense supports key national defense policy objectives. Protecting the U.S. homeland, forces abroad, allies, and partners. Deterring attacks against the United States, its allies, and partners. Assuring allies an strengthening U.S. diplomatic activities in peacetime and crisis. We offer five key findings: 1. The primary benefit of the Ground-based Midcourse Defense (GMD) system is the protection it provides the U.S. homeland against a limited but evolving rogue-state missile threat. While this system has never been tested in combat, it appears thus far to have effectively paced North Korea’s development and deployment of intercontinental ballistic missiles (ICBMs). In the absence of such a defensive capability, the United States would likely have been more heavily exposed to North Korean actions, would have operated at a much higher-risk posture, and would have had less negotiating room with which to navigate coercive tactics and crises. Its allies would have been more concerned about U.S. willingness in time of crisis and war to run the risks of protecting them. The GMD system also serves as a hedge against Iranian breakout. 2. Potential adversaries continue to develop long-range ballistic missiles despite deployment of a U.S. homeland missile defense system. This includes both rogue states and major power rivals whose long-range missile programs predated the deployment of U.S. missile defenses. While North Korea has continued its long-range missile developments and achieved an intercontinental capability, Iran has not yet reached this threshold. The broader proliferation of long-range missiles anticipated in the late 1990s has not materialized. 3. While Russia has used the existence of a U.S. homeland missile defense system as a justification for its substantial and continuing weapon modernization program, neither Russian force modernization nor the limited U.S. homeland missile defense system has altered the strategic balance. Russia has long considered U.S. missile defenses—both theater and homeland—as directed against its strategic forces and as a capability that could rapidly advance, thereby eroding Russian confidence in its nuclear deterrent. Russia’s political and military actions appear excessive and negatively impact areas such as arms control, but they do not undermine the primary benefit of the U.S. system cited in Key Finding 1. 4. China’s expansion and diversification of nuclear and missile forces has been influenced by its concerns about U.S. homeland missile defense, but those concerns are only one of many factors in China’s force planning. Although China’s actions to preserve and expand its assured retaliation prospects have not fundamentally altered the strategic balance, its buildup and lack of transparency about its force modernization goals are troubling. Taken together, the aggregate pattern of China’s modernization activities over the past two decades strongly suggests a concerted effort to develop a modern military force commensurate with its intended geopolitical status. Whatever China’s legacy concerns over the survivability of its strategic nuclear forces, its ability to overwhelm the GMD system appears intact today and will be further strengthened in the years ahead as it continues its long-term modernization program. China’s political and military actions negatively affect U.S. security interests but do not undermine the primary benefit of the U.S. system cited in Key Finding 1. 5. The basic finding that benefits have not so far been undermined by adverse reactions is a function of circumstances that are increasingly in flux. The existing GMD-centered system is under increasing pressure from the pace and scope of North Korean missile deployments. The proposed “layered” system seeks to mitigate this impending capability gap in the near term and to complement future capabilities, such as the Next Generation Interceptor, when they come online. Given the pace of U.S. missile defense developments since 2000, it is possible that adversary advances in capability will outpace the U.S. system’s ability to adapt. Additionally, because this layered system is in principle more readily scaled, it will almost certainly be viewed skeptically by China and Russia. In a context of growing great-power security competition, the Department of Defense (DoD) should consider undertaking a broader net assessment of the U.S., Russian, and Chinese force balance.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

A Technical Evaluation of Self-Protection Dose Rates in U.S. Domestic Nuclear Security Policy

U.S. nuclear security policy uses radiation self-protection as a basis for reducing material attractiveness, but existing dose-rate criteria may not reflect the time scales of theft or sabotage. This report evaluates whether the commonly cited 1 Gy/h at 1 m criterion can plausibly cause adversary task failure during short-duration malicious acts.

Fritchie, Jacob Wesley [Sandia National Laboratori↗

Hawaii Threat Brief

The Digital Energy Transformation is redefining how power systems operate, integrating physical infrastructure with digital technologies to enhance efficiency, visibility, and resilience. For islanded and digitally modernizing systems like Hawai‘i’s, this shift presents both new opportunities and evolving challenges in cybersecurity, supply chain assurance, and environmental resilience. This brief provides an overview of critical infrastructure dependencies and systemic risks associated with increasing digital integration. It summarizes recent energy-sector threat activity and case studies that illustrate adversary tactics and supply chain vulnerabilities, and identifies pathways to strengthen resilience.

25 - ENERGY STORAGE↗