Search NASA⌕ Search

SEARCH · Search NASA

Results for “Policy gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Improved assessment of mangrove forests in Sundarbans East Wildlife Sanctuary using WorldView 2 and TanDEM-X high resolution imagery

Recent developments of remote sensing techniques which can capture both the structure and function of the ecosystem provide a more representative view of the landscape. These unique Earth observations were used to help improve traditional forestry surveys by providing species-specific land cover classes for mangrove forests in the Sundarbans East Wildlife Sanctuary. By combining optical data from WorldView2 (WV2; 2 m pixel) and a canopy height model derived using radar data from TanDEM-X (TDX; 12 m pixel), we identified nine mangrove and five non-mangrove classes by following an Iterative Self-Organizing Data Analysis Algorithm. Three dominant mangrove species accounted for nearly 50% of the sanctuary. Heritieria fomes disproportionately covered the largest area at 43%, overturning previous field-based estimates of Excoecaria agallocha dominance. E. agallocha and Sonneratia apetala, covered 3% and 1.47% of the sanctuary, respectively. Four mixed species classes were also identified with clear vegetation zonation patterns that trended toward species homogeneity with increasing distance from shore. The overall land cover accuracy (WV2: 89.33%; WV2-TDX: 89.89%), the Kappa Coefficient (WV2:0.88; WV2-TDX: 0.89) and change statistics between WV2 and WV2-TDX landcover classifications indicate that the WV2 imagery can separate mangrove community types without structural data. The combination of the land cover classifications and the canopy height model indicated that H. fomes were not only the most dominant forest but also, on average, the tallest (12.3 m) among the other eight mangrove types. Our large-scale mapping with high resolution optical and radar platforms can capture subtle changes in mangrove vegetation and canopy structural gradients more accurately and be used to monitor biodiversity changes and Aichi Biodiversity Targets and Indicators, which would contribute to biodiversity policy updating.

Md Mizanur Rahman↗

Optimal startup control of a jacketed tubular reactor.

The optimal startup policy of a jacketed tubular reactor, in which a first-order, reversible, exothermic reaction takes place, is presented. A distributed maximum principle is presented for determining weak necessary conditions for optimality of a diffusional distributed parameter system. A numerical technique is developed for practical implementation of the distributed maximum principle. This involves the sequential solution of the state and adjoint equations, in conjunction with a functional gradient technique for iteratively improving the control function.

Hahn, D. R.↗

MULTI-OBJECTIVE REINFORCEMENT LEARNING FOR LOW-THRUST TRANSFER DESIGN BETWEEN LIBRATION POINT ORBITS

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to construct low-thrust transfers between periodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Christopher J. Sullivan↗

MULTI-OBJECTIVE REINFORCEMENT LEARNING FOR LOW-THRUST TRANSFER DESIGN BETWEEN LIBRATION POINT ORBITS

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective reinforcement learning algorithm used to construct low-thrust transfers between periodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification

Christopher John Sullivan↗

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L.↗

Multi-objective Reinforcement Learning for Low-thrust Transfer Design Between Libration Point Orbits

Multi-Reward Proximal Policy Optimization (MRPPO) is a multi-objective rein- forcement learning algorithm used to construct low-thrust transfers between pe- riodic orbits in multi-body systems. Previous implementations of MRPPO have relied on a predefined reference transfer to successfully train each policy. In this paper, an algorithmic modification labeled the ‘moving reference’, is introduced to autonomously construct these reference trajectories during training. With this modification, MRPPO is used to recover various low-thrust transfers between two periodic orbits in the Earth-Moon circular restricted three-body problem to solve a multi-objective optimization problem. These results are then compared with the solutions recovered via a gradient descent optimization scheme to validate the performance of MRPPO with the moving reference modification.

Anderson, Rodney L↗

Supplementary active stabilization of nonrigid gravity gradient satellites

The use of active control for stability augmentation of passive gravity gradient satellites is investigated. The reaction jet method of control is the main interest. Satellite nonrigidity is emphasized. The reduction in the Hamiltonian H is used as a control criteria. The velocities, relative to local vertical, of the jets along their force axes are shown to be of fundamental significance. A basic control scheme which satisfies the H reduction criteria is developed. Each jet is fired when its velocity becomes appropriately large. The jet is de-energized when velocity reaches zero. Firing constraints to preclude orbit alteration may be needed. Control is continued until H has been minimized. This control policy is investigated using impulse and rectangular pulse models of the jet outputs.

Keat, J. E.↗

INTEX-NA: Intercontinental Chemical Transport Experiment - North America

INTEX-NA is an integrated atmospheric chemistry field experiment to be performed over North America using the NASA DC-8 and P-3B aircraft as its primary platforms. It seeks to understand the exchange of chemicals and aerosols between continents and the global troposphere. The constituents of interest are ozone and its precursors (hydrocarbons, NOX and HOX), aerosols, and the major greenhouse gases (CO2, CH4, N2O). INTEX-NA will provide the observational database needed to quantify inflow, outflow, and transformations of chemicals over North America. INTEX-NA is to be performed in two phases. Phase A will take place during the period of May-August 2004 and Phase B during March-June 2006. Phase A is in summer when photochemistry is most intense and climatic issues involving aerosols and carbon cycle are most pressing, and Phase B is in spring when Asian transport to North America is at its peak. INTEX-NA will coordinate its activities with concurrent measurement programs including satellites (e. g. Terra, Aura, Envisat), field activities undertaken by the North American Carbon Program (NACP), and other U.S. and international partners. However, it is being designed as a 'stand alone' mission such that its successful execution is not contingent on other programs. Synthesis of the ensemble of observation from surface, airborne, and space platforms, with the help of global/regional models is an important It is anticipated that approximately 175 flight hours for each of the aircraft (DC-8 and P-3B) will be required for each Phase. Principal operational sites are tentatively selected to be Bangor, ME; Wallops Island, VA; Seattle, WA; Rhinelander, WI; Lancaster, CA; and New Orleans, LA. These coastal and continental sites can support large missions and are suitable for INTEX-NA objectives. The experiment will be supported by forecasts from meteorological and chemical models, satellite observations, surface networks, and enhanced O3,-sonde releases. In addition to characterizing Atlantic-outflow and Pacific-inflow, INTEX-NA will characterize air masses transported between the U.S., Canada, and Mexico. INTEX-NA will be the first continental scale inflow, outflow, and transformation experiment to be performed over North America. It will provide the most comprehensive observational data set to date to understand the O3/NOX/HOX/aerosol photochemical system and the carbon cycle. One of the critical needs of the carbon cycle research is to obtain large-scale vertical and horizontal concentration gradients of CO2, throughout the troposphere over continental source/sink regions. INTEX-NA is ideally suited to perform this role. Coastal and continental operational sites will allow us to develop a curtain profile of greenhouse gases (e. g. CO2,) and other key pollutants across North America. Such information is central to our quantitative understanding of chemical budgets on the continental scale. We expect to provide a number of satellite under-flights over land and water to test and validate observations from the appropriate satellite platform (e. g. Aura). We plan to develop strong collaborations with other national and international observational programs. Results from INTEX-NA should directly benefit the development of environmental policy for air quality and climate change.

Singh, Hanwant B.↗

Reevaluating Suitability Estimates Based on Dynamics of Cropland Expansion in the Brazilian Amazon

Agricultural suitability maps are a key input for land use zoning and projections of cropland expansion. Suitability assessments typically consider edaphic conditions, climate, crop characteristics, and sometimes incorporate accessibility to transportation and market infrastructure. However, correct weighting among these disparate factors is challenging, given rapid development of new crop varieties, irrigation, and road networks, as well as changing global demand for agricultural commodities. Here, we compared three independent assessments of cropland suitability to spatial and temporal dynamics of agricultural expansion in the Brazilian state of Mato Grosso during 2001 2012. We found that areas of recent cropland expansion identified using satellite data were generally designated as low to moderate suitability for rainfed crop production. Our analysis highlighted the abrupt nature of suitability boundaries, rather than smooth gradients of agricultural potential, with little additional cropland expansion beyond the extent of the flattest areas (0-2% slope). Satellite-based estimates of the interannual variability in the use of existing crop areas also provided an alternate means to assess suitability. On average, cropland areas in the Cerrado biome had higher utilization (84%) than croplands in the Amazon region of northern Mato Grosso (74%). Areas of more recent expansion had lower utilization than croplands established before 2002, providing empirical evidence for lower suitability or alternative management strategies (e.g., pasture soya rotations) for lands undergoing more recent land use transitions. This unplanted reserve constitutes a large area of potentially available cropland (PAC)without further expansion, within the management limits imposed for pest management and fallow cycles. Using two key constraints on future cropland expansion, slope and restrictions on further deforestation of Amazon or Cerrado vegetation, we found little available flat land for further legal expansion of crop production in Mato Grosso. Dynamics of cropland expansion from more than a decade of satellite observations indicated narrow ranges of suitability criteria, restricting PAC under current policy conditions, and emphasizing the advantages of field-scale information to assess suitability and utilization.

Morton, Douglas C.↗