Search NASA⌕ Search

SEARCH · Search NASA

Results for “vision transformers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Effects of input gradient regularization on neural networks time-series forecasting of thermal power systems

This study proposes using neural networks, specifically gated recurrent unit (GRU), long-short-term memory (LSTM), and transformer networks, to improve control strategies in a 450 MW coal-fired power plant. However, neural networks face issues of becoming overly dependent on just a few variables to make predictions, which negatively impacts control decisions that rely on the model to determine the value of all manipulated variables. The paper introduces regularization techniques, including noise injection and input gradient regularization, during the training phase. Here, the work presents novel contributions in adapting neural networks to control industrial systems and applying regularization techniques from computer vision to industrial process control. Results demonstrate the effectiveness of input gradient regularization in reducing model dependence on subsets of variables, emphasizing the balance between fidelity and controllability. Further exploration is recommended, including the development of recurrent transformers, closed-loop control testing, and a sensitivity analysis on computer models to provide further insight.

20 FOSSIL-FUELED POWER PLANTS↗

3D reconstruction and neural rendering for adversarial machine learning

While evasion attacks on computer vision systems have been widely studied, creating attacks that remain effective under significant changes in viewpoint continues to be challenging. Traditional approaches often rely on affine transformations of images, but these approaches degrade at larger perspective shifts and often produce unrealistic or ineffective perturbations. Recent methods use differentiable renderers to improve viewpoint robustness, but they typically depend on manually constructed 3D models. We introduce a semi-automated pipeline that generates physically printable and perspective-invariant adversarial patches using only a small set of 2D images. Our method integrates 3D reconstruction, neural rendering, adversarial patch optimization, and an object detection victim model into a unified workflow. We use 2D Gaussian Splatting for high fidelity mesh reconstruction and FlexPara for surface parameterization that produces texture maps suitable for patch editing. Together, these components form a fully differentiable pipeline in PyTorch3D that links texture modification to model outputs, enabling efficient optimization of patches that remain effective across many viewpoints. The complete process, from image capture to patch printing and physical evaluation, can be completed within a few hours. We demonstrate the effectiveness of the resulting patches through attacks on the YOLOv8 object detection model and discuss remaining challenges and opportunities for improving robustness and scalability.

Singhvi, Vivaan [ORNL] (ORCID:0009000586288221)↗

Plant Design for a Developing Bioeconomy Workshop Report: Frontier Science for the Bioeconomy Workshop Series

Recent advances in fundamental plant biology research, synthetic biology, and artificial intelligence (AI) are unlocking powerful new capabilities in plant biodesign, offering unprecedented potential to reimagine plants as programmable platforms for resource-efficient production of bioenergy, biomaterials, chemicals, and more. The U.S. Department of Energy (DOE) convened the Plant Design for a Developing Bioeconomy virtual workshop on March 12 through 14, 2025, to bring together leaders across plant science, engineering, and computation to assess the current landscape and define a bold vision for future research. Discussions during the workshop built upon findings included in DOE’s Biological and Environmental Research (BER) workshop report Overcoming Barriers in Plant Transformation: A Focus on Bioenergy Crops (U.S. DOE 2024; genomicscience. energy.gov/plant-transformation). Participants identified critical knowledge gaps, technical barriers, and emerging opportunities in the design and engineering of plant systems to support a robust, resilient domestic bioeconomy aligned with DOE’s mission.

09 BIOMASS FUELS↗

Future foundries: A convergent manufacturing platform

This article introduces the Future Foundries platform developed at Oak Ridge National Laboratory, a first-generation research system designed to demonstrate convergent manufacturing. Convergent manufacturing brings together additive, subtractive, and transformative processes in a digitally interconnected environment to enable end-to-end production workflows. By linking traditionally discrete steps, convergent platforms accelerate production, improve repeatability, and support high-mix, low-volume manufacturing. The Future Foundries platform exemplifies this vision in practice by combining four modular, vendor-agnostic process cells that include robotic WAAM, induction heating, optical metrology, and machining, coordinated through an automated pallet handler and a ROS 2-based digital thread. This architecture provides the flexibility and scalability needed for agile production in small and medium-sized manufacturing enterprises and for field deployable manufacturing. Two case studies illustrate the platform’s capabilities. The first presents an integrated workflow for fabricating, transforming, and repairing critical replacement components, showing how consolidated thermal, additive, inspection, and machining operations reduce manual part handling and streamline process flow. The second case study highlights coordinated multi-part production enabled by automated pallet logistics and multi-cell scheduling. Together, these examples showcase convergent manufacturing as a practical and scalable strategy for strengthening domestic casting and forging capacity, improving supply-chain resilience, and enabling rapid, adaptable production of mission-critical components.

Convergent manufacturing↗

Technical guidance for the development of a solid state image sensor for human low vision image warping

This report surveys different technologies and approaches to realize sensors for image warping. The goal is to study the feasibility, technical aspects, and limitations of making an electronic camera with special geometries which implements certain transformations for image warping. This work was inspired by the research done by Dr. Juday at NASA Johnson Space Center on image warping. The study has looked into different solid-state technologies to fabricate image sensors. It is found that among the available technologies, CMOS is preferred over CCD technology. CMOS provides more flexibility to design different functions into the sensor, is more widely available, and is a lower cost solution. By using an architecture with row and column decoders one has the added flexibility of addressing the pixels at random, or read out only part of the image.

Vanderspiegel, Jan↗

Discrete analysis of spatial-sensitivity models

Procedures for reducing the computational burden of current models of spatial vision are described, the simplifications being consistent with the prediction of the complete model. A method for using pattern-sensitivity measurements to estimate the initial linear transformation is also proposed which is based on the assumption that detection performance is monotonic with the vector length of the sensor responses. It is shown how contrast-threshold data can be used to estimate the linear transformation needed to characterize threshold performance.

Nielsen, Kenneth R. K.↗

Sensing Super-Position: Human Sensing Beyond the Visual Spectrum

The coming decade of fast, cheap and miniaturized electronics and sensory devices opens new pathways for the development of sophisticated equipment to overcome limitations of the human senses. This paper addresses the technical feasibility of augmenting human vision through Sensing Super-position by mixing natural Human sensing. The current implementation of the device translates visual and other passive or active sensory instruments into sounds, which become relevant when the visual resolution is insufficient for very difficult and particular sensing tasks. A successful Sensing Super-position meets many human and pilot vehicle system requirements. The system can be further developed into cheap, portable, and low power taking into account the limited capabilities of the human user as well as the typical characteristics of his dynamic environment. The system operates in real time, giving the desired information for the particular augmented sensing tasks. The Sensing Super-position device increases the image resolution perception and is obtained via an auditory representation as well as the visual representation. Auditory mapping is performed to distribute an image in time. The three-dimensional spatial brightness and multi-spectral maps of a sensed image are processed using real-time image processing techniques (e.g. histogram normalization) and transformed into a two-dimensional map of an audio signal as a function of frequency and time. This paper details the approach of developing Sensing Super-position systems as a way to augment the human vision system by exploiting the capabilities of Lie human hearing system as an additional neural input. The human hearing system is capable of learning to process and interpret extremely complicated and rapidly changing auditory patterns. The known capabilities of the human hearing system to learn and understand complicated auditory patterns provided the basic motivation for developing an image-to-sound mapping system. The human brain is superior to most existing computer systems in rapidly extracting relevant information from blurred, noisy, and redundant images. From a theoretical viewpoint, this means that the available bandwidth is not exploited in an optimal way. While image-processing techniques can manipulate, condense and focus the information (e.g., Fourier Transforms), keeping the mapping as direct and simple as possible might also reduce the risk of accidentally filtering out important clues. After all, especially a perfect non-redundant sound representation is prone to loss of relevant information in the non-perfect human hearing system. Also, a complicated non-redundant image-to-sound mapping may well be far more difficult to learn and comprehend than a straightforward mapping, while the mapping system would increase in complexity and cost. This work will demonstrate some basic information processing for optimal information capture for headmounted systems.

Maluf, David A.↗

Emerging computer technologies and the news media of the future

The media environment of the future may be dramatically different from what exists today. As new computing and communications technologies evolve and synthesize to form a global, integrated communications system of networks, public domain hardware and software, and consumer products, it will be possible for citizens to fulfill most information needs at any time and from any place, to obtain desired information easily and quickly, to obtain information in a variety of forms, and to experience and interact with information in a variety of ways. This system will transform almost every institution, every profession, and every aspect of human life--including the creation, packaging, and distribution of news and information by media organizations. This paper presents one vision of a 21st century global information system and how it might be used by citizens. It surveys some of the technologies now on the market that are paving the way for new media environment.

Vrabel, Debra A.↗

NASA's Nuclear Frontier: The Plum Brook Reactor Facility, 1941-2002

In 1953, President Eisenhower delivered a speech called "Atoms for Peace" to the United Nations General Assembly. He described the emergence of the atomic age and the weapons of mass destruction that were piling up in the storehouses of the American and Soviet nations. Although neither side was aiming for global destruction, Eisenhower wanted to "move out of the dark chambers of horrors into the light, to find a way by which the minds of men, the hopes of men, the souls of men everywhere, can move towards peace and happiness and well-being." One way Eisenhower hoped this could happen was by transforming the atom from a weapon of war into a useful tool for civilization. Many people believed that there were unprecedented opportunities for peaceful nuclear applications. These included hopeful visions of atomic-powered cities, cars, airplanes, and rockets. Nuclear power might also serve as an efficient way to generate electricity in space to support life and machines. Eisenhower wanted to provide scientists and engineers with "adequate amounts of fission- able material with which to test and develop their ideas." But, in attempting to devise ways to use atomic power for peaceful purposes, scientists realized how little they knew about the nature and effects of radiation. As a result, the United States began constructing nuclear test reactors to enable scientists to conduct research by producing neutrons.

Bowles, Mark D.↗

Multi-Vehicle (m:N) Operations in the NAS - NASA's Research Plans

The Advanced Air Mobility movement is occurring across the world with goals of enabling, affordable, efficient, accessible, and safe air transportation at a much larger scale than today’s operations, largely enabled by electrification and automation. Transformative and disruptive innovations are emerging that will support an ecosystem designed to transport goods and people to locations not traditionally served by air transportation. To realize the full vision of AAM, technology will be needed to allow a few operators to operate many vehicles (m:N). The benefits of m:N operations are described, along with the current state-of-the-art, barriers, need, and NASA’s plans to address some of the barriers, including a Multi-Vehicle (m:N) Working Group with goals of producing a community-developed operational approval roadmap for various domains.

multi-vehicle↗

Sensing Super-position: Visual Instrument Sensor Replacement

The coming decade of fast, cheap and miniaturized electronics and sensory devices opens new pathways for the development of sophisticated equipment to overcome limitations of the human senses. This project addresses the technical feasibility of augmenting human vision through Sensing Super-position using a Visual Instrument Sensory Organ Replacement (VISOR). The current implementation of the VISOR device translates visual and other passive or active sensory instruments into sounds, which become relevant when the visual resolution is insufficient for very difficult and particular sensing tasks. A successful Sensing Super-position meets many human and pilot vehicle system requirements. The system can be further developed into cheap, portable, and low power taking into account the limited capabilities of the human user as well as the typical characteristics of his dynamic environment. The system operates in real time, giving the desired information for the particular augmented sensing tasks. The Sensing Super-position device increases the image resolution perception and is obtained via an auditory representation as well as the visual representation. Auditory mapping is performed to distribute an image in time. The three-dimensional spatial brightness and multi-spectral maps of a sensed image are processed using real-time image processing techniques (e.g. histogram normalization) and transformed into a two-dimensional map of an audio signal as a function of frequency and time. This paper details the approach of developing Sensing Super-position systems as a way to augment the human vision system by exploiting the capabilities of the human hearing system as an additional neural input. The human hearing system is capable of learning to process and interpret extremely complicated and rapidly changing auditory patterns. The known capabilities of the human hearing system to learn and understand complicated auditory patterns provided the basic motivation for developing an image-to-sound mapping system.

Maluf, David A.↗

Design of optimal correlation filters for hybrid vision systems

Research is underway at the NASA Johnson Space Center on the development of vision systems that recognize objects and estimate their position by processing their images. This is a crucial task in many space applications such as autonomous landing on Mars sites, satellite inspection and repair, and docking of space shuttle and space station. Currently available algorithms and hardware are too slow to be suitable for these tasks. Electronic digital hardware exhibits superior performance in computing and control; however, they take too much time to carry out important signal processing operations such as Fourier transformation of image data and calculation of correlation between two images. Fortunately, because of the inherent parallelism, optical devices can carry out these operations very fast, although they are not quite suitable for computation and control type operations. Hence, investigations are currently being conducted on the development of hybrid vision systems that utilize both optical techniques and digital processing jointly to carry out the object recognition tasks in real time. Algorithms for the design of optimal filters for use in hybrid vision systems were developed. Specifically, an algorithm was developed for the design of real-valued frequency plane correlation filters. Furthermore, research was also conducted on designing correlation filters optimal in the sense of providing maximum signal-to-nose ratio when noise is present in the detectors in the correlation plane. Algorithms were developed for the design of different types of optimal filters: complex filters, real-value filters, phase-only filters, ternary-valued filters, coupled filters. This report presents some of these algorithms in detail along with their derivations.

Rajan, Periasamy K.↗

The spatial and logical organization of devices in an advanced industrial robot system

This paper describes the geometrical and device organization of a robot system which is based in part upon transformations of Cartesian frames and exchangeable device tree structures. It discusses coordinate frame transformations, geometrical device representation and solution degeneracy along with the data structures which support the exchangeable logical-physical device assignments. The system, which has been implemented in a minicomputer, supports vision, force, and other sensors. It allows tasks to be instantiated with logically equivalent devices and it allows tasks to be defined relative to appropriate frames. Since these frames are, in turn, defined relative other frames this organization provides a significant simplification in task specification and a high degree of system modularity.

Ruoff, C. F.↗

Robot acting on moving bodies (RAMBO): Preliminary results

A robot system called RAMBO is being developed. It is equipped with a camera, which, given a sequence of simple tasks, can perform these tasks on a moving object. RAMBO is given a complete geometric model of the object. A low level vision module extracts and groups characteristic features in images of the object. The positions of the object are determined in a sequence of images, and a motion estimate of the object is obtained. This motion estimate is used to plan trajectories of the robot tool to relative locations nearby the object sufficient for achieving the tasks. More specifically, low level vision uses parallel algorithms for image enchancement by symmetric nearest neighbor filtering, edge detection by local gradient operators, and corner extraction by sector filtering. The object pose estimation is a Hough transform method accumulating position hypotheses obtained by matching triples of image features (corners) to triples of model features. To maximize computing speed, the estimate of the position in space of a triple of features is obtained by decomposing its perspective view into a product of rotations and a scaled orthographic projection. This allows the use of 2-D lookup tables at each stage of the decomposition. The position hypotheses for each possible match of model feature triples and image feature triples are calculated in parallel. Trajectory planning combines heuristic and dynamic programming techniques. Then trajectories are created using parametric cubic splines between initial and goal trajectories. All the parallel algorithms run on a Connection Machine CM-2 with 16K processors.

Davis, Larry S.↗

Robot Acting on Moving Bodies (RAMBO): Interaction with tumbling objects

Interaction with tumbling objects will become more common as human activities in space expand. Attempting to interact with a large complex object translating and rotating in space, a human operator using only his visual and mental capacities may not be able to estimate the object motion, plan actions or control those actions. A robot system (RAMBO) equipped with a camera, which, given a sequence of simple tasks, can perform these tasks on a tumbling object, is being developed. RAMBO is given a complete geometric model of the object. A low level vision module extracts and groups characteristic features in images of the object. The positions of the object are determined in a sequence of images, and a motion estimate of the object is obtained. This motion estimate is used to plan trajectories of the robot tool to relative locations rearby the object sufficient for achieving the tasks. More specifically, low level vision uses parallel algorithms for image enhancement by symmetric nearest neighbor filtering, edge detection by local gradient operators, and corner extraction by sector filtering. The object pose estimation is a Hough transform method accumulating position hypotheses obtained by matching triples of image features (corners) to triples of model features. To maximize computing speed, the estimate of the position in space of a triple of features is obtained by decomposing its perspective view into a product of rotations and a scaled orthographic projection. This allows use of 2-D lookup tables at each stage of the decomposition. The position hypotheses for each possible match of model feature triples and image feature triples are calculated in parallel. Trajectory planning combines heuristic and dynamic programming techniques. Then trajectories are created using dynamic interpolations between initial and goal trajectories. All the parallel algorithms run on a Connection Machine CM-2 with 16K processors.

Davis, Larry S.↗

A modification of the fusion model for log polar coordinates

The fusion mechanism for application in stereo analysis of range restricted the depth of field and therefore required a shift variant mechanism in the peripheral area to find disparity. Misregistration was prevented by restricting the disparity detection range to a neighborhood spanned by the directional edge detection filters. This transformation was essentially accomplished by a nonuniform resampling of the original image in a horizontal direction. While this is easily implemented for digital processing, the approach does not (in the peripheral vision area) model the log-conformal mapping which is known to occur in the human mechanism. This paper therefore modifies the original fusion concept in the peripheral area to include the polar exponential grid-to-log conformal tesselation. Examples of the fusion process resulting in accurate disparity values are given.

Griswold, N. C.↗

Adaptive Changes In Postural Equilibrium And Motion Sickness Following Repeated Exposures To Virtual Environments

Virtual environments offer unique training opportunities, particularly for training astronauts and preadapting them to the novel sensory conditions of microgravity. Two unresolved human factors issues in virtual reality (VR) systems are: 1) potential "cybersickness", and 2) maladaptive sensorimotor performance following exposure to VR systems. Interestingly, these aftereffects are often quite similar to adaptive sensorimotor responses observed in astronauts during and/or following space flight. Changes in the environmental sensory stimulus conditions and the way we interact with the new stimuli may result in motion sickness, and perceptual, spatial orientation and sensorimotor disturbances. Initial interpretation of novel sensory information may be inappropriate and result in perceptual errors. Active exploratory behavior in a new environment, with resulting feedback and the formation of new associations between sensory inputs and response outputs, promotes appropriate perception and motor control in the new environment. Thus, people adapt to consistent, sustained alterations of sensory input such as those produced by microgravity, unilateral labyrinthectomy and experimentally produced stimulus rearrangements. Adaptation is revealed by aftereffects including perceptual disturbances and sensorimotor control disturbances. The purpose of the current study was to compare disturbances in postural control produced by dome and head-mounted virtual environment displays, and to examine the effects of exposure duration, and repeated exposures to VR systems. Forty-one subjects (21 men, 20 women) participated in the study with an age range of 21-49 years old. One training session was completed in order to achieve stable performance on the posture and VR tasks before participating in the experimental sessions. Three experimental sessions were performed each separated by one day. The subjects performed a navigation and pick and place task in either a dome or head-mounted display (HMD) VR system for either 30 or 60 min. The environment was a square room with 15 pedestals on two opposite walls. The objects appeared on one set of pedestals and the subject s objective was to move the objects to the other set of pedestals. After the subject picked up an object, a pathway appeared and they were required to follow the pathway to the other side of the room. The subject was instructed to perform the task as quickly and accurately as possible, avoiding hitting walls and other any obstacles and placing the object on the center of the pedestal. Postural equilibrium was measured (using the Equitest CDP balance system, Neurocom, International) before, immediately after, and at 1 hr, 2 hr, 4 hr and 6 hr following exposure to VR. Postural equilibrium was measured during quiet stance with eyes open, eyes closed and vision and/or ankle proprioceptive inputs selectively altered by servo-controlling the visual surround and/or support surface to the subject s center of mass sway. Posture data was normalized using a log transformation and motion sickness data were normalized using the square root. In general, we found that exposure to VR resulted in decrements in postural stability. The largest decrements were observed in the tests performed immediately following exposure to VR and showed a fairly rapid recovery across the remaining test sessions. In addition, subjects generally showed improvement across days. We found significant main effects for day and time for the composite equilibrium score and for sensory organization tests (SOT) 1, 2 and 6. Significant main effects were observed for day for SOT 3 and 5. Although we found no significant main effects for gender (when center of gravity was used as a covariate), we did observe significant gender X time interaction effects for composite equilibrium and for SOT 1, 3, 4 and 5. Women appeared to show larger decrements in postural stability immediately after exposure to VR than men, but recover more quickly than n. Finally, we found no significant main effects for type of VR device or for exposure duration, however, these factors did interact with other factors during some of the SOTs. Subjects exhibited rapid recovery of motion sickness symptoms across time following exposure to VR and significantly less severe symptoms across days. We did not observe main effects for gender, type of device or duration of exposure. Individuals recovered from the detrimental effects of exposure to virtual reality on postural control and motion sickness within one hour. Sickness severity and initial decrements in postural equilibrium decreases over days, which suggests that subjects become dual-adapted over time. These findings provide some direction for developing training schedules for VR users that facilitate adaptation, and support the idea that preflight training of astronauts may serve as useful countermeasure for the sensorimotor effects of space flight.

Harm, D. L.↗

E-Glider: Active Electrostatic Flight for Airless Body Exploration

The environment near the surface of asteroids, comets, and the Moon is electrically charged due to the Sun's photoelectric bombardment and lofting dust, which follows the Sun illumination as the body spins. Chargeddust is ever present, in the form of dusty plasma, even at high altitudes, following the solar illumination. If abody with high surface resistivity is exposed to the solar wind and solar radiation, sun-exposed areas andshadowed areas become differentially charged. The E-Glider (Electrostatic Glider) is an enabling capability foroperation at airless bodies, a solution applicable to many types of in-situ mission concepts, which leverages thenatural environment. With the E-Glider, we transform a problem (spacecraft charging) into an enablingtechnology, i.e. a new form of mobility in microgravity environments using new mechanisms and maneuveringbased on the interaction of the vehicle with the environment. Consequently, the vision of the E-Glider is toenable global scale airless body exploration with a vehicle that uses, instead of avoids, the local electricallycharged environment. This platform directly addresses the "All Access Mobility" Challenge, one of the NASA'sSpace Technology Grand Challenges. Exploration of comets, asteroids, moons and planetary bodies is limitedby mobility on those bodies. The lack of an atmosphere, the low gravity levels, and the unknown surface soilproperties pose a very difficult challenge for all forms of know locomotion at airless bodies. This E-Gliderlevitates by extending thin, charged, appendages, which are also articulated to direct the levitation force in themost convenient direction for propulsion and maneuvering. The charging is maintained through continuouscharge emission. It lands, wherever it is most convenient, by retracting the appendages or by firing a cold-gasthruster, or by deploying an anchor. The wings could be made of very thin Au-coated Mylar film, which areelectrostatically inflated, and would provide the lift due to electrostatic repulsion with the naturally chargedasteroid surface. Since the E-glider would follow the Sun's illumination, the solar panels on the vehicle wouldconstantly charge a battery. Further articulation at the root of the lateral strands or inflated membrane wings,would generate a component of lift depending on the articulation angle, hence a selective maneuveringcapability which, to all effects, would lead to electrostatic (rather than aerodynamic) flight. Preliminarycalculations indicate that a 1 kg mass can be electrostatically levitated in a microgravity field with a 2 mdiameter electrostatically inflated ribbon structure at 19kV, hence the need for a "balloon-like" system. Due tothe high density and the photo-electron sheath and associate small Debye length, significant power is requiredto levitate even a few kilograms. The power required is in the kilo-Watt range to maintain a constant chargelevel.

Electrostatic Flight↗