Search NASA⌕ Search

SEARCH · Search NASA

Results for “gated convolution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Parallelized Convolutional Interleaver Implementation for Efficient DDR Memory Access

Convolutional interleavers are used in many different communications systems to correct for burst errors due to atmospheric fades and scintillation. The interleaver size is related to the channel coherence time and the data rate. Small convolutional interleavers can be implemented in a field programmable gate array (FPGA) block random access memory (BRAM). However, large interleavers exceeding the size of the BRAM on the FPGA are necessary for channels with longer fades and higher data rates. Therefore, an implementation utilizing double data rate (DDR) memory external to the FPGA is necessary. Wide DDR memory data buses can make the use of DDR memory for convolutional interleavers inefficient when individual symbols are written to and read from the memory. DDR memory operational speeds can also limit the data rate of the interleaver. The Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency (HPE) standard utilizes a convolutional channel symbol interleaver. A previous implementation of the HPE standard utilized BRAM for the convolutional interleaver, but mission requirements for the upcoming Optical Artemis-2 Orion (O2O) communications demonstration dictate the use of an interleaver exceeding the size of the BRAM. An algorithm and method for implementing the convolutional interleaver in the FPGA with DDR memory is described in this paper.

optical communications↗

Parallelized convolutional interleaver implementation for efficient DDR memory access

Convolutional interleavers are used in many different communications systems to correct for burst errors due to atmospheric fades and scintillation. The interleaver size is related to the channel coherence time and the data rate. Small convolutional interleavers can be implemented in a field programmable gate array (FPGA) block random access memory (BRAM). However, large interleavers exceeding the size of the BRAM on the FPGA are necessary for channels with longer fades and higher data rates. Therefore, an implementation utilizing double data rate (DDR) memory external to the FPGA is necessary. Wide DDR memory data buses can make the use of DDR memory for convolutional interleavers inefficient when individual symbols are written to and read from the memory. DDR memory operational speeds can also limit the data rate of the interleaver. The Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency (HPE) standard utilizes a convolutional channel symbol interleaver. A previous implementation of the HPE standard utilized BRAM for the convolutional interleaver, but mission requirements for the upcoming Optical Artemis-2 Orion (O2O) communications demonstration dictate the use of an interleaver exceeding the size of the BRAM. An algorithm and method for implementing the convolutional interleaver in the FPGA with DDR memory is described in this paper.

Optical communications↗

Flight Trajectory Prediction Based on Hybrid-Recurrent Networks

The development of future technologies for the National Airspace System (NAS) will be reliant on a new communications infrastructure capable of managing the limited available spectrum for communications among aircraft and ground systems. Emerging approaches to autonomous allocation of aviation spectrum mostlyrely on machine learning techniques, where 4D (longitude, latitude, altitude, time) trajectory prediction is an important data input to enable real-time resource allocation. This study explores and evaluates effective data sources and deep recurrent neural network techniques when determining flight trajectories. Specifically, data are collected and evaluated in a 100-day and 14-day period. Sources of data include NASA Sherlock Data Warehouse, MIT Lincoln Labs Corridor Integrated Weather Service (CIWS), and assorted NOAA weather datasets. Deep learning models for 4D predictions all utilize a hybrid-recurrent technique. A baseline model is considered via the convolutional-LSTM design from the existing literature. The modified design considers Gated Recurrent Units (GRU), Independently Recurrent Neural Networks (IndRNN), and stand-alone self-attention layers. Results indicatethe effectiveness of LSTM and GRUcells for state-of-the-art data processing (interpolation). Additionally, GRUs may be quickly trained with limited data, allowing for exacting improvements with optimizer selection. Attention mechanisms provide notable performance improvements to convolutional layers and may extend dimensional capabilities of a learning model. Finally, NOAA measurements provide only a supplemental value, requiring support from tailored measurements for Air Traffic Management.

Nathan Schimpf↗

Analysis of GEOS-3 altimeter data and extraction of ocean wave height and dominant wavelength

When the amplitude and timing biases are removed from the GEOS-3 Sample and Hold (S&H) gates, the mean return waveforms can be excellently fitted with a theoretical template which represents the convolution of: (1) the radar point target response; (2) the range noise (jitter) in the altimeter tracking loop; (3) the sea surface height distribution; and (4) the antenna pattern as a function of the range to mean sea level. Several techniques of varying complexity to remove the effect of the tracking loop jitter in computing the wave height are considered. They include: (1) realigning the S&H gates to their actual positions with respect to mean sea level before averaging; (2) using the observed standard deviation on the altitude measurement to remove the integrated effect of the tracking loop jitter, and (3) using a look-up table to correct for the expected value of range noise. Analysis of skewness in the GEOS return waveform demonstrates the potential of a satellite radar altimeter to determine the dominant wavelength of ocean waves.

Walsh, E. J.↗

Encoders for block-circulant LDPC codes

In this paper, we present two encoding methods for block-circulant LDPC codes. The first is an iterative encoding method based on the erasure decoding algorithm, and the computations required are well organized due to the block-circulant structure of the parity check matrix. The second method uses block-circulant generator matrices, and the encoders are very similar to those for recursive convolutional codes. Some encoders of the second type have been implemented in a small Field Programmable Gate Array (FPGA) and operate at 100 Msymbols/second.

encoders↗

Add/Compare/Select Circuit For Rapid Decoding

Prototype decoding system operates at 200 Mb/s. ACS (add/compare/select) gate array is highly integrated emitter-coupled-logic circuit implementing arithmetic operations essential to Viterbi decoding of convolutionally encoded data signals. Principal advantage of circuit is speed. Operates as single unit performing eight additions and finds minimum of eight sums, or operates as two independent units, each performing four additions and finding minimum of four sums. Flexibility enables application to variety of different codes. Includes built-in self-testing circuitry, enabling unit to be tested at full speed with help of only simple test fixture.

Budinger, James M.↗

JTRS/SCA and Custom/SDR Waveform Comparison

This paper compares two waveform implementations generating the same RF signal using the same SDR development system. Both waveforms implement a satellite modem using QPSK modulation at 1M BPS data rates with one half rate convolutional encoding. Both waveforms are partitioned the same across the general purpose processor (GPP) and the field programmable gate array (FPGA). Both waveforms implement the same equivalent set of radio functions on the GPP and FPGA. The GPP implements the majority of the radio functions and the FPGA implements the final digital RF modulator stage. One waveform is implemented directly on the SDR development system and the second waveform is implemented using the JTRS/SCA model. This paper contrasts the amount of resources to implement both waveforms and demonstrates the importance of waveform partitioning across the SDR development system.

Oldham, Daniel R.↗

Constructing LDPC Codes from Loop-Free Encoding Modules

A method of constructing certain low-density parity-check (LDPC) codes by use of relatively simple loop-free coding modules has been developed. The subclasses of LDPC codes to which the method applies includes accumulate-repeat-accumulate (ARA) codes, accumulate-repeat-check-accumulate codes, and the codes described in Accumulate-Repeat-Accumulate-Accumulate Codes (NPO-41305), NASA Tech Briefs, Vol. 31, No. 9 (September 2007), page 90. All of the affected codes can be characterized as serial/parallel (hybrid) concatenations of such relatively simple modules as accumulators, repetition codes, differentiators, and punctured single-parity check codes. These are error-correcting codes suitable for use in a variety of wireless data-communication systems that include noisy channels. These codes can also be characterized as hybrid turbolike codes that have projected graph or protograph representations (for example see figure); these characteristics make it possible to design high-speed iterative decoders that utilize belief-propagation algorithms. The present method comprises two related submethods for constructing LDPC codes from simple loop-free modules with circulant permutations. The first submethod is an iterative encoding method based on the erasure-decoding algorithm. The computations required by this method are well organized because they involve a parity-check matrix having a block-circulant structure. The second submethod involves the use of block-circulant generator matrices. The encoders of this method are very similar to those of recursive convolutional codes. Some encoders according to this second submethod have been implemented in a small field-programmable gate array that operates at a speed of 100 megasymbols per second. By use of density evolution (a computational- simulation technique for analyzing performances of LDPC codes), it has been shown through some examples that as the block size goes to infinity, low iterative decoding thresholds close to channel capacity limits can be achieved for the codes of the type in question having low maximum variable node degrees. The decoding thresholds in these examples are lower than those of the best-known unstructured irregular LDPC codes constrained to have the same maximum node degrees. Furthermore, the present method enables the construction of codes of any desired rate with thresholds that stay uniformly close to their respective channel capacity thresholds.

Divsalar, Dariush↗

Effects of data asymmetry on Shuttle Ku-band communications link performance

This paper systematically analyzes the signal-to-noise ratio degradations which can potentially occur due to data asymmetry in digital transmission systems. Suitable asymmetry models are developed and error probability performance for various types of data detectors (integrate-and-dump filter, filter-sample detector, and gated-integrate-and-dump filter) is derived. Although this work was done to resolve problems being encountered in the Shuttle Ku-band return link design, specifically for the 50 Mbit/s convolutionally encoded channel (NRZ format), generalizations are made which provide results for other cases of interest (other Ku-band return link channels, or other systems entirely). This paper therefore considers Manchester data formats (in addition to NRZ) and uncoded transmission (in addition to convolutionally coded transmission). The effects of bandlimiting are also considered.

Simon, M. K.↗

Independent synchronizer for digital decoders

Logic circuit synchronizes branches of any convolution code-decoder at low signal to noise ratios. Parity checks determine correct node synchronization. Device maintains synchrony as low as -3 dB. Circuit consists of 15 stage shift register, three up down counters, and some logic gates.

Stiffler, J. J.↗

Field-Programmable Gate Array Implementation of a Single Photon-Counting Receive Modem

We present a field-programmable gate array (FPGA) implementation of a single photon-counting receive modem for a pulse position modulated signal. The modem is compliant with the Consultative Committee for Space Data Systems (CCSDS) High Photon Efficiency (HPE) Optical Communications Coding and Synchronization standard and is capable of a maximum data rate of 267 Mbps. The system is designed on a commercial off-the-shelf FPGA platform and utilizes superconducting nanowire single photon counting detectors, analog to digital converters (ADC s) to sample the detectors, and two FPGAs. Symbol timing recovery, photon counting, convolutional deinterleaving, and codeword synchronization areis performed in the first FPGA. The second FPGA performs iterative decoding on each codeword of the serially concatenated pulse position modulated (SCPPM) signal. A digital filter is included to compensate for timing jitter of the detector, and the decoder throughput can be modified adjusted through reconfigurable parallelization. The decoder also implements a resource-efficient, algorithmic polynomial interleaver and deinterleaver. Both FPGAs can be reconfigured to switch between pulse position modulation (PPM)-16 and PPM-32 with code rates 1/3, 1/2, and 2/3. In this paper, we describe the receiver architecture and FPGA implementation of the timing recovery loop and SCPPM decoder, FPGA utilization for the different modes, and receive modem characterization test results.

Field-programmable Gate Array↗

Fast Image Texture Classification Using Decision Trees

Texture analysis would permit improved autonomous, onboard science data interpretation for adaptive navigation, sampling, and downlink decisions. These analyses would assist with terrain analysis and instrument placement in both macroscopic and microscopic image data products. Unfortunately, most state-of-the-art texture analysis demands computationally expensive convolutions of filters involving many floating-point operations. This makes them infeasible for radiation- hardened computers and spaceflight hardware. A new method approximates traditional texture classification of each image pixel with a fast decision-tree classifier. The classifier uses image features derived from simple filtering operations involving integer arithmetic. The texture analysis method is therefore amenable to implementation on FPGA (field-programmable gate array) hardware. Image features based on the "integral image" transform produce descriptive and efficient texture descriptors. Training the decision tree on a set of training data yields a classification scheme that produces reasonable approximations of optimal "texton" analysis at a fraction of the computational cost. A decision-tree learning algorithm employing the traditional k-means criterion of inter-cluster variance is used to learn tree structure from training data. The result is an efficient and accurate summary of surface morphology in images. This work is an evolutionary advance that unites several previous algorithms (k-means clustering, integral images, decision trees) and applies them to a new problem domain (morphology analysis for autonomous science during remote exploration). Advantages include order-of-magnitude improvements in runtime, feasibility for FPGA hardware, and significant improvements in texture classification accuracy.

Thompson, David R.↗

Performance of A Real-Time Photon Counting Optical Receiver in the Presence of Emulated Channel Fading

Free-space optical communication links with terrestrial ground stations experience fading due to atmospheric scintillation and beam pointing. Fiber-coupled receiver systems experience additional fading at the interface between the fiber and free-space optics of the telescope. The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) has characterized a real-time photon-counting optical ground receiver system with an atmospheric fade emulation system. The receiver system is comprised of a fiber interconnect, an array of superconducting nanowire single photon detectors (SNSPDs), and a field programmable gate array (FPGA) based receive modem. Two fiber interconnect/detector architectures have been studied. One architecture uses a 70-mode photonic lantern coupled to seven single pixel SNSPDs. The other architecture uses a 10-mode few-mode fiber (FMF) coupled to a 15-pixel SNSPD array. The receiver system complies with the Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency Coding and Synchronization Standard, which uses serially concatenated convolutionally coded pulse-position modulation (SCPPM). The CCSDS standard is designed for use in low photon flux missions, including the Orion Artemis-II Optical (O2O) communications demonstration. The standard utilizes a convolutional symbol interleaver which can be resized to mitigate different fades. The fade emulation system employed in this work emulates scintillation-induced, pointing-induced, and coupling-induced fading. This paper gives an overview of the real-time optical receiver system and the fade emulation system. It presents tests results which show the impact of fading on the performance on the receiver. The test results show that in the presence of channel fading, the 70-mode photonic lantern outperforms the 10-mode FMF under higher (D/r_0=9) turbulence conditions due to high fiber-coupling-induced fading and fiber coupling loss on the 10-mode FMF. When operating in lower turbulence (D/r_0=4), the 10-mode FMF outperforms the 70-mode photonic lantern. The paper also shows a larger convolutional interleaver improves the system performance as long as the receiver does not lose acquisition.

optical communications↗

Performance of a real-time photon counting optical receiver in the presence of emulated channel fading

Free-space optical communication links with terrestrial ground stations experience fading due to atmospheric scintillation and beam pointing. Fiber-coupled receiver systems experience additional fading at the interface between the fiber and free-space optics of the telescope. The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) has characterized a real-time photon-counting optical ground receiver system with an atmospheric fade emulation system. The receiver system is comprised of a fiber interconnect, an array of superconducting nanowire single photon detectors (SNSPDs), and a field programmable gate array (FPGA) based receive modem. Two fiber interconnect/detector architectures have been studied. One architecture uses a 70-mode photonic lantern coupled to seven single pixel SNSPDs. The other architecture uses a 10-mode few-mode fiber (FMF) coupled to a 15-pixel SNSPD array. The receiver system complies with the Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency Coding and Synchronization Standard, which uses serially concatenated convolutionally coded pulse-position modulation (SCPPM). The CCSDS standard is designed for use in low photon flux missions, including the Orion Artemis-II Optical (O2O) communications demonstration. The standard utilizes a convolutional symbol interleaver which can be resized to mitigate different fades. The fade emulation system employed in this work emulates scintillation-induced, pointing-induced, and coupling-induced fading. This paper gives an overview of the real-time optical receiver system and the fade emulation system. It presents tests results which show the impact of fading on the performance on the receiver. The test results show that in the presence of channel fading, the 70-mode photonic lantern outperforms the 10-mode FMF under higher (D/r_0=9) turbulence conditions due to high fiber-coupling-induced fading and fiber coupling loss on the 10-mode FMF. When operating in lower turbulence (D/r_0=4), the 10-mode FMF outperforms the 70-mode photonic lantern. The paper also shows a larger convolutional interleaver improves the system performance as long as the receiver does not lose acquisition.

optical communications↗

Hardware Implementation of Serially Concatenated PPM Decoder

A prototype decoder for a serially concatenated pulse position modulation (SCPPM) code has been implemented in a field-programmable gate array (FPGA). At the time of this reporting, this is the first known hardware SCPPM decoder. The SCPPM coding scheme, conceived for free-space optical communications with both deep-space and terrestrial applications in mind, is an improvement of several dB over the conventional Reed-Solomon PPM scheme. The design of the FPGA SCPPM decoder is based on a turbo decoding algorithm that requires relatively low computational complexity while delivering error-rate performance within approximately 1 dB of channel capacity. The SCPPM encoder consists of an outer convolutional encoder, an interleaver, an accumulator, and an inner modulation encoder (more precisely, a mapping of bits to PPM symbols). Each code is describable by a trellis (a finite directed graph). The SCPPM decoder consists of an inner soft-in-soft-out (SISO) module, a de-interleaver, an outer SISO module, and an interleaver connected in a loop (see figure). Each SISO module applies the Bahl-Cocke-Jelinek-Raviv (BCJR) algorithm to compute a-posteriori bit log-likelihood ratios (LLRs) from apriori LLRs by traversing the code trellis in forward and backward directions. The SISO modules iteratively refine the LLRs by passing the estimates between one another much like the working of a turbine engine. Extrinsic information (the difference between the a-posteriori and a-priori LLRs) is exchanged rather than the a-posteriori LLRs to minimize undesired feedback. All computations are performed in the logarithmic domain, wherein multiplications are translated into additions, thereby reducing complexity and sensitivity to fixed-point implementation roundoff errors. To lower the required memory for storing channel likelihood data and the amounts of data transfer between the decoder and the receiver, one can discard the majority of channel likelihoods, using only the remainder in operation of the decoder. This is accomplished in the receiver by transmitting only a subset consisting of the likelihoods that correspond to time slots containing the largest numbers of observed photons during each PPM symbol period. The assumed number of observed photons in the remaining time slots is set to the mean of a noise slot. In low background noise, the selection of a small subset in this manner results in only negligible loss. Other features of the decoder design to reduce complexity and increase speed include (1) quantization of metrics in an efficient procedure chosen to incur no more than a small performance loss and (2) the use of the max-star function that allows sum of exponentials to be computed by simple operations that involve only an addition, a subtraction, and a table lookup. Another prominent feature of the design is a provision for access to interleaver and de-interleaver memory in a single clock cycle, eliminating the multiple clock-cycle latency characteristic of prior interleaver and de-interleaver designs.

Moision, Bruce↗

High-speed architecture for the decoding of trellis-coded modulation

Since 1971, when the Viterbi Algorithm was introduced as the optimal method of decoding convolutional codes, improvements in circuit technology, especially VLSI, have steadily increased its speed and practicality. Trellis-Coded Modulation (TCM) combines convolutional coding with higher level modulation (non-binary source alphabet) to provide forward error correction and spectral efficiency. For binary codes, the current stare-of-the-art is a 64-state Viterbi decoder on a single CMOS chip, operating at a data rate of 25 Mbps. Recently, there has been an interest in increasing the speed of the Viterbi Algorithm by improving the decoder architecture, or by reducing the algorithm itself. Designs employing new architectural techniques are now in existence, however these techniques are currently applied to simpler binary codes, not to TCM. The purpose of this report is to discuss TCM architectural considerations in general, and to present the design, at the logic gate level, or a specific TCM decoder which applies these considerations to achieve high-speed decoding.

Osborne, William P.↗

Optimizations of a Hardware Decoder for Deep-Space Optical Communications

The National Aeronautics and Space Administration has developed a capacity approaching modulation and coding scheme that comprises a serial concatenation of an inner accumulate pulse-position modulation (PPM) and an outer convolutional code [or serially concatenated PPM (SCPPM)] for deep-space optical communications. Decoding of this code uses the turbo principle. However, due to the nonbinary property of SCPPM, a straightforward application of classical turbo decoding is very inefficient. Here, we present various optimizations applicable in hardware implementation of the SCPPM decoder. More specifically, we feature a Super Gamma computation to efficiently handle parallel trellis edges, a pipeline-friendly 'maxstar top-2' circuit that reduces the max-only approximation penalty, a low-latency cyclic redundancy check circuit for window-based decoders, and a high-speed algorithmic polynomial interleaver that leads to memory savings. Using the featured optimizations, we implement a 6.72 megabits-per-second (Mbps) SCPPM decoder on a single field-programmable gate array (FPGA). Compared to the current data rate of 256 kilobits per second from Mars, the SCPPM coded scheme represents a throughput increase of more than twenty-six fold. Extension to a 50-Mbps decoder on a board with multiple FPGAs follows naturally. We show through hardware simulations that the SCPPM coded system can operate within 1 dB of the Shannon capacity at nominal operating conditions.

quadratic polynomial interleaver↗

A VLSI decomposition of the deBruijn graph

A new Viterbi decoder for convolutional codes with constraint lengths up to 15, called the Big Viterbi Decoder, is under development for the Deep Space Network. It will be demonstrated by decoding data from the Galileo spacecraft, which has a rate 1/4, constraint-length 15 convolutional encoder on board. Here, the mathematical theory underlying the design of the very-large-scale-integrated (VLSI) chips that are being used to build this decoder is explained. The deBruijn graph B sub n describes the topology of a fully parallel, rate 1/v, constraint length n+2 Viterbi decoder, and it is shown that B sub n can be built by appropriately wiring together (i.e., connecting together with extra edges) many isomorphic copies of a fixed graph called a B sub n building block. The efficiency of such a building block is defined as the fraction of the edges in B sub n that are present in the copies of the building block. It is shown, among other things, that for any alpha less than 1, there exists a graph G which is a B sub n building block of efficiency greater than alpha for all sufficiently large n. These results are illustrated by describing a special hierarchical family of deBruijn building blocks, which has led to the design of the gate-array chips being used in the Big Viterbi Decoder.

Collins, O.↗