Search NASA⌕ Search

SEARCH · Search NASA

Results for “dot product”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Mode-multiplexed photonic integrated vector dot-product core from inverse design

Photonic computing has the potential to harness the full degrees of freedom (DOFs) of the light field, including the wavelength, spatial mode, spatial location, phase quadrature, and polarization, to achieve a higher level of computing parallelism and scalability than digital electronic processors. While multiplexing using the wavelength and other DOFs can be readily integrated on silicon photonics platforms with compact footprints, conventional mode-division multiplexed (MDM) photonic designs occupy areas exceeding tens to hundreds of microns for a few spatial modes, significantly limiting their scalability. Here, we utilize inverse design to demonstrate an ultracompact photonic computing core that calculates vector dot products based on MDM coherent mixing. Our dot-product core integrates the functionalities of two-mode multiplexers and one multimode coherent mixer within a nominal footprint of 5 μm x 3 μm . We have experimentally demonstrated computing examples on the fabricated dot-product core, including complex number multiplication and motion estimation using optical flow. The compact dot-product core design enables large-scale on-chip integration in a parallel photonic computing primitive cluster for high-throughput scientific computing and computer vision tasks.

97 MATHEMATICS AND COMPUTING↗

A Framework for Error-Bounded Approximate Computing, with an Application to Dot Products

Approximate computing techniques, which trade off the computation accuracy of an algorithm for better performance and energy efficiency, have been successful in reducing computation and power costs in several domains. However, error sensitive applications in high-performance computing are unable to benefit from existing approximate computing strategies that are not developed with guaranteed error bounds. While approximate computing techniques can be developed for individual high-performance computing applications by domain specialists, this often requires additional theoretical analysis and potentially extensive software modification. Hence, the development of low-level error-bounded approximate computing strategies that can be introduced into any high-performance computing application without requiring additional analysis or significant software alterations is desirable. In this paper, we provide a contribution in this direction by proposing a general framework for designing error-bounded approximate computing strategies and apply it to the dot product kernel to develop \bf qdot---an error-bounded approximate dot product kernel. Following the introduction of qdot, here we perform a theoretical analysis that yields a deterministic bound on the relative approximation error introduced by qdot. Empirical tests are performed to illustrate the tightness of the derived error bound and to demonstrate the effectiveness of qdot on a synthetic dataset, as well as two scientific benchmarks---the conjugate gradient (CG) and power methods. In some instances, using qdot for the dot products in CG can result in many components being quantized to half precision without increasing the iteration count required for convergence to the same solution as CG using a double precision dot product.

97 MATHEMATICS AND COMPUTING↗

sKokkos: Enabling Kokkos with Transparent Device Selection on Heterogeneous Systems using OpenACC

This paper presents a new feature to enable Kokkos with transparent device selection. For application developers, it is not easy toidentify which device is the most appropriate to use in a heterogeneous system, since this depends on the characteristics of both the application and the hardware. In Kokkos, a backend is associated with one specific programming model/hardware. Programmers decide which backend to use at compilation time. This new feature implemented on the OpenACC backend eliminates the burden of deciding which device to use, providing a highly productive programming solution for Kokkos applications. This work includes implementation details and a performance study conducted with a set of mini-benchmarks (i.e., AXPY and dot product), kernels (Lattice-Bolzmann method), and two mini-apps (LULESH and miniFE) on two heterogeneous systems with different hardware capabilities. This new Kokkos feature provides high accelerations of up to 35× thanks to automatic and transparent device selection.

Lee, Seyong↗

KokkACC: Enhancing Kokkos with OpenACC

Template metaprogramming is gaining popularity as a high-level solution for achieving performance portability on heterogeneous computing resources. Kokkos is a representative approach that offers programmers high-level abstractions for generic programming while most of the device-specific code generation and optimizations are delegated to the compiler through template specializations. For this, Kokkos provides a set of device-specific code specializations in multiple back ends, such as CUDA and HIP. Unlike CUDA or HIP, OpenACC is a high-level and directive-based programming model. This descriptive model allows developers to insert hints (pragmas) into their code that help the compiler to parallelize the code. The compiler is responsible for the transformation of the code, which is completely transparent to the programmer. This paper presents an OpenACC back end for Kokkos: KokkACC. As an alternative to Kokkos’s existing device-specific back ends, KokkACC is a multi-architecture back end providing a high-productivity programming environment enabled by OpenACC’s high-level and descriptive programming model. Moreover, we have observed competitive performance; in some cases, KokkACC is faster (up to 9×) than NVIDIA’s CUDA back end and much faster than OpenMP’s GPU offloading back end. This work also includes implementation details and a detailed performance study conducted with a set of mini-benchmarks (AXPY and DOT product) and three mini-apps (LULESH, miniFE and SNAP, a LAMMPS proxy mini-app).

Valero Lara, Pedro↗

Toward efficient polynomial preconditioning for GMRES

Here, we present a polynomial preconditioner for solving large systems of linear equations. The polynomial is derived from the minimum residual polynomial (the GMRES polynomial) and is more straightforward to compute and implement than many previous polynomial preconditioners. Our current implementation of this polynomial using its roots is naturally more stable than previous methods of computing the same polynomial. We implement further stability control using added roots, and this allows for high degree polynomials. We discuss the effectiveness and challenges of root-adding and give an additional check for stability. In this article, we study the polynomial preconditioner applied to GMRES; however it could be used with any Krylov solver. This polynomial preconditioning algorithm can dramatically improve convergence for some problems, especially for difficult problems, and can reduce dot products by an even greater margin.

97 MATHEMATICS AND COMPUTING↗

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

Vector-Matrix Multiplication Engine for Neuromorphic Computation with a CBRAM Crossbar Array [Slides]

The core function of many neural network algorithms is the dot product, or vector matrix multiply (VMM) operation. Crossbar arrays utilizing resistive memory elements can reduce computational energy in neural algorithms by up to five orders of magnitude compared to conventional CPUs. Moving data between a processor, SRAM, and DRAM dominates energy consumption. By utilizing analog operations to reduce data movement, resistive memory crossbars can enable processing of large amounts of data at lower energy than conventional memory architectures.

97 MATHEMATICS AND COMPUTING↗

Linking emergent phenomena and broken symmetries through one-dimensional objects and their dot/cross products

Abstract The symmetry of the whole experimental setups, including specific sample environments and measurables, can be compared with that of specimens for observable physical phenomena. We, first, focus on one-dimensional (1D) experimental setups, independent from any spatial rotation around one direction, and show that eight kinds of 1D objects (four; vector-like, the other four; director-like), defined in terms of symmetry, and their dot and cross products are an effective way for the symmetry consideration. The dot products form a Z 2 × Z 2 × Z 2 group with Abelian additive operation, and the cross products form a Z 2 × Z 2 group with Abelian additive operation or Q 8 , a non-Abelian group of order eight, depending on their signs. Those 1D objects are associated with characteristic physical phenomena. When a 3D specimen has symmetry operational similarity (SOS) with (identical or lower, but not higher, symmetries than) an 1D object with a particular phenomenon, the 3D specimen can exhibit the phenomenon. This SOS approach can be a transformative and unconventional avenue for symmetry-guided materials designs and discoveries.

Physics↗

Non-equilibrium entropy production and information dissipation in a non-Markovian quantum dot

This study measures trajectory-level entropy production and information dissipation in a driven, non-Markovian quantum dot using time-resolved optical dynamics and machine-learning-based analysis. Although not a 2D-material system, it is relevant because it demonstrates quantitative extraction of nonequilibrium dynamics from nanoscale optical fluctuations, which is conceptually connected to the proposed studies of transient charge and spin dynamics at interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Environmentally Friendly Production of High-Quality and Multifunctional Carbon Quantum Dots from Coal

Researchers from the University of Wyoming and the University of Utah worked jointly to study the valorization of coal to carbon quantum dots (CQDs), a value-added product with a broad spectrum of applications. The CQDs were produced by an environmentally facile hydrothermal method, and the experimental factors influencing the properties of CQDs were investigated. We subsequently explored the applications of CQDs as co-sensitizers of dye-sensitized solar cells (DSSC) and photocatalysts of water treatment. As an outlook, techno-economic and environmental analysis studied the feasibility of mass production of CQDs.

01 COAL, LIGNITE, AND PEAT↗

Trace level of atomic copper in N-doped graphene quantum dots switching the selectivity from C 1 to C 2 products in CO electroreduction

To unravel the relationship between trace Cu on the metal-free catalysts toward CO/CO 2 reduction reaction (CO/CO 2 RR), we investigated the effect of trace Cu loading in N-doped graphene quantum dots (NGQDs) on CO/CO 2 RR. A general trend is that increasing the Cu loading in NGQDs switches the selectivity from C 1 (CH 4 ) to C 2 products in CORR. When 2.5 μg/cm 2 Cu with the atomic size is loaded on NGQDs, the selectivity shifts from 62% Faradaic efficiency (FE) of CH 4 to 52% FE of C 2 products in CORR. Further increasing the atomic Cu loading to 3.8 μg/cm 2 promotes the FE of C 2 products to 78%. CO 2 RR requires one order of magnitude higher Cu loading than CORR to switch the selectivity from C 1 to C 2 products due to the low partial pressure of CO. Finally, this study clarifies the distinct impact of trace (ppm level) Cu on the activity/selectivity between CORR and CO 2 RR.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Alternate InP synthesis with aminophosphines: solution–liquid–solid nanowire growth

Indium phosphide nanowires are important components in high-speed electronics and optoelectronics, including photodetectors and photovoltaics. However, most syntheses either use high-temperature and costly vapor-phase methodology or highly toxic and pyrophoric tris(trimethylsilyl)phosphine. To expand on the success of the aminophosphine-based InP colloidal quantum dot synthesis, we developed a synthesis for thin (~11 nm) zinc blende InP nanowires at 180 °C using indium tris(trifluoroacetate) and tris(diethylamino)phosphine. A flat nanoribbon morphology was identified by transmission electron and atomic force microscopy analysis, with the stoichiometric (110) lattice plane exposed. Nanowire growth proceeded through a solution–liquid–solid mechanism from in situ-formed indium metal nanoparticles. Molecular byproducts of tris(oleylamino)phosphine oxide and N-oleyltrifluoroacetamide observed by 31 P and 19 F NMR spectroscopy inform a proposed mechanism of indium reduction by the aminophosphine. Morphological control over the nanowire product was achieved by varying the phosphorus injection to control the aspect ratio, the In : P ratio to toggle between nanowires and multipods, and the pre-hot injection evacuation step to favor a quantum dot product. Furthermore, replacing the indium precursor with indium tris(trifluoromethanesulfonate) was found to make bulk zinc blende InP nanowires with an average diameter of >250 nm and tens of microns in length.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Gradient synthesis of carbon quantum dots and activated carbon from pulp black liquor for photocatalytic hydrogen evolution and supercapacitor

Black liquor (BL) is a by-product of the chemical pulping industry and is mainly used as a low-value fuel; however, its potential to produce high-value products has not been fully exploited. In this study, a green and simple strategy is reported for the gradient production of Na + -functionalized carbon quantum dots (Na + -CQDs) for the first time, N and S co-doped CQDs (N/S-CQDs), and N and S co-doped KOH-activated carbon (N/S-KAC) from BL by dialysis, hydrothermal carbonization and activation-carbonization, respectively. Due to the good electron trapping ability, photoluminescence and promising up-conversion luminescence of CQDs, the hydrogen evolution efficiency of Na + -CQDs/TiO 2 and N/S-CQDs/TiO 2 photocatalysts was improved by 2.45 and 1.46 times, respectively, compared with pure TiO 2 . N/S-KAC with a high specific surface area of 2294 m 2 g -1 provides an excellent specific capacitance of 253 F g -1 at 0.5 A g -1 and a promising energy density of 26.92 Wh kg -1 under a power density of 566 W kg -1 for the fabricated symmetrical supercapacitor. Moreover, the electrode material has good cycling stability with a capacitance retention of ~ 93.91% after 5000 cycles. In conclusion, this pathway provides a versatile and scalable approach for the construction and co-production of nanostructured materials, photocatalysts and energy storage devices.

36 MATERIALS SCIENCE↗

Spacetime pq theory for AC and DC electric power systems

The 50/60 Hz alternating current (AC) electric power has been the standard and most flexible energy source powering our modern societies for one and a half centuries since the war of the currents: AC versus direct current (DC). A reactive power concept that was introduced at the beginning of the AC power was very useful for circuit/system analysis, design, control, optimization, and ultimately for more efficient and stable generation, transmission, distribution, and consumption. The initial reactive power theory was based on single-phase sinusoidal AC power to capture inductive and capacitive power that yields to net-zero average power over one fundamental cycle. Soon it was expanded to non-sinusoidal AC power and finally to instantaneous three-phase AC power. However, these reactive power theories remain separate and limited to special cases and have never been consolidated and made valid to all cases. Today, more widespread adoption of power electronics and renewable energy is bringing back DC power into the electric grids. The reactive power concept has never been applied to DC power systems. There is no reactive power in DC power systems according to the existing reactive power theories. Do DC power systems really have no reactive power? Capacitors and inductors are widely used in DC just like in AC power systems. Are they not reactive power components? Why are they different from their AC counterparts? Furthermore, are batteries active or reactive power components? What about active devices like power converters (or inverters) with AC (or DC) on one side and DC (or AC) on the other? Do they generate or consume reactive power? Finally, what about AC and DC hybrid power systems? How to define reactive power in such a complex power system that has a multitude of loads, buses, and sources? Is there reactive power between any two loads, any two buses, or any two sources in a power system and what is the total reactive power in such a complex power system as a whole? As the motivation and goal of this paper to answer the above basic questions, to unify the existing AC reactive power theories and to ultimately provide theoretical and insightful guidance for system analysis, design, control, efficiency, optimization, and operation of complex power systems, a concept of spacetime (both spatial and temporal) active and reactive power (pq) theory—the spatiotemporal aspect of active and reactive power—is developed for both AC and DC power systems. The theoretical definitions and physical meanings of the spacetime reactive power will be developed, and real applications and thought experiments/cases/exercises will be explored and discussed. The developed mathematics to define the active (or real) and reactive (or imaginary) power— p and q respectively by dot (scalar) and cross (vector) products of multi-dimension spacetime vectors and time-space mapping principle/law can have some fundamental implications as well.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Tailoring electron transfer pathway for photocatalytic N 2 -to-NH 3 reduction in a CdS quantum dots-nitrogenase system

The combination of abiotic photosensitizers with purified enzymes in a biohybrid system offers a promising pathway to utilizing light to accomplish challenging chemical transformations and provides insights into the rational photocatalytic system design for efficient solar-to-chemical energy conversion. In this work, we demonstrate a hybrid photocatalytic system for ammonia production from N2 by combining cadmium sulfide quantum dots (CdS QDs) and Mo-nitrogenase from Azotobacter vinelandii, composed of the iron protein (FeP) and the molybdenum-iron protein (MoFeP). Photoexcited electrons from the CdS QD are delivered by an electron transfer mediator through the FeP to the catalytic MoFeP. The complete system was optimized for the ligand on the CdS QDs, mediators, and reaction conditions. The best results were achieved with β-mercaptoethanol as a QD ligand. The mediator test revealed that 1,1'-bis(3-sulfonatopropyl)-4,4'-bipyridinium (SPr)2V (-0.4 V vs. NHE) supports the reduction of protons and N 2 to H 2 and ammonia catalyzed by nitrogenase. However, in the presence of 1,1'-trimethylene-2,2'-bipyridinium TQ (-0.54 V vs. NHE) as a mediator, nitrogenase catalysis resulted in remarkably more products. The UV-vis and in situ potentiometric studies revealed that better performance with TQ is achieved due to the significantly more negative solution potential allowing for efficient reduction of FeP. As a result, the quantum yield for conversion of absorbed photons to ammonia attains 16%, far exceeding that of previously reported nitrogenase-based systems. This work reveals the importance of tuning the electron transfer pathways in photocatalytic systems and illustrates a potent strategy for efficient electronic coupling of a photosensitizer and an N 2 reduction catalyst.

36 MATERIALS SCIENCE↗

Nanocomposite Materials for Accelerating Decarbonization

Here, decarbonization is demonstrated by catalytic conversion of CO 2 to fuel by means of exposure of cadmium selenide (CdSe) quantum dots-titania (TiO 2 ) nanophotocatalysts to sunlight illumination. The primary products resulted from this chemical reactions are methanol, carbon monoxide, and hydrogen after several hours of exposure to sun light. The overall CO 2 conversion efficiency of such quantum dot-titania nanostructures was compared with that of pure TiO 2 nanorod array photocatalyst. Data shows an improved conversion efficiency when composite quantum dot-titania nanostructures were used in comparison with titania nanophotocatalysts. It is postulated that this is due to the additional absorbance of visible light by the quantum dots and generation of additional charge separation at the CdSe-TiO 2 interfaces. The conversion efficiency of such an artificial photosynthesis process remains to be optimized for practical applications.

36 MATERIALS SCIENCE↗