Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Enabling the Direct Detection of Earth-Sized Exoplanets with the LBTI HOSTS Project: A Progress Report

NASA has funded a project called the Hunt for Observable Signatures of Terrestrial Systems (HOSTS) to survey nearby solar type stars to determine the amount of warm zodiacal dust in their habitable zones. The goal is not only to determine the luminosity distribution function but also to know which individual stars have the least amount of zodiacal dust. It is important to have this information for future missions that directly image exoplanets as this dust is the main source of astrophysical noise for them. The HOSTS project utilizes the Large Binocular Telescope Interferometer (LBTI), which consists of two 8.4-m apertures separated by a 14.4-m baseline on Mt. Graham, Arizona. The LBTI operates in a nulling mode in the mid-infrared spectral window (8-13 micrometers), in which light from the two telescopes is coherently combined with a 180 degree phase shift between them, producing a dark fringe at the location of the target star. In doing so the starlight is greatly reduced, increasing the contrast, analogous to a coronagraph operating at shorter wavelengths. The LBTI is a unique instrument, having only three warm reflections before the starlight reaches cold mirrors, giving it the best photometric sensitivity of any interferometer operating in the mid-infrared. It also has a superb Adaptive Optics (AO) system giving it Strehl ratios greater than 98% at 10 micrometers. In 2014 into early 2015 LBTI was undergoing commissioning. The HOSTS. project team passed its Operational Readiness Review (ORR) in April 2015. The team recently published papers on the target sample, modeling of the nulled disk images, and initial results such as the detection of warm dust around eta Corvi. Recently a paper was published on the data pipeline and on-sky performance. An additional paper is in preparation on Beta Leo. We will discuss the scientific and programmatic context for the LBTI project, and we will report recent progress, new results, and plans for the science verification phase that started in February 2016, and for the survey.

Hunt for Observable Signatures of Terrestrial Syst↗

The Gaia Catalogue Second Data Release and Its Implications to Optical Observations of Man-Made Earth Orbiting Objects

The Gaia catalogue second data release and its implications to optical observations of man-made Earth orbiting objects. Abstract and not the Final Paper is attached. The Gaia spacecraft was launched in December 2013 by the European Space Agency to produce a three-dimensional, dynamic map of objects within the Milky Way. Gaia's first year of data was released in September 2016. Common sources from the first data release have been combined with the Tycho-2 catalogue to provide a 5 parameter astrometric solution for approximately 2 million stars. The second Gaia data release is scheduled to come out in April 2018 and is expected to provide astrometry and photometry for more than 1 billion stars, a subset of which with a the full 6 parameter astrometric solution (adding radial velocity) and positional accuracy better than 0.002 arcsec (2 mas). In addition to precise astrometry, a unique opportunity exists with the Gaia catalogue in its production of accurate, broadband photometry using the Gaia G filter. In the past, clear filters have been used by various groups to maximize likelihood of detection of dim man-made objects but these data were very difficult to calibrate. With the second release of the Gaia catalogue, a ground based system utilizing the G band filter will have access to 1.5 billion all-sky calibration sources down to an accuracy of 0.02 magnitudes or better. In this talk, we will discuss the advantages and practicalities of implementing the Gaia filters and catalogue into data pipelines designed for optical observations of man-made objects.

Frith, James M.↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Summary of Carbon Dioxide Pipeline Systems and Incident Data in North America

Pipelines are historically seen as the primary transportation mode for carbon dioxide (CO 2 ) streams in the context of carbon capture and storage (CCS) and oil and gas industries. Pipeline transmission of CO 2 over longer distances is regarded as most efficient and economical when the CO 2 is in the dense phase, i.e., in liquid or supercritical regime, due to transporting CO 2 in dense phase that allows for a smaller-diameter pipeline to move a given flow, which optimizes project cost.

42 ENGINEERING↗

Hydropower Capacity Factor Trends & Analytics for the United States

This data repository contains all code, input data, and data generated for Turner et al. (2024)—“Hydropower capacity factors trending down in the United States”. File descriptions: – hydro-cf-trends-inputs.zip: Full set of input data used in this study, organized for direct entry into “/data” directory of hydro-cf-trends data processing pipeline. – hydro-cf-trends.zip: Full data processing pipeline, coded using the R {targets} framework. This is a snapshot release (v1.0) of the code repository stored at https://code.ornl.gov/turnersw/hydro-cf-trends/. – hydro-cf-trends-results.zip: Provides all dam level results required to reproduce results and graphics in Turner et al. (2024). Dams are identified by the “complxID” (root of the hydropower plant ID in the Existing Hydropower Assets Database, inherited from HILARRI). Results include: • dam_CF_trends.csv: Table of long-term trends in annualized capacity factors for 610 dams and modeled annualized capacity factors for 362 modeled dams (naturalized and assimilated flows). • dam_annualized_CF_gen.csv: Annualized time series of the following variables for each of 610 hydropower dams with nameplate > 5MW – Reported nameplate capacity (MW) – Implied maximum annual generation (MWh) – Reported net generation (MWh) – Computed annual capacity factor – Modeled annual capacity factor (362 modeled plants only)

13 HYDRO ENERGY↗

Using XML and Java Technologies for Astronomical Instrument Control

Traditionally, instrument command and control systems have been highly specialized, consisting mostly of custom code that is difficult to develop, maintain, and extend. Such solutions are initially very costly and are inflexible to subsequent engineering change requests, increasing software maintenance costs. Instrument description is too tightly coupled with details of implementation. NASA Goddard Space Flight Center, under the Instrument Remote Control (IRC) project, is developing a general and highly extensible framework that applies to any kind of instrument that can be controlled by a computer. The software architecture combines the platform independent processing capabilities of Java with the power of the Extensible Markup Language (XML), a human readable and machine understandable way to describe structured data. A key aspect of the object-oriented architecture is that the software is driven by an instrument description, written using the Instrument Markup Language (IML), a dialect of XML. IML is used to describe the command sets and command formats of the instrument, communication mechanisms, format of the data coming from the instrument, and characteristics of the graphical user interface to control and monitor the instrument. The IRC framework allows the users to define a data analysis pipeline which converts data coming out of the instrument. The data can be used in visualizations in order for the user to assess the data in real-time, if necessary. The data analysis pipeline algorithms can be supplied by the user in a variety of forms or programming languages. Although the current integration effort is targeted for the High-resolution Airborne Wideband Camera (HAWC) and the Submillimeter and Far Infrared Experiment (SAFIRE), first-light instruments of the Stratospheric Observatory for Infrared Astronomy (SOFIA), the framework is designed to be generic and extensible so that it can be applied to any instrument. Plans are underway to test the framework with other types of instruments, such as remote sensing earth science instruments.

Ames, Troy↗

GeneLab Analysis Working Group Pipelines

GeneLab must establish data processing pipelines for common data types including microarray, RNA-sequencing, and metagenomic profiling. Here we give an overview of current microarray and RNA-seq pipelines and discuss future pipelines including metagenomic profiling pipelines

Galazka, Jonathan M.↗

Summary on the XRT Data Products and Pipeline Production

The new code to generate light curves for GRB observed with the XRT on Swift has been distributed with the latest software release. The code has been extensively tested compared with similar output from Penn State. Adjustments in both codes were made to best represent the latest calibration information and the latest software algorithms. In this talk will be highlighted the steps taken to produce the GRB light-curves, how it is been tested for non GRB sources and last the future adjustments to be made to include a better correction for the bad columns. In addition it will be presented as a new routine developed to automatically collect spectra at different phases of the GRB decay. This new routine will be included in the next software release. The XRT pipeline at GSFC will implement the production of the data files containing the level 3 products. The talk will discuss the current status and the additional data files planned.

Angelini, Lorella↗

Description of the TCERT Vetting Reports for Data Release 25

This document, the Kepler Instrument Handbook (KIH), is for Kepler and K2 observers, which includes the Kepler Science Team, Guest Observers (GOs), and astronomers doing archival research on Kepler and K2 data in NASAs Astrophysics Data Analysis Program (ADAP). The KIH provides information about the design, performance, and operational constraints of the Kepler flight hardware and software, and an overview of the pixel data sets available. The KIH is meant to be read with these companion documents:1. Kepler Data Processing Handbook (KSCI-19081) or KDPH (Jenkins et al., 2016). The KDPH describes how pixels downlinked from the spacecraft are converted by the Kepler Data Processing Pipeline (henceforth just the pipeline) into the data products delivered to the MAST archive. 2. Kepler Archive Manual (KDMC-10008) or KAM (Thompson et al., 2016). The KAM describes the format and content of the data products, and how to search for them.3. Kepler Data Characteristics Handbook (KSCI-19040) or KDCH (Christiansen et al., 2016). The KDCH describes recurring non-astrophysical features of the Kepler data due to instrument signatures, spacecraft events, or solar activity, and explains how these characteristics are handled by the pipeline.4. Kepler Data Release Notes 25 (KSCI-19065) or DRN 25 (Thompson et al., 2015). DRN 25 describes signatures and events peculiar to individual quarters, and the pipeline software changes between a data release and the one preceding it.Together, these documents supply the information necessary for obtaining and understanding Kepler results, given the real properties of the hardware and the data analysis methods used, and for an independent evaluation of the methods used if so desired.

Instrument↗

A path to intelligent watersheds: coordinating the data to decision pipeline

Operations of multi-reservoir systems are challenged in-part by the interplay of complex physical processes functioning within the watershed. The employment of intelligent systems can be of aid by linking environmental sensing, information technology, data analytics, simulation and decision support to achieve a data-to-decision flow of information. A further challenge is that watershed resources are managed for multiple purposes requiring some level of coordination among numerous resource managers, asset operators and users. System intelligence in this context relies on shared community platforms (data portals, community models), and coordinated communication between decision makers. Opportunities to enrich watershed intelligence has been the subject of a roadmapping exercise for the Department of Energy’s Water Power Technologies Office which has relied on broad stakeholder engagement. Initial phases of engagement involved personal interviews and a series of virtual group meetings, which focused on identifying opportunities to improve the intelligence of the physical infrastructure within our watersheds—examples of feedback include improved sensing of snowpack and runoff, data standards for facilitated data sharing, and better forecasting tools. The latter phase of engagement involved the conduct of a case study in the Upper Colorado River basin where key stakeholders were interviewed to map how their decisions are informed by intelligence from other basin stakeholders. Our presentation will highlight the interdisciplinary flow of information in complex watershed systems and identify physical and institutional opportunities toward the strategic operation of water infrastructure.

Colorado River↗

An Efficient GPU-Accelerated Multi-Source Global Fit Pipeline for LISA Data Analysis

The large-scale analysis task of deciphering gravitational wave signals in the LISA data stream will be difficult, requiring a large amount of computational resources and extensive development of computational methods. Its high dimensionality, multiple model types, and complicated noise profile require a global fit to all parameters and input models simultaneously. In this work, we detail our global fit algorithm, called “Erebor,” designed to accomplish this challenging task. It is capable of analysing current state-of-the-art datasets and then growing into the future as more pieces of the pipeline are completed and added. We describe our pipeline strategy, the algorithmic setup, and the results from our analysis of the LDC2A Sangria dataset, which contains Massive Black Hole Binaries, compact Galactic Binaries, and a parameterized noise spectrum whose parameters are unknown to the user. The Erebor algorithm includes three unique and very useful contributions: GPU acceleration for enhanced computational efficiency; ensemble MCMC sampling with multiple MCMC walkers per temperature for better mixing and parallelized sample creation; and special online updates to reversible-jump (or trans-dimensional) sampling distributions to ensure sampler mixing and accurate initial estimates for detectable sources in the data. We recover posterior distributions for all 15 (6) of the injected MBHBs in the LDC2A training (hidden) dataset. We catalog ∼12000 Galactic Binaries (∼8000 as high confidence detections) for both the training and hidden datasets. All of the sources and their posterior distributions are provided in publicly available catalogs.

LISA global fit↗

Efficient GPU-Accelerated MultiSource Global Fit Pipeline for LISA Data Analysis

The large-scale analysis task of deciphering gravitational-wave signals in the LISA data stream will be difficult, requiring a large amount of computational resources and extensive development of computational methods. Its high dimensionality, multiple model types, and complicated noise profile require a global fit to all parameters and input models simultaneously. In this work, we detail our global fit algorithm, called “Erebor,” designed to accomplish this challenging task. It is capable of analyzing current state-of-the-art datasets and then growing into the future as more pieces of the pipeline are completed and added. We describe our pipeline strategy, the algorithmic setup, and the results from our analysis of the LDC2A Sangria dataset, which contains massive black hole binaries, compact galactic binaries, and a parametrized noise spectrum whose parameters are unknown to the user. The Erebor algorithm includes three unique and very useful contributions: GPU acceleration for enhanced computational efficiency; ensemble Markov Chain Monte Carlo (MCMC) sampling with multiple MCMC walkers per temperature for better mixing and parallelized sample creation; and special online updates to reversible-jump (or transdimensional) sampling distributions to ensure sampler mixing and accurate initial estimates for detectable sources in the data.We recover posterior distributions for all 15 (6) of the injected massive black hole binaries (MBHB) in the LDC2A training (hidden) dataset. We catalog ∼12000 galactic binaries (∼8000 as high confidence detections) for both the training and hidden datasets. All of the sources and their posterior distributions are provided in publicly available catalogs.

LISA↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗

Artificial Intelligence Transforming Post-Translational Modification Research

Post-Translational Modifications (PTMs) are covalent changes to amino acids that occur after protein synthesis, including covalent modifications on side chains and peptide backbones. Many PTMs profoundly impact cellular and molecular functions and structures, and their significance extends to evolutionary studies as well. In light of these implications, we have explored how artificial intelligence (AI) can be utilized in researching PTMs. Initially, rationales for adopting AI and its advantages in understanding the functions of PTMs are discussed. Then, various deep learning architectures and programs, including recent applications of language models, for predicting PTM sites on proteins and the regulatory functions of these PTMs are compared. Finally, our high-throughput PTM-data-generation pipeline, which formats data suitably for AI training and predictions is described. We hope this review illuminates areas where future AI models on PTMs can be improved, thereby contributing to the field of PTM bioengineering.

59 BASIC BIOLOGICAL SCIENCES↗

Kepler: A Search for Terrestrial Planets. K2 Handbook

The Kepler spacecraft was repurposed for the K2 mission a year after the failure of the second of Kepler's four reaction wheels in 2013 May. The purpose of this document, the K2 Handbook (K2H), is to describe features of K2 operations, performance, data analysis, and archive products which are common to most K2 campaigns, but different in degree or kind from the corresponding features of the Kepler mission.The K2 Handbook is meant to be read with the following companion documents, which are all publicly available:1. Kepler Instrument Handbook (KSCI-19033) provides information about the design, performance and operational constraints of the instrument and an overview of the types of pixel data that are available.2. Kepler Data Processing Handbook (KSCI-19081) describes how pixels downloaded from the spacecraft are converted by the Kepler Data Processing Pipeline into the data products available at the MAST archive3. Kepler Archive Manual (KDMC-100008) describes the format and content of the data products and how to search for them.4. Kepler Data Characteristics Handbook (KSCI-19040) describes recurring non-astrophysical features of the Kepler data due to instrument signatures, spacecraft events or solar activity and explains how these characteristics are handled by the Kepler pipeline.5. The Ecliptic Plane Input Catalog describes the provenance of the positions and Kepler magnitudes used for target management and aperature photometry.6. K2 Data Release Notes (DRN) are on-line documents available on the K2 science website which describe the data inventory, instrumental signatures and events peculiar to individual observing campaigns.

K2↗