Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

TESS Data Release Notes: Sector 21 DR 29

These Data Release Notes provide information on the processing and export of data from the Transiting Exoplanet Survey Satellite (TESS). The data products included in this data release are full frame images (FFIs), target pixel les, light curve les, collateral pixel les, cotrending basis vectors (CBVs), and Data Validation (DV) reports, time series, and associated xml les. These data products were generated by the TESS Science Processing Operations Center (SPOC, Jenkins et al., 2016) at NASA Ames Research Center from data collected by the TESS instrument, which is managed by the TESS Payload Operations Center (POC) at Massachusetts Institute of Technology (MIT). The format and content of these data products are documented in the Science Data Products Description Document (SDPDD)1. The SPOC science algorithms are based heavily on those of the Kepler Mission science pipeline, and are described in the Kepler Data Processing Handbook (Jenkins, 2017).2 The Data Validation algorithms are documented in Twicken et al. (2018) and Li et al. (2019). The TESS Instrument Handbook (Vanderspek et al., 2018) contains more information about the TESS instrument design, detector layout, data properties, and mission operations. The TESS Mission is funded by NASA's Science Mission Directorate.

Michael M. Fausnaugh↗

TESS Data Release Notes: Sectors 1 – 36, Multi-sector Search, DR53

These Data Release Notes provide information on the processing and export of data from the Transiting Exoplanet Survey Satellite (TESS). This data release is a combined, multi-sector transit search only. The underlying data products from individual observing sectors have been previously released. The data products included in this data release are the Data Validation (DV) reports, time series, and associated xml files for the threshold crossing events (TCEs) found by searching a combined data set including data from multiple observing sectors. These data products were generated by the TESS Science Processing Operations Center (SPOC, Jenkins et al., 2016) at NASA Ames Research Center from data collected by the TESS instrument, which is managed by the TESS Payload Operations Center (POC) at Massachusetts Institute of Technology (MIT). The format and content of these data products are documented in the Science Data Products Description Document (SDPDD)1. The SPOC science algorithms are based heavily on those of the Kepler Mission science pipeline, and are described in the Kepler Data Processing Handbook (Jenkins, 2020)2. The Data Validation algorithms are documented in Twicken et al. (2018) and Li et al. (2019). The TESS Instrument Handbook (Vanderspek et al., 2018) contains more information about the TESS instrument design, detector layout, data properties, and mission operations. The TESS Mission is funded by NASA's Science Mission Directorate.

TESS↗

NASA Sounding Rocket Program Educational Outreach

Educational and public outreach is a major focus area for the National Aeronautics and Space Administration (NASA). The NASA Sounding Rocket Program (NSRP) shares in the belief that NASA plays a unique and vital role in inspiring future generations to pursue careers in science, mathematics, and technology. To fulfill this vision, the NSRP engages in a variety of educator training workshops and student flight projects that provide unique and exciting hands-on rocketry and space flight experiences. Specifically, the Wallops Rocket Academy for Teachers and Students (WRATS) is a one-week tutorial laboratory experience for high school teachers to learn the basics of rocketry, as well as build an instrumented model rocket for launch and data processing. The teachers are thus armed with the knowledge and experience to subsequently inspire the students at their home institution. Additionally, the NSRP has partnered with the Colorado Space Grant Consortium (COSGC) to provide a "pipeline" of space flight opportunities to university students and professors. Participants begin by enrolling in the RockOn! Workshop, which guides fledgling rocketeers through the construction and functional testing of an instrumentation kit. This is then integrated into a sealed canister and flown on a sounding rocket payload, which is recovered for the students to retrieve and process their data post flight. The next step in the "pipeline" involves unique, user-defined RockSat-C experiments in a sealed canister that allow participants more independence in developing, constructing, and testing spaceflight hardware. These experiments are flown and recovered on the same payload as the RockOn! Workshop kits. Ultimately, the "pipeline" culminates in the development of an advanced, user-defined RockSat-X experiment that is flown on a payload which provides full exposure to the space environment (not in a sealed canister), and includes telemetry and attitude control capability. The RockOn! and RockSat-C elements of the "pipeline" have been successfully demonstrated by five annual flights thus far from Wallops Flight Facility. RockSat-X has successfully flown twice, also from Wallops. The NSRP utilizes launch vehicles comprised of military surplus rocket motors (Terrier-Improved Orion and Terrier-Improved Malemute) to execute these missions. The NASA Sounding Rocket Program is proud of its role in inspiring the "next generation of explorers" and is working to expand its reach to all regions of the United States and the international community as well.

Rosanova, G.↗

pathSQE : an automated workflow for single-crystal inelastic neutron scattering data processing and analysis

Inelastic neutron scattering (INS) experiments utilizing modern time-of-flight spectrometers enable the comprehensive mapping of the energy (E)- and momentum (Q)-resolved dynamical structure factor of single crystals, probing both the lattice and magnetic excitations. Yet, the large size and complexity of four-dimensional INS data are challenging current analysis workflows, often resulting in an underutilization of the measured information. To help address this issue, this paper introduces new software interfaced with the Mantid framework, pathSQE, designed to streamline the processing, analysis and interpretation of 4D single-crystal INS data. By automating key tasks such as 1D/2D slicing, symmetrization, Brillouin zone folding, data visualization, prioritization and filtering, and comparisons with simulations, pathSQE facilitates and accelerates INS data analysis workflows. Here, this paper outlines the features and implementation and provides several illustrations of the use of pathSQE on data collected on single crystals using direct-geometry time-of-flight spectrometers at the Spallation Neutron Source, including Ge, FeSi, MnO and SnS single-crystal measurements on the ARCS, HYSPEC and CNCS neutron spectrometers. Beyond streamlining post-experiment data processing, pathSQE establishes an automated and modular processing pipeline that could support future real-time experiment steering.

36 MATERIALS SCIENCE↗

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele↗

The DECADE cosmic shear project I: A new weak lensing shape catalog of 107 million galaxies

We present the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. This catalog was assembled from public DECam data including survey and standard observing programs. These data were consistently processed with the Dark Energy Survey Data Management pipeline as part of the DECADE campaign and serve as the basis of the DECam Local Volume Exploration survey (DELVE) Early Data Release 3 (EDR3). We apply the Metacalibration measurement algorithm to generate and calibrate galaxy shapes. After cuts, the resulting cosmology-ready galaxy shape catalog covers a region of $5,\!412 \,\,{\rm deg}^2$ with an effective number density of $4.59\,\, {\rm arcmin}^{-2}$. The coadd images used to derive this data have a median limiting magnitude of $r = 23.6$, $i = 23.2$, and $z = 22.6$, estimated at ${\rm S/N} = 10$ in a 2 arcsecond aperture. We present a suite of detailed studies to characterize the catalog, measure any residual systematic biases, and verify that the catalog is suitable for cosmology analyses. In parallel, we build an image simulation pipeline to characterize the remaining multiplicative shear bias in this catalog, which we measure to be $m = (-2.454 \pm 0.124) \times10^{-2}$ for the full sample. Despite the significantly inhomogeneous nature of the data set, due to it being an amalgamation of various observing programs, we find the resulting catalog has sufficient quality to yield competitive cosmological constraints.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

GEONEX: Challenges in Producing MODIS-Like Land Products from a New Generation of Geostationary Sensors

The new generation geostationary (GEO) remote sensors (GOES-R ABI, Himawari AHI, and FY4 AGRI) provide high frequency (5-15 minute) observations spatially/spectrally similar to MODIS/VIIRS for land monitoring. These new features of GEO satellite sensors make producing MODIS like land products for terrestrial monitoring possible. The NASA Earth Exchange (NEX) team developed the GEONEX pipeline that is containerized, deployable on NASA Pleiades supercomputer as well as public cloud platforms (e.g. AWS). The processing pipeline is designed to take Himawari Standard Data (HSD) and GOES-16 L1b to generate surface reflectance (SR) and other high-level land remote sensing products. In order to produce low-Earth-orbiting (LEO) remote sensing compatible land products, inter-comparison between Himawari AHI and MODIS Terra/Aqua has been conducted in this research work. Comparisons of TOA reflectance and surface reflectance between AHI and Terra/Aqua are presented. Ray-Matching method was used to locate the co-located pixels, where GEO and LEO sensors look at the land target with similar Viewing Zenith Angle (VZA) and Viewing Azimuth Angle (VAA) simultaneously. Here, we address challenges associated with the selection of qualified pixels of similar solar illumination condition and atmosphere path. We used strict criterion to constrain the pixel selection: the time difference between GEO and LEO observations is less than +-2.5 mins, the cosine of VZA difference is less than 1%, and the VAA difference is less than 10 deg. We also discuss the strong radiometric consistency that the new generation GEO sensors along with the popular LEO sensors would benefit the environmental remote sensing community.

Li, Shuang↗

Kepler Science Operations Center Architecture

We give an overview of the operational concepts and architecture of the Kepler Science Data Pipeline. Designed, developed, operated, and maintained by the Science Operations Center (SOC) at NASA Ames Research Center, the Kepler Science Data Pipeline is central element of the Kepler Ground Data System. The SOC charter is to analyze stellar photometric data from the Kepler spacecraft and report results to the Kepler Science Office for further analysis. We describe how this is accomplished via the Kepler Science Data Pipeline, including the hardware infrastructure, scientific algorithms, and operational procedures. The SOC consists of an office at Ames Research Center, software development and operations departments, and a data center that hosts the computers required to perform data analysis. We discuss the high-performance, parallel computing software modules of the Kepler Science Data Pipeline that perform transit photometry, pixel-level calibration, systematic error-correction, attitude determination, stellar target management, and instrument characterization. We explain how data processing environments are divided to support operational processing and test needs. We explain the operational timelines for data processing and the data constructs that flow into the Kepler Science Data Pipeline.

Middour, Christopher↗

WFC3 Calibration and Data Processing

Wide Field Camera 3 (WFC3), a panchromatic imager being developed for the Hubble Space Telescope (HST), is now fully integrated and over the past year has completed first rounds of extensive ground testing at Goddard Space Flight Center (GSFC), in both ambient and thermal-vacuum test environments. This report summarizes the results of those tests and describes the pipeline processing methods that will be used to calibrate WFC3 data. WFC3 is designed to ensure that the superb imaging performance of HST is maintained through the end of the mission and takes advantage of recent developments in detector technology to provide new and unique capabilities for HST. WFC3 contains ultraviolet/visible (UVIS) and near-infrared (IR) imaging channels, offering high sensitivity and wide field of view over the broadest wavelength range of any HST instrument. It is slated to replace the current Wide Field and Planetary Camera 2 during Servicing Mission 4. The WFC3 UVIS channel is based on elements from the Advanced Camera for Surveys (ACS)Wide Field Camera (WFC), with a 4096x4096 pixel Marconi CCD covering a 160x160 arcsecond field of view. The WFC3 UVIS channel is optimized for maximum sensitivity in the near-UV and contains a complement of 48 spectral filters and a grism. The WFC3 IR channel uses a 1024x1024 pixel HgCdTe Hawaii-1R detector array covering a 135x135 arcsecond field of view. The array sensitivity is optimized in the 0.8-1.7micron spectral range. The IR channel accomodates 15 filters and 2 grisms for slitless spectroscopy.

Bushouse, H.↗

Data Accountability and Uncertainty Analysis for the Mars Science Laboratory

This paper presents machine learning-based approaches to automate and optimize the detection of volume loss for the downlink process of telemetry data from the Mars Curiosity Rover. The Curiosity observes volume loss and data corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the data flow. To resolve this issue, we created a data pipeline to accumulate data from various data sources in the downlink process and detect where the data is missed. In this paper, we benchmarked different methodologies based on the accuracy and excitability of them to identify whether a downlink data that is received to the ground system is complete or incomplete. Our results show that machine learning methods can improve the performance of the GDSA by 55% while the user can diagnose why data is missed and provide an explanation for the data accountability problem.

Chowdhury, Ameera↗

Video Mosaicking for Inspection of Gas Pipelines

A vision system that includes a specially designed video camera and an image-data-processing computer is under development as a prototype of robotic systems for visual inspection of the interior surfaces of pipes and especially of gas pipelines. The system is capable of providing both forward views and mosaicked radial views that can be displayed in real time or after inspection. To avoid the complexities associated with moving parts and to provide simultaneous forward and radial views, the video camera is equipped with a wide-angle (>165 ) fish-eye lens aimed along the axis of a pipe to be inspected. Nine white-light-emitting diodes (LEDs) placed just outside the field of view of the lens (see Figure 1) provide ample diffuse illumination for a high-contrast image of the interior pipe wall. The video camera contains a 2/3-in. (1.7-cm) charge-coupled-device (CCD) photodetector array and functions according to the National Television Standards Committee (NTSC) standard. The video output of the camera is sent to an off-the-shelf video capture board (frame grabber) by use of a peripheral component interconnect (PCI) interface in the computer, which is of the 400-MHz, Pentium II (or equivalent) class. Prior video-mosaicking techniques are applicable to narrow-field-of-view (low-distortion) images of evenly illuminated, relatively flat surfaces viewed along approximately perpendicular lines by cameras that do not rotate and that move approximately parallel to the viewed surfaces. One such technique for real-time creation of mosaic images of the ocean floor involves the use of visual correspondences based on area correlation, during both the acquisition of separate images of adjacent areas and the consolidation (equivalently, integration) of the separate images into a mosaic image, in order to insure that there are no gaps in the mosaic image. The data-processing technique used for mosaicking in the present system also involves area correlation, but with several notable differences: Because the wide-angle lens introduces considerable distortion, the image data must be processed to effectively unwarp the images (see Figure 2). The computer executes special software that includes an unwarping algorithm that takes explicit account of the cylindrical pipe geometry. To reduce the processing time needed for unwarping, parameters of the geometric mapping between the circular view of a fisheye lens and pipe wall are determined in advance from calibration images and compiled into an electronic lookup table. The software incorporates the assumption that the optical axis of the camera is parallel (rather than perpendicular) to the direction of motion of the camera. The software also compensates for the decrease in illumination with distance from the ring of LEDs.

Magruder, Darby↗

The DEEP2 Galaxy Redshift Survey: Design, Observations, Data Reduction, and Redshifts

We describe the design and data analysis of the DEEP2 Galaxy Redshift Survey, the densest and largest high-precision redshift survey of galaxies at z approx. 1 completed to date. The survey was designed to conduct a comprehensive census of massive galaxies, their properties, environments, and large-scale structure down to absolute magnitude MB = −20 at z approx. 1 via approx.90 nights of observation on the Keck telescope. The survey covers an area of 2.8 Sq. deg divided into four separate fields observed to a limiting apparent magnitude of R(sub AB) = 24.1. Objects with z approx. < 0.7 are readily identifiable using BRI photometry and rejected in three of the four DEEP2 fields, allowing galaxies with z > 0.7 to be targeted approx. 2.5 times more efficiently than in a purely magnitude-limited sample. Approximately 60% of eligible targets are chosen for spectroscopy, yielding nearly 53,000 spectra and more than 38,000 reliable redshift measurements. Most of the targets that fail to yield secure redshifts are blue objects that lie beyond z approx. 1.45, where the [O ii] 3727 Ang. doublet lies in the infrared. The DEIMOS 1200 line mm(exp −1) grating used for the survey delivers high spectral resolution (R approx. 6000), accurate and secure redshifts, and unique internal kinematic information. Extensive ancillary data are available in the DEEP2 fields, particularly in the Extended Groth Strip, which has evolved into one of the richest multiwavelength regions on the sky. This paper is intended as a handbook for users of the DEEP2 Data Release 4, which includes all DEEP2 spectra and redshifts, as well as for the DEEP2 DEIMOS data reduction pipelines. Extensive details are provided on object selection, mask design, biases in target selection and redshift measurements, the spec2d two-dimensional data-reduction pipeline, the spec1d automated redshift pipeline, and the zspec visual redshift verification process, along with examples of instrumental signatures or other artifacts that in some cases remain after data reduction. Redshift errors and catastrophic failure rates are assessed through more than 2000 objects with duplicate observations. Sky subtraction is essentially photon-limited even under bright OH sky lines; we describe the strategies that permitted this, based on high image stability, accurate wavelength solutions, and powerful B-spline modeling methods. We also investigate the impact of targets that appear to be single objects in ground-based targeting imaging but prove to be composite in Hubble Space Telescope data; they constitute several percent of targets at z approx. 1, approaching approx. 5%-10% at z > 1.5. Summary data are given that demonstrate the superiority of DEEP2 over other deep high-precision redshift surveys at z approx. 1 in terms of redshift accuracy, sample number density, and amount of spectral information. We also provide an overview of the scientific highlights of the DEEP2 survey thus far.

Galaxy↗

GEONEX Data Products: Geostationary Satellite Derived Land Surface and Atmospheric Public Data Products

The latest generation of geostationary satellites carry sensors such as the Advanced Baseline Imager (GOES-16/17) and the Advanced Himawari Imager (Himawari-8/9) that closely mimic the spatial and spectral characteristics of MODIS and VIIRS, useful for monitoring land surface conditions. The NASA Earth Exchange (NEX) team at Ames Research Center has embarked on a collaborative effort among scientists from NASA and NOAA exploring the feasibility of producing operational land surface products similar to those from MODIS/VIIRS. The team built a processing pipeline called GEONEX that is capable of converting raw geostationary data into routine products of Fires, surface reflectances, vegetation indices, LAI/FPAR, ET and GPP/NPP using algorithms adapted from both NASA/EOS and NOAA/GOES-R programs. The GEONEX pipeline will begin to produce provisional data products to be consumed by external collaborators and the academic community. In order to better inform and introduce the GEONEX products to the science community, the provisional products shall be distributed from the NAS data portal, located at data.nas.nasa.gov, and simple webpage at www.nasa.gov/geonex has been deployed, which describes any algorithms used in deriving the products, user manuals and data file information. We will also update the status of the data processing, on the website and provide links to the latest datasets, and use a geonex mailing list, using lists.nasa.gov.

Wang, Weile↗

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri↗

The Mars 2020 Ground Data System Architecture

The Mars 2020 Mission’s primary objective is to collect 20 geographically unique samples during its prime mission of one and a quarter Martian years, or just over 2 Earth years. Mission planners determined the project needed to develop a system that would enable the operations team to analyze engineering and science data, make science decisions, select viable rover targets at a millimeter resolution and validate an uplink bundle for a car sized rover with more complex science instruments than any previous Mars surface mission. All this had to be done within a five hour time frame. Doing this with a small team would be a challenge, but this had to be accomplished by a large team of engineers and scientists located across North America and Europe. Achieving this level of operational efficiency was unheard of in the prime mission. In addition, the mission had another set of requirements that had nothing to do with surface operations; the Mars 2020 Ground Data System (GDS) was also expected to comply with a new set of security requirements to keep up with the ever changing cybersecurity landscape. The Mars 2020 Ground Data System (GDS) is a re-architected version of the Mars Science Laboratory GDS. The primary goal was to integrate the lessons learned from previous Mars surface missions, accommodate a set of new requirements and capabilities required to ensure mission success, and comply with a new set of cybersecurity controls. The new architecture includes several unique qualities including a data lake, language-agnostic system-wide event-based operations, containerization, automated deployment, network segmentation, infrastructure-as-code, API-driven interfaces, and the first Mars surface GDS to operate primarily in the cloud. The new architecture enabled greater access to the system’s data, tighter integration with the operations team, and a higher level of traceability. The availability of the data also enabled a new set of capabilities previously not possible on surface missions. These new capabilities include an autonomous data to information, pipeline for downlink analysis, horizontal scaling of science data processing capabilities, autonomous round trip data tracking of science and engineering data, integration of flight system state into the tactical planning cycle, high fidelity targeting utilizing kinematic data, and hierarchical image and 3d meshes data representations. This paper will introduce the requirements for the Mars 2020 Mission, the heritage architecture, and the rationale for the changes to achieve the new architecture. The paper will continue to describe the fundamental changes made to the GDS architecture, how these changes enabled a more tightly integrated GDS, and the new capabilities that were enabled by the new architecture. The paper will conclude with the lessons learned from the process of rearchitecting a heritage GDS system and from the first 200 days of operations supporting over 800 users from around the world.

Lopez-Roig, Reynaldo↗

Machine Learning for Predicting Team Functioning in HERA Missions

Team functioning is integral to success in future long term space exploration missions. Proactively detecting declines in team functioning can mitigate conflict and ensure mission success. This project developed a speech-based artificial intelligence (AI) system that unobtrusively predicts degradation in team functioning, including performance and cohesion, in the Human Exploration Research Analog (HERA) Campaigns 4 and 5. The AI system conducted automated analysis of the prosodic (tone of voice) and linguistic (language content) components of speech, modeling interpersonal dynamics at both the turn-taking and day-wide levels. We investigated team functioning via observing structured interactions (i.e., multi-mission space exploration vehicle-extra vehicular activity [MMSEV-EVA], team interaction battery [TIB]) and unstructured interactions before the MMSEV-EVA task. We developed machine learning models to predict team functioning (objective task accuracy, self reported team efficacy and self reported team cohesion) by analyzing OpenSmile acoustic features, linguistic descriptors extracted via the linguistic inquiry and word count (LIWC) dictionary, and semantic embeddings. In the TIB, static models using logistic regression and random forests were not able to predict task accuracy, but predicted team efficacy and cohesion during both the decision making and relational tasks to a moderate level (60-70%). Majority voting on the individual turns to predict day long team efficacy further increased accuracies (70-80%). Finally, long short-term memory (LSTM) models showed the best performance across all variables (80-91%), including task performance. In the MMSEV-EVA, static models achieved an accuracy of 60% with majority voting, which increased to 80% through the incorporation of mission day as a variable, accounting for the learning effect. A key finding across both tasks was the "team-dependent" nature of these interactions; models achieved much higher accuracy when trained on prior days of the same team's data rather than attempting to generalize across entirely different teams, with even 1-2 days of prior data per team achieving 5-15% improvement over team-independent models. In addition, the incorporation of pre-task data from the same team also improves model performance, e.g., incorporating data from the decision-making task of the TIB, which preceded the relational task, improved the prediction of team efficacy and cohesion during the latter. We compared model performance when trained on machine-generated data compared to data that had been further corrected by human annotators. Overall, models trained on human-corrected data exhibited a modest improvement in performance, particularly when acoustic features were used. We found no significant correlation between word error rate (WER) and model accuracy (r(55) = -0.08, p = 0.51), but model’s accuracy was significantly higher for medium/high quality transcription (0.74 (SD = 0.48)) compared to the low-quality group (0.64 (SD = 0.36)) (t(63)=2.82, p = 0.006). Based on these, several design recommendation emerge, that could inform Standards at NASA. Models predicting team functioning should incorporate at least one to two days of historical interaction data, include brief pre-task discussions, and explicitly model temporal learning effects, especially for longer operational tasks. Minimum quality standards for automated speech-processing pipelines are needed, given the performance gains observed with manually corrected acoustic data. Finally, systems should leverage both acoustic features and language embeddings in complementary ways, with modality choices and fusion strategies tailored to mission context, task demands, and data quality requirements.

Shrivatsa Mishra↗