Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scientific workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Large-Scale Visualization of 3D Unstructured Groundwater Model Using Cave Automated Virtual Environment

The immersive three-dimensional (3D) virtual reality (VR) visualization of groundwater models allows us to deepen our understanding of aquifer systems and provide better solutions to present groundwater-related problems, such as groundwater recharge, water quality, and sustainability. Visualization assists in accurately developing groundwater models and revealing important subsurface features, including faulting, folding, and unconformity. However, assessing model accuracy poses challenges due to the complexity of geology and groundwater systems. This research demonstrates a workflow to visualize and analyze raw 3D unstructured groundwater model data using an immersive Cave Automated Virtual Environment (CAVE). To visualize the unstructured groundwater model data, the raw dataset is converted into interactive CAVE-compatible formats utilizing a set of tools: ParaView, Blender, and Unity. This enables researchers to immerse themselves in the data, identifying influential patterns and relationships. e resulting insights can inform the development of sophisticated machine-learning models for groundwater level prediction. The CAVE’s immersive capabilities allow intuitive exploration from various perspectives, providing a more holistic understanding of the factors affecting groundwater levels. These insights are crucial to improve predictive models. The CAVE results also facilitate collaborative analysis and have potential applications in training and education. is research demonstrates the value of immersive VR tools such as the CAVE for unraveling intricacies within high-dimensional scientific data to drive real-world forecasting and modeling applications.

54 ENVIRONMENTAL SCIENCES↗

Integration of the Biot–Gassmann Fluid Substitution Method and Machine Learning-Based Velocity–Stress Relationship for Estimating In Situ Stresses

Recent advancements have shown that in situ stresses can be reliably estimated through an integrated machine/deep learning (ML/DL)-based framework, which relies on models trained and validated using true triaxial ultrasonic velocity (TUV) experimental data that involve measurements of ultrasonic velocity in saturated rocks under varying stress configurations. However, when the goal is to interpret lower frequency measurements, it may be more appropriate to run experiments on dry rocks and then obtain Biot–Gassmann-derived equivalent saturated velocities (low-frequency approximation) and employ these quantities for training ML/DL models to predict in situ stress. Whether the dispersion effect of frequency on the velocity–stress relationship substantially impacts in situ stress prediction is an important and unresolved question. This work presents an enhancement of ML/DL-based workflow by training and implementing ML/DL models using equivalent saturated acoustic velocities (low-frequency) obtained by applying Biot–Gassmann fluid substitution on the ultrasonic velocities of dry cores. The models were trained on TUV data sets derived from three subsurface cores extracted from the geothermal well 16B(78)-32 at the Utah FORGE site. Each core was subjected to 75 unique stress configurations for velocity measurement in the dry state. The ML/DL trained on the TUV data set with equivalent saturated velocities demonstrated promising performance to predict in situ stress in subsurface geological rocks using velocity–stress relationships with R 2 of 0.86, 0.971, and 0.975 and root mean squared error (RMSE) of 2.59, 1.92, and 1.80 for validation/testing phases of vertical, minimum horizontal, and maximum horizontal stress models, respectively. Additionally, interpretation and explanation by Shapley additive explanations (SHAP) analysis further improved scientific validation and model reliability for estimating in situ stresses.

colloids↗

Hadoop for High-Performance Climate Analytics: Use Cases and Lessons Learned

Scientific data services are a critical aspect of the NASA Center for Climate Simulations mission (NCCS). Hadoop, via MapReduce, provides an approach to high-performance analytics that is proving to be useful to data intensive problems in climate research. It offers an analysis paradigm that uses clusters of computers and combines distributed storage of large data sets with parallel computation. The NCCS is particularly interested in the potential of Hadoop to speed up basic operations common to a wide range of analyses. In order to evaluate this potential, we prototyped a series of canonical MapReduce operations over a test suite of observational and climate simulation datasets. The initial focus was on averaging operations over arbitrary spatial and temporal extents within Modern Era Retrospective- Analysis for Research and Applications (MERRA) data. After preliminary results suggested that this approach improves efficiencies within data intensive analytic workflows, we invested in building a cyber infrastructure resource for developing a new generation of climate data analysis capabilities using Hadoop. This resource is focused on reducing the time spent in the preparation of reanalysis data used in data-model inter-comparison, a long sought goal of the climate community. This paper summarizes the related use cases and lessons learned.

analytics↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

Object Proxy Patterns for Accelerating Distributed Applications

Workflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. As a result, we evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage.

Distributed Computing↗

An Integrated Data Analytics Platform

An Integrated Science Data Analytics Platform is an environment that enables the confluence of resources for scientific investigation. It harmonizes data, tools and computational resources which subsequently enable the research community to focus on the investigation rather than spending time on security, data preparation, management, etc. OceanWorks is a NASA technology integration project to establish a cloud-based Integrated Ocean Science Data Analytics Platform at NASA’s Physical Oceanography Distributed Active Archive Center (PO.DAAC) for big ocean science. It focuses on advancement and maturity by bringing together several NASA open-source, big data projects for parallel analytics, anomaly detection, in-situ to satellite data matchup, quality-screened data subsetting, search relevancy, and data discovery. Our communities are relying on data distributed through data centers such as the PO.DAAC, COAPS, NCAR, and many others to conduct their research. In typical investigations, scientists would engage in: search for data, evaluate the relevance of that data, download it, and then apply algorithms to identify trends. Such workflow cannot scale if the research involves a massive amount of data or multi-variate measurements. NASA’s Surface Water and Ocean Topography (SWOT) mission is expected to produce massive amount of observational data during its 3-year nominal mission. Collections like SWOT challenges all existing Earth Science data archival, distribution and analysis paradigms. In this paper, we will discuss how OceanWorks enhances the analysis of physical ocean data where the computation is done on an elastic cloud platform next to the archive to deliver fast, web-accessible services for working with oceanographic measurements.

Yang, Chaowei↗

The Scientific Importance of Returning Airfall Dust as a Part of Mars Sample Return (MSR)

Dust transported in the martian atmosphere is of intrinsic scientific interest and has relevance for the planning of human missions in the future. The MSR Campaign, as currently designed, presents an important opportunity to return serendipitous, airfall dust. The tubes containing samples collected by the Perseverance rover would be placed in cache depots on the martian surface perhaps as early as 2023–24 for recovery by a subsequent mission no earlier than 2028–29, and possibly as late as 2030–31. Thus, the sample tube surfaces could passively collect dust for multiple years. This dust is deemed to be exceptionally valuable as it would inform our knowledge and understanding of Mars' global mineralogy, surface processes, surface-atmosphere interactions, and atmospheric circulation. Preliminary calculations suggest that the total mass of such dust on a full set of tubes could be as much as 100 mg and, therefore, sufficient for many types of laboratory analyses. Two planning steps would optimize our ability to take advantage of this opportunity: (1) the dust-covered sample tubes should be loaded into the Orbiting Sample container (OS) with minimal cleaning and (2) the capability to recover this dust early in the workflow within an MSR Sample Receiving Facility (SRF) would need to be established. A further opportunity to advance dust/atmospheric science using MSR, depending upon the design of the MSR Campaign elements, may lie with direct sampling and the return of airborne dust.

Monica M. Grady↗

Integrated fluorescence light microscopy-guided cryo-focused ion beam-milling for in situ montage cryo-ET

Cryogenic-electron tomography (cryo-ET) permits the in situ visualization of biological macromolecules at the molecular level. Owing to the variable thickness of cells, tissues and organisms, frozen specimens may need to be thinned by cryo-focused ion beam (FIB) milling to produce thin (<500 nm) cryo-lamellae suitable for cryo-ET. Locating regions of interest remains a challenge because untargeted milling can lead to inadvertent ablation and removal of regions of interest. Correlative light and electron microscopy, combined with cryo-FIB milling, can guide the identification of labeled targets in the cellular milieu. Multiple transfers between cryo-imaging instruments, cumbersome correlation algorithms, limited accuracy and low throughput have hindered the routine adoption of cryo-FIB milling within a multimodal correlative workflow for in situ structural biology. Here, in this study, we present a workflow for 3D correlative cryo-fluorescence light microscopy-FIB-ET that streamlines fluorescence light microscopy-guided FIB milling, improving throughput while preserving both structural and contextual information. The complete integration of hardware and software described here minimizes sample contamination from cross-platform exchanges and greatly enhances the efficiency of 3D targeting in cryo-milling. We then describe procedures for implementing montage parallel array cryo-ET (MPACT), which can be easily adapted to any modern life-science transmission electron microscope. MPACT supports high-throughput cryo-ET acquisitions (10 tilt series in 1.5 h) for structure determination and comprehensive contextual understanding of macromolecules within their native surroundings. A complete session from sample preparation to MPACT data processing takes 5−7 d for an individual experienced in both cryo-EM and cryo-FIB milling.

Yang, Jie E. [Univ. of Wisconsin, Madison, WI (Uni↗

Evolution of the Scope and Capabilities of Uplink Support Software for Mars Surface Operations

In January of 2004 both of the Mars Exploration Rover spacecraft landed safely, initiating daily surface operations at the Jet Propulsion Laboratory for what was anticipated to be approximately three months of mobile exploration. The longevity of this mission, still ongoing after ten years, has provided not only a tremendous return of scientific data but also the opportunity to refine and improve the methodology by which robotic Mars surface missions are commanded. Since the landing of the Mars Science Laboratory spacecraft in August of 2012, this methodology has been successfully applied to operate a Martian rover which is both similar to, and quite different from, its predecessors. For MER and MSL, daily uplink operations can be most broadly viewed as converting the combined interests of both the science and engineering teams into a spacecraft-safe set of transmittable command files. In order to accomplish these ends a discrete set of mission-critical software tools were developed which not only allowed for conformation to established JPL standards and practices but also enabled innovative technologies specific to each mission. Although these primary programs provided the requisite capabilities for meeting the high-level goals of each distinct phase of the uplink process, there was little in the way of secondary software to support the smooth flow of data from one phase to the next. In order to address this shortcoming a suite of small software tools was developed to aid in phase transitions, as well as to automate some of the more laborious and error-prone aspects of uplink operations. This paper describes the evolution of this software suite, from its initial attempts to merely shorten the duration of the operator's shift, to its current role as an indispensable tool enforcing workflow of the uplink operations process and agilely responding to the new and unexpected challenges of missions which can, and have, lasted many years longer than originally anticipated.

CoUGAR↗

A bi-channel aided stitching of atomic force microscopy images

Microscopy is an essential tool in scientific research, enabling the visualization of structures at micro- and nanoscale resolutions. However, the field of microscopy often encounters limitations in field-of-view (FOV), restricting the amount of sample that can be imaged in a single capture. To overcome this limitation, image stitching techniques have been developed to seamlessly merge multiple overlapping images into a single, high-resolution composite. The images collected from microscope need to be optimally stitched before accurate physical information can be extracted from post analysis. However, the existing stitching tools either struggle to stitch images together when the microscopy images are feature sparse or cannot address all the transformations of images when performing image stitching. To address these issues, we propose a bi-channel aided feature-based image stitching method and demonstrate its use on Atomic Force Microscopy (AFM) generated Pantoea sp. YR343 biofilm and PTO thin film sample images as experimental data. The topographical channel image of AFM data captures the morphological details of the sample, and a stitched topographical image is desired for researchers. We utilize the amplitude and phase channels of AFM data to maximize the matching features and to estimate the position of the original topographical images and show that the proposed bi-channel aided stitching method outperforms the traditional direct stitching approach in AFM topographical image stitching task. Here, we demonstrated the application on AFM, but similar approaches could be employed of optical microscopy with brightfield and fluorescence channels. We believe this proposed workflow can serve as a valuable augmentation strategy for microscopy image stitching tasks and will benefit the experimentalist to avoid erroneous analysis and discovery due to incorrect stitching.

Atomic force microscopy↗

Evolution of storage monitoring – update in response to commercial and regulatory drivers

Carbon Capture and Storage (CCS) is in transition from first-of-a kind projects and research-orientated pilots to commercially-motivated applications. Monitoring results from many newly developed and planned large scale commercial projects are limited; however, it is worthwhile to assess their evolution and consider new strategies as part of an effort to assess and document best practices. Commercial monitoring is targeted to activities that comply with regulatory drivers and de-risk investments. Commercial monitoring also supports accounting that storage has occurred and is tied to project financing. It deals with long time frames and large volumes injected into multiple wells and multiple projects in favorable areas. We see developing trends toward reproducible workflows that systematically reduce risks and clarify expectations for oversight and long-term surveillance. Monitoring techniques showing increasing trends include injection zone pressure as a history-matching and compliance tool. To reduce cost and environmental impact of time-lapse seismic data collection, deploying new approaches and tools, such as use of fibre and installed sources are increasingly applied. Concern over the risk of induced seismicity by regulatory bodies and the general public has increased, which has also resulted in increased monitoring. Some techniques used in the early research phases have been sidelined or used only in restricted applications. For example, geochemical analyses in the injection zone as well as the environment are now being deployed less than it was in research-oriented programs, except in the US where it is required by the permitting process. Expectations of frequent area-wide near surface monitoring have also decreased.

25 ENERGY STORAGE↗

Tools for unbinned unfolding

Machine learning has enabled differential cross section measurements that are not discretized. Going beyond the traditional histogram-based paradigm, these unbinned unfolding methods are rapidly being integrated into experimental workflows. Here, in order to enable widespread adaptation and standardization, we develop methods, benchmarks, and software for unbinned unfolding. For methodology, we demonstrate the utility of boosted decision trees for unfolding with a relatively small number of high-level features. This complements state-of-the-art deep learning models capable of unfolding the full phase space. To benchmark unbinned unfolding methods, we develop an extension of existing dataset to include acceptance effects, a necessary challenge for real measurements. Additionally, we directly compare binned and unbinned methods using discretized inputs for the latter in order to control for the binning itself. Lastly, we have assembled two software packages for the OmniFold unbinned unfolding method that should serve as the starting point for any future analyses using this technique. One package is based on the widely-used RooUnfold framework and the other is a standalone package available through the Python Package Index (PyPI).

47 OTHER INSTRUMENTATION↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

WETO Software Stack Best Practices

Wind energy researchers typically share one key characteristic: a passion for increasing wind energy in the global energy mix. The U.S. Department of Energy (DOE) supports this mission in a number of ways including allocating funding directly to various aspects of wind energy research through the Office of Energy Efficiency and Renewable Energy (EERE) via the Wind Energy Technologies Office (WETO). While the traditional output of research is academic publication, software development efforts are increasingly a major focus. Software tools in the research environment allow researchers to describe an idea and quickly increase the scope and scale as they study it further. As a product of research, these tools represent a direct pipeline from researcher to industry practitioners since they are the implementation of ideas described in academic publications. Given this vital role in wind energy research and commercial development, the broad research software portfolio supported by WETO must maintain a minimum level of quality to support the wind energy field in the growing transition to renewable energy. This report outlines a series o f best practices to be adopted by all WETO-supported software projects, as well as expectations that the communities interacting with these projects should have of the developers and tools themselves. Wind energy research software has a unique standing in the field of scientific software. The stakeholders are varied with a subset being: (1) DOE EERE leadership, (2) DOE WETO leadership and program managers, (3) National lab leadership, (4) Associated project principle investigators, (5) Research software engineers, (6) Wind energy researchers in academia (including graduate students, post docs, and national lab staff), (7) Industry researchers and practitioners, (8) Commercial software developers, and (9) The general public interested in wind energy. These software are typically the end-user of other generic software libraries, so the funding cycles are often tied to applied research rather than the development of the software itself. Since the developers are also wind energy researchers, these tools are typically designed in a way that closely resembles the application in which they're used. Additionally, the expertise and incentives for the developers have a high variability, and often neither are aligned with software engineering or computer science. Given the unique environment in which wind energy research software is produced and consumed, it is critical for model owners to understand the context of their software. A framework for developing this understanding is to answer the following questions of a given software project: What is it's purpose? What is its role in the field of wind energy? What is the profile of the expected users? For how long will it be relevant? What is the expected impact? These questions allow model owners to identify the appropriate methods for the design, development, and long term maintenance of their software. Additionally, the answer provide context for future planners to understand why particular decisions were made and discern the consequences of changing course. The information is aggregated from experience within WETO-supported software development groups as well as external organizations and efforts to define the craft of research software engineering. These best practices aim to make the collaborative development process efficient and effective while improving the model understanding across stakeholders. Additionally, the general adoption of a common framework for software quality ensures that the end users of WETO software can trust these tools and accurately understand the risks to workflow integration.

17 WIND ENERGY↗

Applying the FAIR Principles to computational workflows

Recent trends within computational and data sciences show an increasing recognition and adoption of computational workflows as tools for productivity and reproducibility that also democratize access to platforms and processing know-how. As digital objects to be shared, discovered, and reused, computational workflows benefit from the FAIR principles, which stand for Findable, Accessible, Interoperable, and Reusable. The Workflows Community Initiative’s FAIR Workflows Working Group (WCI-FW), a global and open community of researchers and developers working with computational workflows across disciplines and domains, has systematically addressed the application of both FAIR data and software principles to computational workflows. We present recommendations with commentary that reflects our discussions and justifies our choices and adaptations. These are offered to workflow users and authors, workflow management system developers, and providers of workflow services as guidelines for adoption and fodder for discussion. The FAIR recommendations for workflows that we propose in this paper will maximize their value as research assets and facilitate their adoption by the wider community.

97 MATHEMATICS AND COMPUTING↗

Assessing the Needs of NASA's Near Real-Time Earth Observation Products

"The 2017-2027 Decadal Survey for Earth Science and Applications from Space stated that NASA's Earth Science with planned implementation of applications provides sustained earth observations for societal benefits [1]. The Decadal Survey indicated that data latency is invaluable for time-sensitive applications including disaster risk reduction, wildland fire carbon emissions quantification, real-time measurements of the state of the hydrologic systems and many more. Data latency refers to the time between earth observation and data products available to users. During the past 13 years, NASA's Land, Atmosphere Near Real-Time Capability for Earth Observing Systems (LANCE) continues to provide free access to earth observation products that are made available much quicker than routine processing allows. The latency of most LANCE data products is Near Real-time (NRT) which is defined as less than three hours from satellite observations [2]. LANCE is managed by the Earth Science Data and Information System (ESDIS) Project at NASA Goddard Space Flight Center [3], and a User Working Group (UWG) is responsible for providing guidance to LANCE. LANCE data are used by direct users and brokers who add value to the data [4]. NASA Earth Applied Sciences Program (ASP) is one of the primary users of LANCE, which collaborates with partner organizations and provides support to scientists to solve problems in applications of earth observations. ASP promotes the use of LANCE NRT data products to demonstrate applications in decision making, facilitates end-user feedback to the science team to improve data products, and provides information on future demands for research. LANCE supports applications that need a rapid response including detecting wildland fires and volcanic eruptions, tracking smoke, ash and dust plumes, monitoring air quality and tracking extreme weather events such as hurricanes, landslides, and floods. To gather feedback regarding the availability, accessibility and actionability of NASA's NRT data products for societal benefit, three surveys and a few discussions with experts involved in the topic within ASP were conducted from the perspective of users. Feedback has been collected from users who are interested in using low latency NASA data within application communities of agriculture, disasters, water resources, health and air quality, ecological conservation, wildland fires and capacity building. Analysis-ready NRT data products in a variety of formats have been mentioned many times in the collected feedback, especially for applied users with little to no experience using research-grade earth observation products. Users prefer to have products that can be easily integrated into their existing workflows and take their analysis to the data. HDF5 is a commonly used data format for research, but typically requires some conversion to a more friendly format for applications and regular use in decision-making. Users prefer the GeoTIFF data format that can be directly ingested into a GIS mapping software and platform for data analysis and visualization. For example, LANCE’s fire, flood, SO2 and Black Marble Nighttime Blue/Yellow Composite data products have been integrated into NASA Disasters Mapping Portal, which is an GIS-based open data portal, for users in the disaster management community. There are 291 LANCE NRT layers available through GIBS and Worldview, where users can download a snapshot in GeoTIFF format. Operational users expect data to be processed as close to the user as possible. The collected feedback indicates that LANCE fire products within 3 hours latency would meet the needs of the wildland fire community. The ideal latency for volcanic application is 10-15 minutes. Users in Volcanic Ash Advisory Centers (VAAC) reported that the first forecast volcanic product should be issued within 75 minutes from the volcano eruption [5]. Overall, for disaster applications, data latency within 3 hours is useful while latency greater than 12 hours is not timely enough for operational use. Capacity building and training are critical for users to be able to access, interpret and use data products and tools for their decision making, especially for applied users with limited experience using earth observation products. LANCE data products have been used in a number of capacity building projects domestically and internationally [6]. As LANCE continues to bring new products into the system, users request training to utilize LANCE new and upcoming data products and capabilities in their applications. Due to the limitation of bandwidth and downstream flow paths, users in some developing countries need tools to select and download data for a specific area of interest instead of bulk downloads. The collected feedback also shows the lack of available SAR satellite low latency data products. The advantages of SAR to monitor conditions and changes on the ground through darkness, clouds, volcanic ash, and other atmospheric conditions, are appealing to low latency users. For example, terabytes of low latency but cloudy optical images are not helpful in rapidly identifying the extent of flood or fire impacts. LANCE could be complemented with low latency measurements via the upcoming NASA-ISRO Synthetic Aperture Radar (NISAR) mission [7]. Requests for higher spatial resolution products are expressed. A user from the wildland fire management community reported that products with 30-m spatial resolution could be used to detect small fires. The 30-m Landsat OLI fire data is now part of NASA’s Fire Information for Resource Management System (FIRMS) US/Canada [8]. Within the open and free NASA resources, LANCE disseminates NRT data products in a manner that allows them to be accessible and understandable to both scientific and applied users. In many application areas, latency plays an important or even decisive role where low latency earth observations help people to observe areas of interest, detect and track changes in the environment and make timely decisions. NASA’s Earth Applied Sciences Program promotes the use of LANCE NRT products and builds a bridge between application users and research teams. The collected feedback indicates data latency within 3 hours is useful for most of the applications, and shows the needs of user-friendly, analysis-ready products, and requests training on LANCE’s new and upcoming data products. User feedback has been provided to LANCE UWG for guidance and recommendations, and for translating findings into something actionable.

Tian Yao↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗