Search NASA⌕ Search

SEARCH · Search NASA

Results for “software: data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An R Shiny graphical user interface for analyzing, visualizing, and interpreting high precision mass spectrometric data

There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2 , IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

Labone, Elizabeth↗

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

National User Resource for Biological Accelerator Mass Spectrometry (Final Report)

The National User Resource for Biological Accelerator Mass Spectrometry (User Resource) will provide isotopic analysis (primarily radiocarbon or 14C) by accelerator mass spectrometry (AMS) for NIH- funded researchers across the United States and will be the only User Resource of its type in the United States. The User Resource will provide measurement capability and expertise to a research community that requires highly sensitive, quantitative isotope analyses. Since commissioning a new accelerator mass spectrometer in June 2014, we have measured over 4000 samples a year for collaborators and service users. The User Resource will enable us to continue to meet these research needs, as well as provide for new users whose research programs would benefit from AMS as a measurement tool. The User Resource’s forte will be ultra-high sensitivity quantitation of radiocarbon and selected other radioisotopes for research studies where isotopes are required. Radioisotope labeling studies have been and will continue to be an important tool for addressing many complex biomedical science problems. AMS is a specialized and unique type of mass spectrometry that provides absolute quantitation of radiocarbon and other relevant radioisotopes with extreme sensitivity, having limits of detection in real samples on the order of a few attomol/mg of sample at measurement precisions of ~3%. It is the only instrumental method capable of quantifying radioisotope-labeled agents routinely in real-world samples with such precision and sensitivity. The sensitivity of AMS allows for the quantification of radiolabeled metabolites in extremely complex matrices of cells and organisms at very low concentrations and in small samples. AMS allows studies to be conducted without perturbing metabolism leading to more relevant quantification of metabolic rates and pathways. In addition, it enables quantification of pharmacokinetic and metabolic properties of toxicants at environmentally relevant concentrations in model systems as well as the ability to quantify pharmacokinetics and other molecular endpoints directly in humans. Such quantitative assessments can 1) improve risk assessment for toxicants, 2) address safety and efficacy considerations for therapeutic entities, 3) deepen understanding of xenobiotic and intermediary metabolism, 4) help understand the interactions between critical molecular pathways, and 5) improve efforts to model and predict various metabolic and biological states. These capabilities have been applied in a number of areas including research in carcinogenesis, toxicology, nutrition, pharmacology/drug development and basic biological science. As a NIGMS National Resource the National User Resource for Biological Accelerator Mass Spectrometry will help NIH funded scientists achieve a deeper understanding of the etiology of human health concerns by (1) enabling the quantification of pharmacokinetics and other molecular endpoints directly in humans; (2) offering the ability to conduct quantitative studies using biologics such as proteins or lipids, and thereby reducing the amount of radioisotope usage in biomedical labs; and (3) enabling more relevant studies of metabolic pathways in health and disease through the use of much lower, more biologically-relevant, concentrations of metabolic substrates in cells and intact organisms. Such studies support NIGMS’s basic biomedical research areas that contribute to the understanding of fundamental cellular and physiological principles and enable research supported by the Biophysics, Biomedical Technology, and Computational Biosciences (BBCB); Genetics and Molecular, Cellular, and Developmental Biology (GMCDB); Pharmacology, Physiology, Biological Chemistry (PPBC) and Training, Workforce Development, and Diversity (TWD) Divisions. Over the next five years, our goals are to: 1. Improve the efficiency of operation for AMS measurements through installation of new interfaces to our AMS systems, technical modifications to improve gas accepting ion source efficiency and upgrading our data analysis software for improved ease of use and data reporting. 2. Increase the accessibility and visibility of ultra-sensitive 14C measurements for the biomedical research community by training of new investigators and expanding our national user base. 3. Provide high throughput, ultra-sensitive 14C analysis for the NIGMS and NIH user community.

47 OTHER INSTRUMENTATION↗

Pando

SAND2025-02006O Pando is a distributed data analysis software tool. It is designed to handle large-scale graph analysis problems, often with a specific focus on blockchain/cryptocurrency data. Pando handles scalability by running on a distributed cluster of servers. Users can customize the output using the program’s plugin/extension design methodology. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Gabert, Kasimir↗

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

An R shiny graphical user interface for highprecision mass spectrometric data analysis

• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

LABONE, ELIZABETH↗

nautilus : boosting Bayesian importance nested sampling with deep learning

ABSTRACT We introduce a novel approach to boost the efficiency of the importance nested sampling (INS) technique for Bayesian posterior and evidence estimation using deep learning. Unlike rejection-based sampling methods such as vanilla nested sampling (NS) or Markov chain Monte Carlo (MCMC) algorithms, importance sampling techniques can use all likelihood evaluations for posterior and evidence estimation. However, for efficient importance sampling, one needs proposal distributions that closely mimic the posterior distributions. We show how to combine INS with deep learning via neural network regression to accomplish this task. We also introduce nautilus, a reference open-source python implementation of this technique for Bayesian posterior and evidence estimation. We compare nautilus against popular NS and MCMC packages, including emcee, dynesty, ultranest, and pocomc, on a variety of challenging synthetic problems and real-world applications in exoplanet detection, galaxy SED fitting and cosmology. In all applications, the sampling efficiency of nautilus is substantially higher than that of all other samplers, often by more than an order of magnitude. Simultaneously, nautilus delivers highly accurate results and needs fewer likelihood evaluations than all other samplers tested. We also show that nautilus has good scaling with the dimensionality of the likelihood and is easily parallelizable to many CPUs.

97 MATHEMATICS AND COMPUTING↗

Training and onboarding initiatives in high energy physics experiments

In this article we document the current analysis software training and onboarding activities in several High Energy Physics (HEP) experiments: ATLAS, CMS, LHCb, Belle II and DUNE. Fast and efficient onboarding of new collaboration members is increasingly important for HEP experiments. With rapidly increasing data volumes and larger collaborations the analyses and consequently, the related software, become ever more complex. This necessitates structured onboarding and training. Recognizing this, a meeting series was held by the HEP Software Foundation (HSF) in 2022 for experiments to showcase their initiatives. Here we document and analyze these in an attempt to determine a set of key considerations for future HEP experiments.

analysis software↗

The 200 Gbps Challenge: Imagining HL-LHC analysis facilities

The IRIS-HEP software institute, as a contributor to the broader HEP Python ecosystem, is developing scalable analysis infrastructure and software tools to address the upcoming HL-LHC computing challenges with new approaches and paradigms, driven by our vision of what HL-LHC analysis will require. The institute uses a "Grand Challenge" format, constructing a series of increasingly large, complex, and realistic exercises to show the vision of HL-LHC analysis. Recently, the focus has been demonstrating the IRIS-HEP analysis infrastructure at scale and evaluating technology readiness for production. As a part of the Analysis Grand Challenge activities, the institute executed a "200 Gbps Challenge", aiming to show sustained data rates into the event processing of multiple analysis pipelines. The challenge integrated teams internal and external to the institute, including operations and facilities, analysis software tools, innovative data delivery and management services, and scalable analysis infrastructure. The challenge showcases the prototypes - including software, services, and facilities - built to process around 200 TB of data in both the CMS NanoAOD and ATLAS PHYSLITE data formats with test pipelines. The teams were able to sustain the 200 Gbps target across multiple pipelines. The pipelines focusing on event rate were able to process at over 30 MHz. These target rates are demanding; the activity revealed considerations for future testing at this scale and changes necessary for physicists to work at this scale in the future. The 200 Gbps Challenge has established a baseline on today's facilities, setting the stage for the next exercise at twice the scale.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

WETO Stack [SWR-24-81]

The WETO Stack software is a collection of data, data analysis scripts, and website content produced through the Holistic Modeling Project’s Portfolio Coordination task. The purpose of the software is to characterize the collection of wind energy software projects supported by the U.S. Department of Energy’s Wind Energy Technologies Office. The included data contains characteristics of the actively funded software projects, and the data analysis scripts provide insights into the breadth, depth, and maturity of the collection of software projects. The website content presents the results of the analyses as well as reports from workshops and other events. It can be viewed at https://nrel.github.io/WETOStack/index.html.

Mudafort, Rafael↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

CDB-AP: An application for coincidence Doppler broadening spectroscopy analysis

Coincidence Doppler Broadening (CDB) Positron Annihilation Spectroscopy (PAS) is a material analysis technique that can be used to non-destructively measure characteristics of structural defects in samples. Analyzing and comparing large datasets obtained using this technique, however, can be complicated and time intensive. The Coincidence Doppler Broadening Analysis Program (CDB-AP) is a graphical user interface that facilitates rapid analysis of many data files while using transparent processes. It is already used in three laboratories at Idaho National Laboratory and can be used in laboratories worldwide.

36 MATERIALS SCIENCE↗

Periodicity significance testing with null-signal templates: reassessment of PTF’s SMBH binary candidates

Periodograms are widely employed for identifying periodicity in time series data, yet they often struggle to accurately quantify the statistical significance of detected periodic signals when the data complexity precludes reliable simulations. We develop a data-driven approach to address this challenge by introducing a null-signal template (NST). The NST is created by carefully randomizing the period of each cycle in the periodogram template, rendering it non-periodic. It has the same frequentist properties as a periodic signal template, and we show with simulations that the distribution of false positives is the same as with the original periodic template, regardless of the underlying data. Thus, performing a periodicity search with the NST acts as an effective simulation of the null (no-signal) hypothesis, without having to simulate the noise properties of the data. We apply the NST method to the supermassive black hole binaries (SMBHB) search in the Palomar Transient Factory (PTF), where Charisi et al. had previously proposed 33 high signal-to-noise candidates utilizing simulations to quantify their significance. Our approach reveals that these simulations do not capture the complexity of the real data. There are no statistically significant periodic signal detections above the non-periodic background. To improve the search sensitivity, we introduce a Gaussian quadrature based algorithm for the Bayes Factor with correlated noise as a test statistic. We show with simulations that this improves sensitivity to true signals by more than an order of magnitude. However, the Bayes Factor approach also results in no statistically significant detections in the PTF data.

79 ASTRONOMY AND ASTROPHYSICS↗

Plutonium Basket Counter Measurement Control Q3-Q4 2023

The Plutonium Basket Counter (PBC) is a spent fuel nuclear safeguards instrument that was developed to measure spent-fuel elements from MAGNOX-type research reactors and determine their plutonium content. The instrument is designed to measure fuel elements in spent fuel cooling pools through underwater operation, but it is also able to measure radiation sources in air. The PBC determines fuel plutonium content by detecting neutrons with an array of Helium-3 detectors. The electronics make use of a JSR-15 shift register for data collection and IAEA Neutron Coincidence Counting (INCC) software for data analysis. The PBC is used to determine the 240 Pu content in spent MAGNOX fuel elements grouped into basket-like bundles, thus the instrument’s name. Reactor burnup calculations can be used in conjunction with the data from the PBC to estimate the total plutonium content in the fuel.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Nucleus++ : a new tool bridging AME and NUBASE for advancing nuclear data analysis

The newly developed software, Nucleus++ , is an advanced tool for displaying basic nuclear physics properties from NUBASE and integrating comprehensive mass information for each nuclide from Atomic Mass Evaluation. Additionally, it allows users to compare experimental nuclear masses with predictions from different mass models. Building on the success and learning experiences of its predecessor, Nucleus , this enhanced tool introduces improved functionality and compatibility. With its user-friendly interface, Nucleus++ was designed as a valuable tool for scholars and practitioners in the field of nuclear science. Finally, this article offers an in-depth description of Nucleus++ , highlighting its main features and anticipated impacts on nuclear science research.

AME↗