Search NASA⌕ Search

SEARCH · Search NASA

Results for “science data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

High-throughput single-cell transcriptomics of bacteria using combinatorial barcoding

Microbial split-pool ligation transcriptomics (microSPLiT) is a high-throughput single-cell RNA sequencing method for bacteria. With four combinatorial barcoding rounds, microSPLiT can profile transcriptional states in hundreds of thousands of Gram-negative and Gram-positive bacteria in a single experiment without specialized equipment. As bacterial samples are fixed and permeabilized before barcoding, they can be collected and stored ahead of time. During the first barcoding round, the fixed and permeabilized bacteria are distributed into a 96-well plate, where their transcripts are reverse transcribed into cDNA and labeled with the first well-specific barcode inside the cells. The cells are mixed and redistributed two more times into new 96-well plates, where the second and third barcodes are appended to the cDNA via in-cell ligation reactions. Finally, the cells are mixed and divided into aliquot sub-libraries, which can be stored until future use or prepared for sequencing with the addition of a fourth barcode. It takes 4 days to generate sequencing-ready libraries, including 1 day for collection and overnight fixation of samples. Here, the standard plate setup enables single-cell transcriptional profiling of up to 1 million bacterial cells and up to 96 samples in a single barcoding experiment, with the possibility of expansion by adding barcoding rounds. The protocol requires experience in basic molecular biology techniques, handling of bacterial samples and preparation of DNA libraries for next-generation sequencing. It can be performed by experienced undergraduate or graduate students. Data analysis requires access to computing resources, familiarity with Unix command line and basic experience with Python or R.

59 BASIC BIOLOGICAL SCIENCES↗

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

Advanced Data Science Model for Detecting Intelligent Malware

This study focused on developing a robust artificial intelligence (AI) model capable of detecting and characterizing advanced malware in Internet of Things (IoT) devices using network data. By analyzing network traffic with various machine learning (ML) models, our AI model can identify and characterize malicious activities to significantly improve malware detection accuracy and reliability as compared to traditional methods. The developed AI/ML model was trained using network data from IoT devices, leveraging classifiers such as Random Forest, Gradient Boosting, AdaBoost, and others to optimize detection performance. This project demonstrates a scalable framework for real-time malware detection and characterization in IoT networks, capable of identifying infected devices and facilitating the necessary steps to remove or isolate them, thereby preventing further infections. Although digital twin (DT) integration is not yet implemented in the current model, it represents a promising future enhancement. By creating a virtual replica of physical IoT devices, DT technology would allow for real-time monitoring and analysis without directly accessing operational technology, thus reducing the risk of compromising or reducing the performance of actual devices. This integration would further enhance the security of IoT ecosystems, combining AI technology to better flag and detect indications of malware-infected devices within a nuclear system environment.

42 ENGINEERING↗

Are Atmospheric Models Too Cold in the Mountains? The State of Science and Insights from the SAIL Field Campaign

Mountains play an outsized role in water resource availability, and the amount and timing of water they provide depend strongly on temperature. To that end, we ask the question: How well are atmospheric models capturing mountain temperatures? We synthesize results showing that high-resolution, regionally relevant climate models produce 2-m air temperature (T2m) measurements colder than what is observed (a “cold bias”), particularly in snow-covered midlatitude mountain ranges during winter. We find common cold biases in 44 studies across global mountain ranges, including single-model and multimodel ensembles. We explore the factors driving these biases and examine the physical mechanisms, data limitations, and observational uncertainties behind T2m. Our analysis suggests that the biases are genuine and not due to observation sparsity or resolution mismatches. Cold biases occur primarily on mountain peaks and ridges, whereas valleys are often warm biased. Our literature review suggests that increasing model resolution does not clearly mitigate the bias. By analyzing data from the Surface Atmosphere Integrated Field Laboratory (SAIL) field campaign in the Colorado Rocky Mountains, we test various hypotheses related to cold biases and find that local wind circulations, longwave (LW) radiation, and surface-layer parameterizations contribute to the T2m biases in this particular location. We conclude by emphasizing the value of coordinated model evaluation and development efforts in heavily instrumented mountain locations for addressing the root cause(s) of T2m biases and improving predictive understanding of mountain climates.

54 ENVIRONMENTAL SCIENCES↗

Are Atmospheric Models Too Cold in the Mountains? The State of Science and Insights from the SAIL Field Campaign

Mountains play an outsized role in water resource availability, and the amount and timing of water they provide depend strongly on temperature. To that end, we ask the question: How well are atmospheric models capturing mountain temperatures? We synthesize results showing that high-resolution, regionally relevant climate models produce 2-m air temperature (T2m) measurements colder than what is observed (a “cold bias”), particularly in snow-covered midlatitude mountain ranges during winter. We find common cold biases in 44 studies across global mountain ranges, including single-model and multimodel ensembles. We explore the factors driving these biases and examine the physical mechanisms, data limitations, and observational uncertainties behind T2m. Our analysis suggests that the biases are genuine and not due to observation sparsity or resolution mismatches. Cold biases occur primarily on mountain peaks and ridges, whereas valleys are often warm biased. Our literature review suggests that increasing model resolution does not clearly mitigate the bias. By analyzing data from the Surface Atmosphere Integrated Field Laboratory (SAIL) field campaign in the Colorado Rocky Mountains, we test various hypotheses related to cold biases and find that local wind circulations, longwave (LW) radiation, and surface-layer parameterizations contribute to the T2m biases in this particular location. We conclude by emphasizing the value of coordinated model evaluation and development efforts in heavily instrumented mountain locations for addressing the root cause(s) of T2m biases and improving predictive understanding of mountain climates.

54 ENVIRONMENTAL SCIENCES↗

Field Validation of Cloud Properties Sensor Field Campaign Report

The purpose of this campaign is to deploy Aerodyne Research Inc.’s extended wavelength cloud optical properties sensor (TWST-EN) in an operationally relevant environment with co-located, validated sensors. The Atmospheric Radiation Measurement (ARM) user facility’s Southern Great Plains (SGP) observatory is ideal for this deployment because of the variety of operational sensors that can measure some of the same cloud properties using different modalities. While cloud property sensors have long existed, they tend to be costly to produce and maintain. Our sensor measures absolute spectral radiance in the two bands and will retrieve cloud optical depth (COD), droplet effective radius, and thermodynamic phase. Our prototype is built predominantly from off-the-shelf components and uses uncooled spectrometers. A lower-cost, easy-to-use sensor such as this could allow deployment at many more sites for greater spatial coverage. Analysis and retrieval algorithm development using the data from this deployment has been a central technical objective of our U.S. Department of Energy Small Business Innovative Research (SBIR) Phase 2 contract (DE-SC0020473: Low-Cost Shortwave Spectroradiometer for Retrieval of Cloud Properties).

54 ENVIRONMENTAL SCIENCES↗

Fractionation of Filamentous Algae from Mixed Biofilms

Filamentous algae, which grow in long, hair-like filaments within biofilms, play a crucial role in wastewater treatment due to their ability to produce significant biomass and their resistance to predation compared to traditional microalgal treatments. These algae can effectively uptake and utilize pollutants, particularly excessive nitrogen (ammonia, nitrate, nitrite) and phosphorus (phosphate), making filamentous algae valuable for wastewater treatment, as well as bioethanol and biodiesel production due to high lipid productions. However, each algal species possesses different capacities, necessitating a thorough genetic identification and understanding of each community. A major challenge in accurately assessing these communities is the lack of coverage in large sequencing databases which can lead to misrepresentation of the true composition and abundance of organisms and overall sequencing bias. To address this, I evaluated chemical and physical techniques for separating filamentous algae from mixed biofilms to achieve clean genetic sequencing results. I employed pH washing (0.001M HCl, 0.001M HCl, DiH2O, 0.0001M HCl, 0.001M HCl) for chemical treatment, followed by physical separation through centrifugation (5000rpm, 6500rpm) or filtration (2mm, 250um, 75um). The most successful method was deionized water washing, which yielded clear differences across stacked filters; the 2mm filtrate showed high levels of filamentous algae, with microalgae eluting in the 75um filtrate or remaining within agglutinations of algae larger filters. Base washing eluted the highest concentrations of microalgae, with larger filter sizes retaining more filamentous algae, indicating the breakdown of extracellular polymeric substances (EPS). Our downstream plans include sending the high-throughput next-generation sequencing to confirm the purity and ratios of filamentous and non-filamentous algae, as well as bacteria present, thereby validating the success of our treatments. Potential applications include creating community-based fractions for analysis, refining current sequencing data with clearer isolations, and generating designer biofilms to enhance our understanding of community interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Training and onboarding initiatives in high energy physics experiments

In this article we document the current analysis software training and onboarding activities in several High Energy Physics (HEP) experiments: ATLAS, CMS, LHCb, Belle II and DUNE. Fast and efficient onboarding of new collaboration members is increasingly important for HEP experiments. With rapidly increasing data volumes and larger collaborations the analyses and consequently, the related software, become ever more complex. This necessitates structured onboarding and training. Recognizing this, a meeting series was held by the HEP Software Foundation (HSF) in 2022 for experiments to showcase their initiatives. Here we document and analyze these in an attempt to determine a set of key considerations for future HEP experiments.

analysis software↗

G-Mapper: Learning a Cover in the Mapper Construction

The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a “nice” Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in 𝑘-means by iteratively applying the Anderson–Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. In conclusion, experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.

G-means clustering↗

Optimizing Batch Crystallization with Model-based Design of Experiments

Adaptive and self-optimizing intelligent systems such as digital twins are increasingly important in science and engineering. Digital twins utilize mathematical models to provide added precision to decision-making. However, physics-informed models are challenging to build, calibrate, and validate with existing data science methods. Model-based design of experiments (MBDoE) is a popular framework for optimizing data collection to maximize parameter precision in mathematical models and digital twins. In this work, we apply MBDoE, facilitated by the open-source package Pyomo.DoE, to train and validate mathematical models for batch crystallization. We quantitatively examined the estimability of the model parameters for experiments with different cooling rates. This analysis provides a quantitative explanation for the heuristic of using multiple experiments at different cooling rates.

Lynch, Hailey↗

Compressed sensing methods with applications to advanced air sampling

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

54 ENVIRONMENTAL SCIENCES↗

Overview of the SCEC/USGS Community Stress Drop Validation Study Using the 2019 Ridgecrest Earthquake Sequence

We present initial findings from the ongoing Community Stress Drop Validation Study to compare spectral stress-drop estimates for earthquakes in the 2019 Ridgecrest, California, sequence. This study uses a unified dataset to independently estimate earthquake source parameters through various methods. Stress drop, which denotes the change in average shear stress along a fault during earthquake rupture, is a critical parameter in earthquake science, impacting ground motion, rupture simulation, and source physics. Spectral stress drop is commonly derived by fitting the amplitude-spectrum shape, but estimates can vary substantially across studies for individual earthquakes. Sponsored jointly by the U.S. Geological Survey and the Statewide (previously, Southern) California Earthquake Center our community study aims to elucidate sources of variability and uncertainty in earthquake spectral stress-drop estimates through quantitative comparison of submitted results from independent analyses. The dataset includes nearly 13,000 earthquakes ranging from M 1 to 7 during a two-week period of the 2019 Ridgecrest sequence, recorded within a 1° radius. Here, in this article, we report on 56 unique submissions received from 20 different groups, detailing spectral corner frequencies (or source durations), moment magnitudes, and estimated spectral stress drops. Methods employed encompass spectral ratio analysis, spectral decomposition and inversion, finite-fault modeling, ground-motion-based approaches, and combined methods. Initial analysis reveals significant scatter across submitted spectral stress drops spanning over six orders of magnitude. However, we can identify between-method trends and offsets within the data to mitigate this variability. Averaging submissions for a prioritized subset of 56 events shows reduced variability of spectral stress drop, indicating overall consistency in recovered spectral stress-drop values.

58 GEOSCIENCES↗

Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry

Advancing automated classification of atmospheric aerosols from Single-Particle Mass Spectrometry (SPMS) data remains challenging due to overlapping ion signatures, compositional diversity, and limited labeled data. This study evaluates supervised and semi-supervised learning frameworks to enhance aerosol identification by jointly leveraging labeled and unlabeled spectra. Four models were compared: a supervised Support Vector Machine (SVM), a self-training SVM, a stacked autoencoder classifier, and a stacked autoencoder trained using a temporal-ensembling Mean Teacher approach. All models achieved high and stable accuracies (90.0 %–91.1 %), surpassing previous results on the same dataset (87 %) and matching the performance of state-of-the-art deep learning methods. Despite small global metric differences (≤ 1 %), semi-supervised variants yielded up to 5 %–10 % improvements for compositionally rare particle types – such as soot (0.77 % of spectra, F1-score: 0.93–0.97) and hazelnut pollen (0.98 % of spectra, F1-score: 0.97–1.00) – equating to roughly ∼ 187 additional correctly classified spectra. These gains are scientifically significant, as such rare particles exert disproportionate influence on radiative absorption and ice nucleation processes; their improved detection reduces modeled uncertainties in aerosol absorption optical depth and mixed-phase cloud ice nucleation rates. The models' residual misclassifications (≈ 9 %) largely arise from true spectral overlap among chemically adjacent species (e.g., Na- vs. K-feldspar, coated vs. uncoated feldspars), reflecting physical compositional continuity rather than algorithmic error. Collectively, these findings demonstrate that leveraging unlabeled data to learn robust spectral representations and refine classification enhances both fidelity and interpretability, bridging data-driven analysis with aerosol–climate process understanding.

54 ENVIRONMENTAL SCIENCES↗

Enabling Early Transient Discovery in LSST via Difference Imaging with DECam

We present SLIDE, a pipeline that enables transient discovery in data from the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), using archival images from the Dark Energy Camera as templates for difference imaging. We apply this pipeline to the recently released Data Preview 1 (DP1; the first public release of Rubin commissioning data) and search for transients in the resulting difference images. The image subtraction, photometry extraction, and transient detection are all performed on the Rubin Science Platform. We demonstrate that SLIDE effectively extracts clean photometry by circumventing poor or missing LSST templates. We identified 29 previously unreported transients, 12 of which would not have been detected based on the DP1 DiaObject catalog. SLIDE will be especially useful for transient analysis in the early years of LSST, when template coverage will be largely incomplete or when templates may be contaminated by transients present at the time of acquisition. We present multiband light curves for a sample of known transients, along with new transient candidates identified through our search. Finally, we discuss the prospects of applying this pipeline during the main LSST survey. Our pipeline is broadly applicable and will support studies of all transients with slowly evolving phases.

Dong, Yize 一泽董 [Harvard-Smithsonian Center for Ast↗

Graphite Oxidation Rate Study on ET-10 and ETU-10 Grades - Task 4: QA Support and Testing for Structural Graphite Oxidation

INL performed targeted oxidation tests to measure oxidation rates for samples of ET-10 and ETU-10 graphite under CRADA No. 21CRA22 Mod. 3, Annex A, “Tritium Testing to Support Kairos Power Advanced Reactor Demonstration” (04/02/2024). All testing was conducted within INL’s Carbon Characterization Laboratory (CCL) using test standard ASTM D7542-21 "Standard Test Method for Air Oxidation of Carbon and Graphite in the Kinetic Regime" [ASTM International, 2021]. Kairos Power provided all test specimens through its graphite vendor Ibiden, Inc. to INL and ASTM specimen specified dimensions. Information within this report only provides the Arrhenius oxidation rate plots as a function of temperature for each graphite grade tested. The raw mass loss per time data will be provided on the Nuclear Data Management and Analysis System (NDMAS) portal located on the INL information system.

36 MATERIALS SCIENCE↗

FY25 status report on the addition of candidate materials in Class B Code Case

This report provides the time-dependent allowable stress calculation strategy leveraging the limited creep rupture tests data generated to support the allowable stress for 100,000 hours in American Society of Mechanical Engineers (ASME) Boiler and Pressure Vessel Code (BPVC), Section II, Part D. A variable confidence index procedure to extrapolate material properties to support 500,000 hours design life is discussed. Time-dependent allowable stresses for Class B component design and analysis are presented for Grade 1 and Grade 2 of Alloy 625. The presented data extrapolation and allowable stress calculation method will support new material addition using limited creep rupture data in the new ASME Boiler and BPVC, Section III, Division 5, Class B rules.

Part D↗

Increasing Mosquito Abundance Under Global Warming

Mosquitoes are a key virus vector that poses significant health threats globally, affecting 700 million individuals and causing 1 million deaths annually. Accurately predicting mosquito abundance and dispersion remains a challenge. Complex interactions between mosquito dynamics and various environmental factors, notably hydrology, contribute to this challenge. Existing models typically focus on precipitation and temperature and often overlook further impacts of hydrological variables within mosquito modeling. In this study, we developed an artificial intelligence‐based model for mosquito dynamics, explicitly accounting for different hydrological variables, such as precipitation, soil moisture and streamflow. Using Toronto, Canada, as a case study, we identified causal relationships between changes in mosquito populations, hydrological factors, vegetation (e.g., leaf area index), and climate variables (e.g., daylight length, precipitation, and temperature). We embedded these relationships into a Long Short‐Term Memory (LSTM) Neural Network Model capable of accurately detecting mosquito dynamics across annual, seasonal, and monthly time scales. The LSTM is able to explain, on average, approximately 40% of the variance in the observed mosquito abundance data. Using the calibrated model, we predicted that the summer season mosquito abundance would increase by ∼16% and ∼19% under an intermediate greenhouse emission scenario, Shared Socioeconomic Pathway (SSP) 2–4.5, and a high greenhouse emission scenario, SSP5‐8.5, respectively. We expect that this model can serve as a valuable tool and inform science‐based decisions affecting mosquito dynamics and public health. It can also build a foundation for future risk analysis at the regional and larger scales.

54 ENVIRONMENTAL SCIENCES↗