Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Advances in the photon avalanche luminescence of inorganic lanthanide-doped nanomaterials

Photon avalanche (PA)—where the absorption of a single photon initiates a ‘chain reaction’ of additional absorption and energy transfer events within a material—is a highly nonlinear optical process that results in upconverted light emission with an exceptionally steep dependence on the illumination intensity. Over 40 years following the first demonstration of photon avalanche emission in lanthanide-doped bulk crystals, PA emission has been achieved in nanometer-scale colloidal particles. The scaling of PA to nanomaterials has resulted in significant and rapid advances, such as luminescence imaging beyond the diffraction limit of light, optical thermometry and force sensing with (sub)micron spatial resolution, and all-optical data storage and processing. In this review, we discuss the fundamental principles underpinning PA and survey the studies leading to the development of nanoscale PA. Finally, we offer a perspective on how this knowledge can be used for the development of next-generation PA nanomaterials optimized for a broad range of applications, including mid-IR imaging, luminescence thermometry, (bio)sensing, optical data processing and nanophotonics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BNF Radar b1 Data Processing Report: Spring 2025

The U.S. Department of Energy’s Atmospheric Radiation Measurement (ARM) User Facility supports atmospheric science through an integrated network of fixed and mobile observatories. These facilities collect continuous and campaign-based observations of atmospheric properties, with the goal of improving the representation of clouds, aerosols, precipitation, and radiation in Earth system models. The Bankhead National Forest (BNF) site, established as an ARM Mobile Facility (AMF) on 1 October 2024, is situated in a forested region of northern Alabama. Its strategic location in a southeastern U.S. environment characterized by complex terrain, diverse land cover, and frequent convective storms provides a valuable opportunity to examine coupled land-atmosphere processes under natural variability.

54 ENVIRONMENTAL SCIENCES↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

Quantum-centric supercomputing for materials science: A perspective on challenges and future directions

Computational models are an essential tool for the design, characterization, and discovery of novel materials. Computationally hard tasks in materials science stretch the limits of existing high-performance supercomputing centers, consuming much of their resources for simulation, analysis, and data processing. Quantum computing, on the other hand, is an emerging technology with the potential to accelerate many of the computational tasks needed for materials science. In order to do that, the quantum technology must interact with conventional high-performance computing in several ways: approximate results validation, identification of hard problems, and synergies in quantum-centric supercomputing. Here in this paper, we provide a perspective on how quantum-centric supercomputing can help address critical computational problems in materials science, the challenges to face in order to solve representative use cases, and new suggested directions.

36 MATERIALS SCIENCE↗

Addressing Limitations of the Endpoint Slippage Analysis

Some rate of oxidation and reduction side-reactions will inevitably coexist in most rechargeable batteries. While parasitic reduction traps electrons, parasitic oxidation donates electrons to the cell’s inventory and may cause temporary capacity gain. Consequently, capacity measurements can provide unreliable information about the total extent of side-reactions occurring in the cell. The most widely used method to determine the rate of both these parasitic processes involves analyzing the slippage of endpoints, which consists in tracking the termination of cell charge and discharge when data is represented along a cumulative capacity axis. Here, we argue that this approach could lead to inaccuracies when applied to certain systems, which includes Si electrodes in Li-ion batteries and hard carbon in Na-ion batteries. This inaccuracy originates from the smooth nature of the voltage profiles of these materials at low and high alkali-ion content, causing the termination of charge and discharge to be dictated by voltage changes at both the positive and negative electrodes. We analyze this issue in quantitative terms and propose equations that can provide true rates of parasitic processes from experimental endpoint slippage data. This work shows that, in battery science, well-established analytical approaches may not be directly transferrable to new electrode systems.

25 ENERGY STORAGE↗

The DECam MAGIC Survey $-$ Mapping the Ancient Galaxy in CaHK: Overview and Summary of Early Science

We present the DECam Mapping the Ancient Galaxy in CaHK (MAGIC) survey, a 54-night NOIRLab Survey Program to image $\gtrsim$5,000$\,$deg$^2$ of the southern hemisphere using a metallicity-sensitive narrow-band filter covering the Ca$\,$ii$\,$H&K lines centered at 3955$\,$A. This filter is installed on the Dark Energy Camera (DECam), mounted on the 4-m NSF Víctor M. Blanco Telescope. The survey reaches typical $10σ$ depths of $\text{mag}_{\text{CaHK}} \approx 22.5$, 3$-$4$\,$mag deeper than comparable surveys in the southern hemisphere. By combining photometry from this Ca$\,$ii$\,$H&K filter with existing DECam $g,r,i$ broadband photometry from the DECam Local Volume Exploration (DELVE) survey, MAGIC is deriving photometric metallicities for red giant branch stars down to the magnitude limit of usable proper motions from Gaia data release 3 (DR3). MAGIC has already imaged $\sim$3,000$\,$deg$^2$, supplemented by other affiliated observing programs that have used this filter to image star clusters, dwarf galaxies, and stellar streams. We overview MAGIC's survey strategy, describe data processing through the derivation of metallicities and photometric distances, and summarize early science results that have been published with this dataset. In addition, we present several new results, including the confirmation of a distant ($>5\,r_h$) member of the Reticulum II ultra-faint dwarf galaxy, on-sky density maps of low-metallicity stars into the distant Milky Way halo ($\sim150\,$kpc) recovering 13/14 ultra-faint dwarf galaxies in the current footprint, and a validation of our initial targeting of extremely metal-poor stars. Collectively, these results demonstrate that the MAGIC dataset enables cutting-edge studies of the faint, low-metallicity regime of the Milky Way and its substructures.

Chiti, A. [KIPAC, Menlo Park] (ORCID:0000000271556↗

Salk Institute for Biological Studies Requirements (Analysis Report)

EPOC uses the Deep Dive process to discuss and analyze current and planned science, research, or education activities and the anticipated data output of a particular use case, site, or project to help inform the strategic planning of a campus or regional networking environment. This includes understanding future needs related to network operations, network capacity upgrades, and other technological service investments. A Deep Dive comprehensively surveys major research stakeholders’ plans and processes in order to investigate data management requirements over the next 5–10 years. Between February and March 2024, staff members from the Engagement and Performance Operations Center (EPOC) met with researchers and staff from the Salk Institute for Biological Studies (Salk) for the purpose of a Deep Dive into scientific and research drivers. The goal of this activity was to help characterize the requirements for a number of campus use cases, and to enable cyberinfrastructure support staff better to understand the needs of the researchers within the community. Material for this event included the written documentation from each of the profiled research areas, documentation about the current state of technology support, and a write-up of the discussion that took place via e-mail and video conferencing. The case studies highlighted the ongoing challenges and opportunities that Salk Institute for Biological Studies have in supporting a cross-section of established and emerging research use cases. Each case study mentioned unique challenges which were summarized into common needs.

59 BASIC BIOLOGICAL SCIENCES↗

New York-Presbyterian and Columbia University Irving Medical Center Requirements (Analysis Report)

EPOC uses the Deep Dive process to discuss and analyze current and planned science, research, or education activities and the anticipated data output of a particular use case, site, or project to help inform the strategic planning of a campus or regional networking environment. This includes understanding future needs related to network operations, network capacity upgrades, and other technological service investments. A Deep Dive comprehensively surveys major research stakeholders’ plans and processes in order to investigate data management requirements over the next 5–10 years. Between February and June 2024, staff members from the Engagement and Performance Operations Center (EPOC) met with researchers and staff from New York-Presbyterian (NYP), Columbia University Irving Medical Center (CUIMC), and NYSERNet for the purpose of a Deep Dive into scientific and research drivers. The goal of this activity was to help characterize the requirements for a number of campus use cases, and to enable cyberinfrastructure support staff to better understand the needs of the researchers within the community. Material for this event included the written documentation from each of the profiled research areas, documentation about the current state of technology support, and a write-up of the discussion that took place via e-mail and video conferencing. The case studies highlighted the ongoing challenges and opportunities that NYP and CUIMC have in supporting a cross-section of established and emerging research use cases. Each case study mentioned unique challenges which were summarized into common needs.

97 MATHEMATICS AND COMPUTING↗

Particle Markov Chain Monte Carlo Approach to Inference in Transient Surface Kinetics

Here, in this work, we develop a novel Bayesian approach to study the adsorption and desorption of CO onto a Pd(111) surface, a process of great importance in natural sciences. The motivation for this work comes from the recent availability of time-resolved infrared spectroscopy data and the need for model interpretability and uncertainty quantification in chemical processes. The objective is to learn the relevant parameters that characterize the process: coverage with time, rate constants, activation energies, and pre-exponential factors. Our approach consists of three main schemes: (i) a problem design and probabilistic model for the whole system, (ii) a particle Markov chain Monte Carlo sampler to learn the hidden coverages and rate constant parameters, and (iii) two Bayesian formulations to infer the activation energies and pre-exponential factors. The flexibility of the Bayesian framework allows for uncertainty quantification where possible and integration of mathematical constraints in the model to reflect the system physically. We found that our results for the activation energies and pre-exponential factor are in agreement with those reported in the experimental literature, independently, and we provide discussions on the advantages and disadvantages as well as applicability to other systems.

36 MATERIALS SCIENCE↗

Data-scarce surrogate modeling of shock-induced pore collapse process

Understanding the mechanisms of shock-induced pore collapse is of great interest in various disciplines in sciences and engineering, including materials science, biological sciences, and geophysics. However, numerical modeling of the complex pore collapse processes can be costly. To this end, a strong need exists to develop surrogate models for generating economic predictions of pore collapse processes. Here, in this work, we study the use of a data-driven reduced-order model, namely dynamic mode decomposition, and a deep generative model, namely conditional generative adversarial networks, to resemble the numerical simulations of the pore collapse process at representative training shock pressures. Since the simulations are expensive, the training data are scarce, which makes training an accurate surrogate model challenging. To overcome the difficulties posed by the complex physics phenomena, we make several crucial treatments to the plain original form of the methods to increase the capability of approximating and predicting the dynamics. In particular, physics information is used as indicators or conditional inputs to guide the prediction. In realizing these methods, the training of each dynamic mode composition model takes only around 30 s on CPU. In contrast, training a generative adversarial network model takes 8 h on GPU. Moreover, using dynamic mode decomposition, the final-time relative error is around 0.3% in the reproductive cases. We also demonstrate the predictive power of the methods at unseen testing shock pressures, where the error ranges from 1.3 to 5% in the interpolatory cases and 8 to 9% in extrapolatory cases.

97 MATHEMATICS AND COMPUTING↗

Mind the gap: Bridging the divide between AI aspirations and the reality of autonomous microscopy

What does materials science look like in the “Age of Artificial Intelligence?” Each material’s domain—synthesis, characterization, and modeling—has a different answer to this question, motivated by unique challenges and constraints. This work focuses on the tremendous potential of autonomous characterization within electron microscopy. We present our recent advancements in developing domain-aware, multimodal models for microscopy analysis capable of describing complex atomic systems. We then address the critical gap between the theoretical promise of autonomous microscopy and its current practical limitations, showcasing recent successes while highlighting the necessary developments to achieve robust, real-world autonomy.

2D materials↗

Increasing the Reproducibility and Replicability of Supervised AI/ML in the Earth Systems Science by Leveraging Social Science Methods

Artificial intelligence (AI) and machine learning (ML) pose a challenge for achieving science that is both reproducible and replicable. The challenge is compounded in supervised models that depend on manually labeled training data, as they introduce additional decision-making and processes that require thorough documentation and reporting. We address these limitations by providing an approach to hand labeling training data for supervised ML that integrates quantitative content analysis (QCA)—a method from social science research. The QCA approach provides a rigorous and well-documented hand labeling procedure to improve the replicability and reproducibility of supervised ML applications in Earth systems science (ESS), as well as the ability to evaluate them. Specifically, the approach requires (a) the articulation and documentation of the exact decision-making process used for assigning hand labels in a “codebook” and (b) an empirical evaluation of the reliability” of the hand labelers. In this paper, we outline the contributions of QCA to the field, along with an overview of the general approach. We then provide a case study to further demonstrate how this framework has and can be applied when developing supervised ML models for applications in ESS. With this approach, we provide an actionable path forward for addressing ethical considerations and goals outlined by recent AGU work on ML ethics in ESS.

58 GEOSCIENCES↗

Constrained GAN-Generated X-Ray CT Data For Self-Supervised And Foundation-Model Segmentation Of Concrete Microstructures

Three-dimensional characterization of materials using X-ray computed tomography (XCT) is challenging due to the complexity of internal structures, noise, and variations in resolution. Traditional computer vision models often struggle to accurately segment these images, particularly in domain-specific applications like materials science. While supervised deep learning approaches have been developed to address the limitations of conventional algorithms, they typically require large amounts of labeled training data and often fail to generalize across different datasets. Self-supervised, few-and zero-shot learning methods have gained prominence in natural image processing and segmentation tasks, but their application to scientific imaging remains limited due to the unique structural complexity, noise, and textural artifacts present in materials science data. In this work, we investigate how domain adaptation, leveraging physics-based and GAN-generated synthetic data, impacts segmentation performance. We introduce a modified Contrastive Unpaired Translation (CUT) model designed to generate realistic labeled data, which can be used for training, pre-training, and fine-tuning segmentation models for real XCT microstructure data. We evaluate the performance of two segmentation approaches: a self-supervised network (SSL-ALPNet) and a foundation model (Segment Anything Model), assessing their improvements when pre-trained and/or fine-tuned on the synthesized data. Our results demonstrate that leveraging synthetic data significantly enhances segmentation performance, particularly in challenging materials science applications.

Ziabari, Amir [ORNL] (ORCID:000000034776457X)↗

Overview of the distributed image processing infrastructure to produce the Legacy Survey of Space and Time

The Vera C. Rubin Observatory is preparing to execute the most ambitious astronomical survey ever attempted, the Legacy Survey of Space and Time (LSST). Currently the final phase of construction is under way in the Chilean Andes, with the Observatory’s ten-year science mission scheduled to begin in 2025. Rubin’s 8.4-meter telescope will nightly scan the southern hemisphere collecting imagery in the wavelength range 320–1050 nm covering the entire observable sky every 4 nights using a 3.2 gigapixel camera, the largest imaging device ever built for astronomy. Automated detection and classification of celestial objects will be performed by sophisticated algorithms on high-resolution images to progressively produce an astronomical catalog eventually composed of 20 billion galaxies and 17 billion stars and their associated physical properties. In this article we present an overview of the system currently being constructed to perform data distribution as well as the annual campaigns which reprocess the entire image dataset collected since the beginning of the survey. These processing campaigns will utilize computing and storage resources provided by three Rubin data facilities (one in the US and two in Europe). Each year a Data Release will be produced and disseminated to science collaborations for use in studies comprising four main science pillars: probing dark matter and dark energy, taking inventory of solar system objects, exploring the transient optical sky and mapping the Milky Way. Also presented is the method by which we leverage some of the common tools and best practices used for management of large-scale distributed data processing projects in the high energy physics and astronomy communities. We also demonstrate how these tools and practices are utilized within the Rubin project in order to overcome the specific challenges faced by the Observatory.

79 ASTRONOMY AND ASTROPHYSICS↗

Density estimation via measure transport: Outlook for applications in the biological sciences

Abstract One among several advantages of measure transport methods is that they allow or a unified framework for processing and analysis of data distributed according to a wide class of probability measures. Within this context, we present results from computational studies aimed at assessing the potential of measure transport techniques, specifically, the use of triangular transport maps, as part of a workflow intended to support research in the biological sciences. Scenarios characterized by the availability of limited amount of sample data, which are common in domains such as radiation biology, are of particular interest. We find that when estimating a distribution density function given limited amount of sample data, adaptive transport maps are advantageous. In particular, statistics gathered from computing series of adaptive transport maps, trained on a series of randomly chosen subsets of the set of available data samples, leads to uncovering information hidden in the data. As a result, in the radiation biology application considered here, this approach provides a tool for generating hypotheses about gene relationships and their dynamics under radiation exposure.

gene expression data↗

Direct neutrino-mass measurement based on 259 days of KATRIN data

That neutrinos carry a nonvanishing rest mass is evidence of physics beyond the Standard Model of elementary particles. Their absolute mass holds relevance in fields from particle physics to cosmology. We report on the search for the effective electron antineutrino mass with the KATRIN experiment. KATRIN performs precision spectroscopy of the tritium β-decay close to the kinematic endpoint. On the basis of the first five measurement campaigns, we derived a best-fit value of $m^{2}_{v} = -0.14^{+0.13}_{-0.15}$ eV 2 , resulting in an upper limit of m ν < 0.45 eV at 90% confidence level. Stemming from 36 million electrons collected in 259 measurement days, a substantial reduction of the background level, and improved systematic uncertainties, this result tightens KATRIN’s previous bound by a factor of almost two.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗