Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Constraining the impact of chlorine as a neutron absorber in next-gen fast reactor designs

The role of chlorine as a neutron poison and as a seed for producing radioactive waste in nuclear systems has driven a renewed interest to improve its nuclear data uncertainties. Additionally, basic and applied science programs that use CLYC (Cs 2 LiYCl 6 :Ce) detectors for neutron spectroscopy and monitoring are also very sensitive to any change in chlorine nuclear data for simulations of the detector response. In this work, sensitivities relevant for these different applications are addressed through simulations of the efficiency of CLYC detectors in a fast fission spectrum when applying new chlorine nuclear data as input. These simulations are validated by an experimental measurement using CLYC detectors coupled to an ionization chamber loaded with a 252 Cf spontaneous fission source. The results are then used to obtain the first reliable direct measurement of the 35 Cl(n,p 0 ) and summed Cl(n,p+n,α) fission spectrum average cross sections, found to be 54.7(32) and 105.0(98) mb, respectively. The results are within uncertainty of calculated fission spectrum averaged cross sections based on recently re-evaluated chlorine nuclear data, which confirm recent impact studies performed for the Molten Chloride Reactor Experiment. Meanwhile, there currently exists only one published criticality benchmark experiment that is sufficiently sensitive to chlorine nuclear data. Discrepancies are found with this set of criticality safety benchmarks, which are more sensitive to thermal and epithermal neutron energies than the energies, above 100 keV, tested in this current work. Hence, there is still a need to re-evaluate the chlorine nuclear data at lower energies to assess these discrepancies. Interpretation of the data from future “faster” criticality benchmarks, which are needed for next-gen fast reactor designs, benefit from the improved constraints on the chlorine nuclear data validated in this work.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Reflections on the Shifting Experiences of Scientific Infrastructure

Infrastructure of all types is fundamental to modern work and life. Computing for scientific work, especially, extends from distributed local research sites, often at the edges of other major systems, outward into globally connected high-performance facilities and infrastructures. This commentary reviews longstanding research on the social characteristics of infrastructure. We reflect on social concerns that affect the ongoing development, use, and maintenance of a wide range of scientific computing and data resources. Reflecting on the social nature of infrastructure is timely for Computing in Science & Engineering readers, given continued emphasis on developing even more expansive platforms for data and artificial intelligence work in science (e.g., the United States’ Genesis Mission). We assert that, regardless of technological advances, the complex nature of scientific research and data will require continued understanding of longstanding and nascent social practices across varied communities. This is fundamentally necessary to build and sustain usable infrastructure or platforms that can productively advance scientific research.

Paine, Drew [Lawrence Berkeley National Laboratory↗

Beta-delayed gamma spectra compilation and analysis following the thermal-neutron induced fission of 235 U, 239,241 Pu

The integral gamma and electron spectra emitted by fission products, also known as delayed gamma and electron spectra, were measured at Oak Ridge National Laboratory in the 1970s for the thermal-neutron induced fission of 235 U and 239,241 Pu. Scintillator detectors were used to measure these spectra, data used later on to obtain decay heat values - that is, the spectra mean values per unit time as function of time - work that was published in the Nuclear Science and Technology journal; the spectral data, however, was only published in laboratory reports. Here, in this work, we analyze the gamma spectra data using modern methods and nuclear databases to reveal the signature of individual fission products as well as to gauge the performance of the ENDF/B-VIII.0 decay data sub-library, concluding about possible future enhancements in predictive capabilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Intelligent experiments through real-time AI: Fast Data Processing and Autonomous Detector Control for sPHENIX and future EIC detectors (Phase-I)

With an ever increasing demand for high precision data from modern detectors for discovery science and precision measurements, all major high energy nuclear and particle experiments, current and future, are facing the challenge on how to deal with the large volume of raw data generated from sophisticated state-of-the-art detectors in high rate collisions. These goals need to be balanced with available hardware and cost limits on DAQ (Data AcQuisition system) bandwidth and offline computing resources to capture, store and process the signal events. Two prototypical examples are the upcoming sPHENIX experiment, the DOE next generation heavy ion physics experiment at the Relativistic Heavy Ion Collider at BNL, and the future EIC experiments that are planned to be online circa 2030.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Computational Modeling of Atmospheric Processes at Texas Southern University

Texas Southern University (TSU) is strengthening its research program in atmospheric chemistry and physics with a climate science emphasis by leveraging partnerships with the U.S. Department of Energy’s Atmospheric Radiation Measurement (ARM) Facility, Brookhaven National Laboratory (BNL), and the Tracking Aerosol Convection Interactions ExpeRiment (TRACER). This RDPP-supported program focuses on secondary organic aerosols (SOAs) and reactive atmospheric species that influence cloud formation, precipitation processes, and radiative forcing. SOAs play a critical role in cloud microphysics and Earth’s energy balance, yet the chemical and physical mechanisms governing SOA–cloud interactions remain a significant source of uncertainty in predictive climate models. Through computational modeling, observational data analysis, and national laboratory collaboration, this program develops a skilled cohort of students trained in atmospheric science, environmental data analysis, and climate-relevant modeling. These research experiences build technical competencies that are transferable to careers in government laboratories, academia, and industry. By engaging students from historically underrepresented communities in high-impact climate research, TSU expands participation in the atmospheric sciences workforce while contributing meaningful scientific insights to DOE-supported ARM research activities. This partnership strengthens national capacity in climate science and supports the development of the next generation of atmospheric researchers.

54 ENVIRONMENTAL SCIENCES↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Cracking the code of multi-layer films to promote circularity in single-use plastic packaging

Multi-layer film packaging (MLF) revolutionized food preservation by combining diverse material layers to optimize barrier properties, mechanical strength, and shelf-life. These materials are essential for transporting perishables across various climates and allow for access to fresh goods in “food deserts”, but they pose significant recycling challenges due to their structural complexity. This perspective examines key structure-property relationships governing barrier performance and highlights innovations in material design. We explore how machine learning can predict performance metrics and propose recyclable alternatives, integrating data-driven approaches with material science insights. By challenging the status quo of MLF design, we advocate for circularity in food packaging, inspiring innovation at the intersection of sustainability, material science, and artificial intelligence.

36 MATERIALS SCIENCE↗

First Constraint on Atmospheric Millicharged Particles with the LUX-ZEPLIN Experiment

We report on a search for millicharged particles (mCPs) produced in cosmic ray atmospheric interactions using data collected during the first science run of the LUX-ZEPLIN experiment. The mCPs produced by two processes—meson decay and proton bremsstrahlung—are considered in this study. This search utilized a novel signature unique to liquid xenon (LXe) time projection chambers, allowing sensitivity to mCPs with masses ranging from 10 to 1000 MeV/c 2 and fractional charges between 0.001 and 0.02 of the electron charge (𝑒). With an exposure of 60 live days and a 5.5 metric ton fiducial mass, we observed no significant excess over background. This represents the first experimental search for atmospheric mCPs and the first search for mCPs using an underground LXe experiment.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine learning approaches for crystallographic classification from synthetic 2D X-ray diffraction data

Crystallographic structure identification is crucial for understanding material properties; however, current methodologies often depend on labor-intensive and time-consuming analyses of 2D X-ray diffraction (XRD) patterns. To address these limitations, this study employs synthetic 2D XRD patterns combined with deep learning (DL) techniques to enable automated and high-throughput classification of the seven crystal systems and 230 space groups. We introduce the novel Auto Diffraction Pipeline, designed to generate synthetic 2D XRD spot patterns from crystallographic information files under diverse conditions, including varying zone axes, atomic substitution, atomic depletion and mechanical loading. These conditions enhance the realism of synthetic data, mitigating the scarcity of experimental datasets and enabling the creation of large representative training sets. Convolutional neural networks were trained and validated on these synthetic datasets to classify crystallographic structures across multiple scenarios. Our results demonstrate that integrating synthetic 2D XRD patterns with DL facilitates rapid, accurate and automated crystallographic classification, promoting the wider adoption of data-driven approaches in materials science.

Shahnazari, Ayoub [Univ. of Rochester, NY (United ↗

Prediction of plant complex traits via integration of multi-omics data

The formation of complex traits is the consequence of genotype and activities at multiple molecular levels. However, connecting genotypes and these activities to complex traits remains challenging. Here, we investigate whether integrating genomic, transcriptomic, and methylomic data can improve prediction for six Arabidopsis traits. We find that transcriptome- and methylome-based models have performances comparable to those of genome-based models. However, models built for flowering time using different omics data identify different benchmark genes. Nine additional genes identified as important for flowering time from our models are experimentally validated as regulating flowering. Gene contributions to flowering time prediction are accession-dependent and distinct genes contribute to trait prediction in different genotypes. Models integrating multi-omics data perform best and reveal known and additional gene interactions, extending knowledge about existing regulatory networks underlying flowering time determination. These results demonstrate the feasibility of revealing molecular mechanisms underlying complex traits through multi-omics data integration.

59 BASIC BIOLOGICAL SCIENCES↗

Real-time data processing for serial crystallography experiments

We report the use of streaming data interfaces to perform fully online data processing for serial crystallography experiments, without storing intermediate data on disk. The system produces Bragg reflection intensity measurements suitable for scaling and merging, with a latency of less than 1 s per frame. Our system uses the CrystFEL software in combination with the ASAP::O data framework. In a series of user experiments at PETRA III, frames from a 16 megapixel Dectris EIGER2 X detector were searched for peaks, indexed and integrated at the maximum full-frame readout speed of 133 frames per second. The computational resources required depend on various factors, most significantly the fraction of non-blank frames ('hits'). The average single-thread processing time per frame was 242 ms for blank frames and 455 ms for hits, meaning that a single 96-core computing node was sufficient to keep up with the data, with ample headroom for unexpected throughput reductions. Further significant improvements are expected, for example by binning pixel intensities together to reduce the pixel count. We discuss the implications of real-time data processing on the `data deluge' problem from recent and future photon-science experiments, in particular on calibration requirements, computing access patterns and the need for the preservation of raw data.

47 OTHER INSTRUMENTATION↗

EPCAPE Radar b1 Data Processing: Corrections, Calibrations, and Processing Report

The U.S. Department of Energy (DOE)’s Atmospheric Radiation Measurement (ARM) user facility recently deployed its First ARM Mobile Facility (AMF1) to La Jolla, California as part of the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) campaign. Some of the goals behind EPCAPE were to characterize the diurnal and seasonal cycles of stratocumulus clouds and to investigate the cloud-aerosol-radiation interactions and feedbacks in the area. The deployment of the AMF1 for a full year from 15 February 2023 to 14 February 2024 aided in addressing these scientific questions. While AMF1 collected data year-round, enhanced measurements were taken during two intensive operational periods (IOPs). The first IOP occurred from April to June and focused on the chemistry of low clouds (EPCAPE_Chem), while the second IOP occurred from July to September and was focused on the radiation of high clouds (EPCAPE_Radiation). Several cloud radars were deployed with AMF1 to collect valuable data on cloud properties that will help users address key science objectives. As in past ARM campaigns, a1-level radar data is extensively analyzed and calibration techniques are performed to generate b1-level data (Matthews et al. 2023, Feng et al. 2024). Radar data at the b1-level are of the highest quality and thus can be used to examine scientific questions. The status of the a1-level data and the a1-to-b1 process for the EPCAPE radars is subsequently detailed in this document.

54 ENVIRONMENTAL SCIENCES↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software

Scientific software is essential to scientific innovation and in many ways it is distinct from other types of software. Abandoned (or unmaintained), buggy, and hard to use software, a perception often associated with scientific software can hinder scientific progress, yet, in contrast to other types of software, its longevity is poorly understood. Existing data curation efforts are fragmented by science domain and/or are small in scale and lack key attributes. We use large language models to classify public software repositories in World of Code into distinct scientific domains and layers of the software stack, curating a large and diverse collection of over 18,000 scientific software projects. Using this data, we estimate survival models to understand how the domain, infrastructural layer, and other attributes of scientific software affect its longevity. We further obtain a matched sample of non-scientific software repositories and investigate the differences. We find that infrastructural layers, downstream dependencies, mentions of publications, and participants from government are associated with a longer lifespan, while newer projects with participants from academia had shorter lifespan. Against common expectations, scientific projects have a longer lifetime than matched non-scientific open-source software projects. We expect our curated attribute-rich collection to support future research on scientific software and provide insights that may help extend longevity of both scientific and other projects.

Malviya Thakur, Addi [ORNL] (ORCID:000000022681999↗