Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Deep Learning Models for Planetary Seismicity Detection

Research in planetary seismology is fundamentally constrained by a lack of data. Seismo-logical science products of future missions can typically only be informed by theoretical signal/noise characteristics of the environment or likely Earth-analogues. Although objectives can be re-assessed after some initial data-collection upon lander arrival, transfer of high-resolution data back to Earth is costly on lander power usage. Over the last several years, development of GPU computing techniques and open-source high-level APIs have led to rapid advances in deep learning within the fields of computer vision, natural language processing, and collaborative filtering. These techniques are actively being adapted in seismology for a variety of tasks, including: earthquake detection, seismic phase discrimination, and ground-motion prediction. Until the recent detection of mars quakes during the Mars InSight mission, the only other measurements of seismicity recorded outside of Earth was on the Moon during the Apollo missions between 1969 to 1977. These unique data sets have been periodically revisited using new seismological methods, including ambient noise interferometry and Hidden Markov Models. Our objective is to develop a deep learning seismic detector and use it to catalog moonquakes from the Apollo 17 Lunar Seismic Profiling Experiment (LSPE) and compare the results with those obtained by other methods. Additionally, we will assess the accuracy tradeoff between using a training set of lunar data and one composed of Earth seismicity. In this document, we present preliminary results using a prototype classifier trained on a small set of earthquakes that was able to obtain detections for LSPE moonquakes with a greater accuracy than a recent study using Hidden Markov Models.

Civilini, F.↗

Uncertainty in inventories for life cycle assessment: State‐of‐the‐art, challenges, and new technologies

Uncertainty is a critical factor that can hinder the quality and potential applications of life cycle assessment (LCA) results. A prominent source of uncertainty stems from the life cycle inventory (LCI) data. Various methodologies exist to estimate the uncertainty associated with LCI data, primarily based on the widely used structured pedigree matrix approach or the computationally intensive Monte Carlo simulation. This perspective review explores how new technologies (e.g., computational algorithms and data collection methods) from data science and related fields can contribute to identifying, quantifying, and reducing uncertainty in LCI modeling. A brief overview of the sources of uncertainty in LCI modeling and how they are addressed in current LCA practice is provided. Additionally, several new technologies are identified, and the potential benefits of their implementation in reducing uncertainties in LCI modeling are discussed. This perspective review concludes by identifying potential areas that require further development for these technologies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Probing the scalar WIMP-pion coupling with the first LUX-ZEPLIN data

Weakly interacting massive particles (WIMPs) may interact with a virtual pion that is exchanged between nucleons. This interaction channel is important to consider in models where the spin-independent isoscalar channel is suppressed. Using data from the first science run of the LUX-ZEPLIN dark matter experiment, containing 60 live days of data in a 5.5 tonne fiducial mass of liquid xenon, we report the results on a search for WIMP-pion interactions. We observe no significant excess and set an upper limit of 1.5 × 10$^{−46}$ cm$^{2}$ at a 90% confidence level for a WIMP mass of 33 GeV/c$^{2}$ for this interaction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A comparison of probabilistic generative frameworks for molecular simulations

Generative artificial intelligence is now a widely used tool in molecular science. Despite the popularity of probabilistic generative models, numerical experiments benchmarking their performance on molecular data are lacking. Here, in this work, we introduce and explain several classes of generative models, broadly sorted into two categories: flow-based models and diffusion models. We select three representative models: neural spline flows, conditional flow matching, and denoising diffusion probabilistic models, and examine their accuracy, computational cost, and generation speed across datasets with tunable dimensionality, complexity, and modal asymmetry. Our findings are varied, with no one framework being the best for all purposes. In a nutshell, (i) neural spline flows do best at capturing mode asymmetry present in low-dimensional data, (ii) conditional flow matching outperforms other models for high-dimensional data with low complexity, and (iii) denoising diffusion probabilistic models appear the best for low-dimensional data with high complexity. Our datasets include a Gaussian mixture model and the dihedral torsion angle distribution of the Aib9 peptide, generated via a molecular dynamics simulation. We hope our taxonomy of probabilistic generative frameworks and numerical results may guide model selection for a wide range of molecular tasks.

Artificial intelligence↗

Novel application of neutrinos to evaluate U.S. nuclear weapons performance

There is a growing realization that neutrinos can be used as a diagnostic tool to better understand the inner workings of a nuclear weapon. Robust estimates demonstrate that an Inverse Beta Decay (IBD) neutrino scintillation detector built at the Nevada Test Site with a 1000-ton active target mass at a standoff distance of 500 m would detect thousands of antineutrino events per nuclear test. This would provide less than 4% statistical error on the measured antineutrino rate and 5% error on antineutrino energy. Extrapolating this to an error on the test device explosive yield requires knowledge from evaluated nuclear databases, non-equilibrium fission rates, and assumptions on internal neutron fluxes. Initial calculations demonstrate that the total number of neutrinos emitted per fission in the first 10 3 s after a short pulse of 239 Pu fission is about a factor of two less than that from Pu fissioning under steady state conditions. Furthermore, there are significant energy spectral differences as a function of time after the pulse that must be considered. These and other model dependencies will be discussed in the paper. In the absence of nuclear weapons testing, many of the technical and theoretical challenges of a full nuclear test could be mitigated with a low cost smaller scale 20 ton fiducial mass IBD demonstration detector placed near a pulsed reactor. Potential reactors include the Texas A&M University TRIGA 1 GW–10 ms pulsed facility or the Sandia Annular Core Research Reactor. The short duty cycle and repeatability of pulses would provide critical real environment testing and measurements, which would be valuable for planning a possible real test shot in the future. Furthermore, the antineutrino rate as a function of time data would provide unique constraints on fission databases and model assumptions. Finally, there are impactful science drivers such as sensitive searches for ∼1 eV 2 sterile neutrinos and ∼MeV scale axions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Tropical Rainfall Measuring (TRMM) - What Have We Learned and What Does the Future Hold?

Rainfall is important in the hydrological cycle and to the lives and welfare of humans. In addition to being a life-giving resource, rainfall processes also plays a crucial role in the dynamics of the global atmospheric circulation. Three-fourths of the energy that drives the atmospheric wind circulation comes from the latent heat released by tropical precipitation. It varies greatly in space and time. The rain-producing cloud systems may last several hours or days. Their dimensions range from 10 km to several hundred km. This makes it difficult to incorporate rainfall directly large-scale weather and climate models. Until the end of 1997, precipitation in the global tropics was not known to within a factor of two. Regarding "global warming", the various large-scale models differed among themselves in the predicted magnitude of the warming and in the expected regional effects of these temperature and moisture changes. The Tropical Rainfall Measuring Mission (TRMM) satellite has yielded important interim results related to rainfall observations, data assimilation and model forecast skills when rainfall data is assimilated. This talk will summarize where the TRMM science team is with regards to answering some of these important scientific challenges, as well as discuss the future Global Precipitation Mission which will provide 3 hourly rainfall coverage and offers some unique collaborative potential for NOAA and NASA.

Kummerow, C.↗

SN1987A: The Birth of a Supernova Remnant

This grant was intended to support the development of theoretical models needed to interpret and understand the observations by the Hubble Space Telescope and the Chandra X-ray telescope of the rapidly developing remnant of Supernova 1987A. In addition, we carried out a few investigations of related topics. The project was spectacularly successful. The models that we developed provide the definitive framework for predicting and interpreting this phenomenon. Following is a list of publications based on our work. Some of these papers include results of both theoretical modeling supported by this project and also analysis of data supported by the Space Telescope Science Institute and the Chandra X-ray Observatory. We first list papers published in refereed journals, then conference proceedings and book chapters, and also an educational web site.

McCray, Richard↗

Improving an Atlantic Fisheries DSS using Sea Surface Salinity Data from NASA's Aquarius Mission

This report assesses the capacity of incorporating NASA#s Aquarius SSS (sea surface salinity) data into the SMAST (School of Marine Science and Technology) DSS for Fisheries Science. This data will enhance the SMAST DSS by providing SSS over a large area. Aquarius is a focused satellite mission designed to measure global SSS. SSS mapping is limited because conventional in situ SSS sampling is too sparse to give a large-scale view of the salinity variability. Aquarius will resolve missing physical processes that link the water cycle, the climate, and the ocean. The SMAST Fisheries program provides a DSS for fisheries science. It collects fisheries and environmental data, integrates them into a suite of data assimilation ocean models, and provides hindcasts, nowcasts, and forecasts for fisheries research, fisheries management, and the fishery industry. Currently, SMAST is using SSS data from the National Oceanic and Atmospheric Administration#s National Data Buoy Center. The SMAST DSS would be enhanced with SSS data from the Aquarius mission.

Guest, DeNeice↗

Dynamic Black-Level Correction and Artifact Flagging for Kepler Pixel Time Series

Methods applied to the calibration stage of Kepler pipeline data processing [1] (CAL) do not currently use all of the information available to identify and correct several instrument-induced artifacts. These include time-varying crosstalk from the fine guidance sensor (FGS) clock signals, and manifestations of drifting moire pattern as locally correlated nonstationary noise, and rolling bands in the images which find their way into the time series [2], [3]. As the Kepler Mission continues to improve the fidelity of its science data products, we are evaluating the benefits of adding pipeline steps to more completely model and dynamically correct the FGS crosstalk, then use the residuals from these model fits to detect and flag spatial regions and time intervals of strong time-varying black-level which may complicate later processing or lead to misinterpretation of instrument behavior as stellar activity.

Kolodziejczak, J. J.↗

Dynamic Black-Level Correction and Artifact Flagging in the Kepler Data Pipeline

Instrument-induced artifacts in the raw Kepler pixel data include time-varying crosstalk from the fine guidance sensor (FGS) clock signals, manifestations of drifting moiré pattern as locally correlated nonstationary noise and rolling bands in the images which find their way into the calibrated pixel time series and ultimately into the calibrated target flux time series. Using a combination of raw science pixel data, full frame images, reverse-clocked pixel data and ancillary temperature data the Keplerpipeline models and removes the FGS crosstalk artifacts by dynamically adjusting the black level correction. By examining the residuals to the model fits, the pipeline detects and flags spatial regions and time intervals of strong time-varying blacklevel (rolling bands ) on a per row per cadence basis. These flags are made available to downstream users of the data since the uncorrected rolling band artifacts could complicate processing or lead to misinterpretation of instrument behavior as stellar. This model fitting and artifact flagging is performed within the new stand-alone pipeline model called Dynablack. We discuss the implementation of Dynablack in the Kepler data pipeline and present results regarding the improvement in calibrated pixels and the expected improvement in cotrending performances as a result of including FGS corrections in the calibration. We also discuss the effectiveness of the rolling band flagging for downstream users and illustrate with some affected light curves.

Clarke, B. D.↗

Multi-Objective Multi-User Scheduling for Space Science Missions

We have developed an architecture called MUSE (Multi-User Scheduling Environment) to enable the integration of multi-objective evolutionary algorithms with existing domain planning and scheduling tools. Our approach is intended to make it possible to re-use existing software, while obtaining the advantages of multi-objective optimization algorithms. This approach enables multiple participants to actively engage in the optimization process, each representing one or more objectives in the optimization problem. As initial applications, we apply our approach to scheduling the James Webb Space Telescope, where three objectives are modeled: minimizing wasted time, minimizing the number of observations that miss their last planning opportunity in a year, and minimizing the (vector) build up of angular momentum that would necessitate the use of mission critical propellant to dump the momentum. As a second application area, we model aspects of the Cassini science planning process, including the trade-off between collecting data (subject to onboard recorder capacity) and transmitting saved data to Earth. A third mission application is that of scheduling the Cluster 4-spacecraft constellation plasma experiment. In this paper we describe our overall architecture and our adaptations for these different application domains. We also describe our plans for applying this approach to other science mission planning and scheduling problems in the future.

science planning↗

Treating gridded geospatial data as point data to simplify analytics

Gridded geospatial remote sensing (satellite) data has traditionally been stored in file-based multidimensional arrays to preserve the locality of data. Measurements from locations that are physically next to each other on earth remain next to each other in the arrays. Maintaining this locality is useful when running calculations like reprojection, but unnecessary for many other calculations. This talk will go through a real world example of a tool redesign at the Goddard Earth Sciences Data and Information Services Center (GES DISC), showing the advantages of using the data frame model for calculating summary statistics, where measurement proximity is unimportant.

Analysis-ready data↗

MEDLI2: MISP Inferred Aerothermal Environment and Flow Transition Assessment

The Mars Entry, Descent, and Landing Instrumentation 2 (MEDLI2) sensor suite on the Mars2020 mission contained multiple sensors on the aeroshell to measure the aerothermal environment during entry into the Martian atmosphere. These sensors performed superbly and successfully returned forebody and aftbody heating measurements. Analysis of MEDLI2 data indicated flow transitioning from a laminar to turbulent state on the heatshield. No evidence of flow transition was observed on the backshell. Two methods were used to estimate flow transition times on the heatshield: (1) temperature gradient of near-surface thermocouple data and (2) heat flux gradient from an inverse reconstruction approach using thermocouple data and material response modeling. Both methods produced similar transition times with an estimated accuracy of ±1 s. To assess various transition criteria, transition parameters were evaluated at each sensor location using flow field solutions from computational fluid dynamics (CFD) simulations. The idea was to use conservative values inferred from MEDLI2 data as transition criteria for other Mars missions. To test this hypothesis, MEDLI data from the Mars Science Laboratory (MSL) mission was used to compare predicted vs. actual flow transition times. The comparisons suggest smooth wall transition criteria are not well-suited in modeling the rapid progression of a turbulent transition front. Transition criteria containing a roughness element parameter agreed better with the flight data. In summary, critical transition values derived from MEDLI2 data may be used as a starting point in constructing a flow transition model for future Mars missions.

Chun Y Tang↗

32 examples of LLM applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery

Abstract Large language models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 32 total projects developed during the second annual LLM hackathon for applications in materials science and chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

Computer Science↗

Legacy Survey of Space and Time Data Preview 2: object_scarlet_models dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the object_scarlet_models dataset type. These are scarlet deblender models for detected sources in the deep coadds. This release contains 195,366 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software

Scientific software is essential to scientific innovation and in many ways it is distinct from other types of software. Abandoned (or unmaintained), buggy, and hard to use software, a perception often associated with scientific software can hinder scientific progress, yet, in contrast to other types of software, its longevity is poorly understood. Existing data curation efforts are fragmented by science domain and/or are small in scale and lack key attributes. We use large language models to classify public software repositories in World of Code into distinct scientific domains and layers of the software stack, curating a large and diverse collection of over 18,000 scientific software projects. Using this data, we estimate survival models to understand how the domain, infrastructural layer, and other attributes of scientific software affect its longevity. We further obtain a matched sample of non-scientific software repositories and investigate the differences. We find that infrastructural layers, downstream dependencies, mentions of publications, and participants from government are associated with a longer lifespan, while newer projects with participants from academia had shorter lifespan. Against common expectations, scientific projects have a longer lifetime than matched non-scientific open-source software projects. We expect our curated attribute-rich collection to support future research on scientific software and provide insights that may help extend longevity of both scientific and other projects.

Malviya Thakur, Addi [ORNL] (ORCID:000000022681999↗

An Earth System Digital Twin for Flood Prediction and Analysis

An Earth System Digital Twin (ESDT) is a dynamic, interactive, digital replica of the state and temporal evolution of Earth systems. It integrates multiple models along with observation data, and connecting them with analysis, AI, and visualization tools. Together, these enable users to explore the current state of the Earth system, predict future conditions, and run hypothetical scenarios to understand how the system would evolve under various assumptions. The NASA’s Advanced Information Systems Technology (AIST)’s Integrated Digital Earth Analysis System (IDEAS) project is to establish an extensible architectural solution to develop digital twins of our physical environment for Earth Science. IDEAS delivers a formal system architecture with mechanisms for the outputs of one model to feed into others; for driving models with observation data; and for harmonizing observation data and model outputs for analysis. To validate and demonstrate the IDEAS architecture, this project collaborates with the Space Climate Observatory (SCO)’s FloodDAM project and the Centre National d’Etudes Spatiales (CNES) to focus on floods detection, prediction and their impacts.

Kettig, Peter↗

NASA Observations and Modeling During ICE-POP

Recap: NASA-Specific Objectives for ICE-POP: Provide real-time observational and NWP data in support of ICE-POP, participate in significant international science effort; GPM (Global Precipitation Measurement) Ground Validation and NASA Weather Program -Direct/physical validation of active/passive satellite-based snowfall retrieval algorithms over coastline and mountains; melting layer interaction with terrain -Physics of snow, coupling to snow water equivalent rate and satellite remote sensor retrieval algorithm assumptions - -Size distributions, types/habit, water equivalent, profiles -NU-WRF (NASA-Unified Weather Research and Forecasting) Model plus Observational analyses: Movement toward “level IV products” leverage intensive and multi-faceted NWP (Numerical Weather Prediction) component -Model precipitation processes (liquid, mixed phase and frozen); Build model testing database for further active/passive remote sensing algorithm development (e.g., satellite data simulators) -"Integrated" validation of products in operational context.

Precipitation Science↗