Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Preface to the Special Issue on Modeling and Data Analysis Methods for the SMILE mission

The SMILE (Solar wind Magnetosphere Ionosphere Link Explorer) project (http://www.nssc.cas.cn/smile/, https://www.cosmos.esa.int/web/smile/mission) is a joint spacecraft mission of the European Space Agency (ESA) and the Chinese Academy of Sciences (CAS) with an expected launch in 2025. SMILE aims to study the global interactions of solar wind–magnetosphere–ionosphere innovatively by imaging the Earth’s magnetosheath and cusps in soft X-rays and the northern auroral region in ultraviolet (UV) while simultaneously measuring plasma and magnetic field parameters in the solar wind and magnetosheath along a highly-elliptical and highly-inclined orbit. This special issue is composed of 22 articles, presenting recent progress in modeling and data analysis techniques developed for the SMILE mission. In this preface, we categorize the articles into the following seven topics and provide brief summaries: (1) instrument descriptions of the Soft X-ray Imager (SXI), (2) numerical modeling of the X-ray signals, (3) data processing of the X-ray images, (4) boundary tracing methods from the simulated images, (5) physical phenomena and a mission concept related to the scientific goals of SMILE-SXI, (6) studies of the aurora, and (7) ground-based support for SMILE.

SMILE↗

Evolution of the Earth Observing System (EOS) Data and Information System (EOSDIS)

One of the strategic goals of the U.S. National Aeronautics and Space Administration (NASA) is to "Develop a balanced overall program of science, exploration, and aeronautics consistent with the redirection of the human spaceflight program to focus on exploration". An important sub-goal of this goal is to "Study Earth from space to advance scientific understanding and meet societal needs." NASA meets this subgoal in partnership with other U.S. agencies and international organizations through its Earth science program. A major component of NASA s Earth science program is the Earth Observing System (EOS). The EOS program was started in 1990 with the primary purpose of modeling global climate change. This program consists of a set of space-borne instruments, science teams, and a data system. The instruments are designed to obtain highly accurate, frequent and global measurements of geophysical properties of land, oceans and atmosphere. The science teams are responsible for designing the instruments as well as scientific algorithms to derive information from the instrument measurements. The data system, called the EOS Data and Information System (EOSDIS), produces data products using those algorithms as well as archives and distributes such products. The first of the EOS instruments were launched in November 1997 on the Japanese satellite called the Tropical Rainfall Measuring Mission (TRMM) and the last, on the U.S. satellite Aura, were launched in July 2004. The instrument science teams have been active since the inception of the program in 1990 and have participation from Brazil, Canada, France, Japan, Netherlands, United Kingdom and U.S. The development of EOSDIS was initiated in 1990, and this data system has been serving the user community since 1994. The purpose of this chapter is to discuss the history and evolution of EOSDIS since its beginnings to the present and indicate how it continues to evolve into the future. this chapter is organized as follows. Sect. 7.2 provides a discussion of EOSDIS, its elements and their functions. Sect. 7.3 provides details regarding the move towards more distributed systems for supporting both the core and community needs to be served by NASA Earth science data systems. Sect. 7.4 discusses the use of standards and interfaces and their importance in EOSDIS. Sect. 7.5 provides details about the EOSDIS Evolution Study. Sect. 7.6 presents the implementation of the EOSDIS Evolution plan. Sect. 7.7 briefly outlines the progress that the implementation has made towards the 2015 Vision, followed by a summary in Sect. 7.8.

Ramapriyan, Hampapuram K.↗

Wildfire Segmentation From Remotely Sensed Data Using Quantum-Compatible Conditional Vector Quantized-Variational Autoencoders

Wildfires represent a critical environmental hazard with multifaceted implications for ecosystems, communities, and public health [1]. The escalating frequency and intensity of wildfires globally have intensified the urgency for robust segmentation methodologies to facilitate effective mitigation, response, and recovery strategies [2]. Accurate wildfire segmentation is pivotal for delineating fire boundaries, assessing progression patterns, and prioritizing resource allocation during emergency scenarios. Furthermore, precise segmentation enables stakeholders, including policymakers, environmental scientists, and emergency responders, to formulate evidence-based strategies, thereby minimizing socio-economic disruptions and ecological degradation. Consequently, advancing wildfire segmentation techniques through innovative technological interventions remains a paramount research imperative. Although foundational in wildfire segmentation, traditional deterministic models exhibit inherent limitations that compromise their efficacy in dynamic and uncertain environments. These models often operate on rigid algorithms prioritizing deterministic classifications, thereby overlooking the inherent complexities and uncertainties associated with wildfire behavior and satellite data variability. Such deterministic frameworks tend to produce oversimplified representations that fail to capture the intricate nuances of evolving fire dynamics, spatial heterogeneity, and environmental interactions [1]. Consequently, the deterministic approach’s propensity for uncertainty collapsing [1, 3] hampers the accuracy, reliability, and applicability of segmentation outcomes in real-world scenarios. Contrastingly, stochastic models offer a more nuanced and adaptable framework for wildfire segmentation. By integrating probabilistic elements into the modeling paradigm, stochastic approaches, particularly probabilistic approaches such as variational auto encoders (VAEs) [4], facilitate comprehensive uncertainty assessment, enabling researchers to quantify and incorporate uncertainties into segmentation outcomes effectively. This probabilistic nature empowers stochastic models to encapsulate variability, account for data inconsistencies, and adapt to evolving environmental conditions, enhancing segmentation accuracy, reliability, and robustness. Embracing stochastic methodologies thus catalyzes advancements in wildfire science by fostering a more holistic, adaptive, and resilient segmentation framework. Despite VAEs demonstrating significant promise in various applications, they come with inherent limitations that have garnered attention within the machine learning community. One of the primary drawbacks lies in their reliance on static priors, which essentially assume a fixed distribution for latent variables, thereby limiting the model’s flexibility to capture complex data structures effectively [5]. This static nature leads to suboptimal representations, especially when dealing with complex and high-dimensional data. Additionally, VAEs often struggle with generating sharp and realistic samples, a phenomenon commonly referred to as mode collapse [5, 7, 6]. Furthermore, the optimization process in VAEs, which involves balancing the reconstruction loss and the regularization term, can sometimes be challenging to fine-tune [7]. In recent efforts to address these shortcomings, alternative approaches like Vector Quantized Variational Auto encoders(VQ-VAEs) [7], address the challenges by incorporating discrete latent variables and leveraging techniques that enhance the quality and diversity of generated samples while maintaining efficient training dynamics. VQ-VAEs propose a dynamic prior distribution generation mechanism that diverges from the static priors commonly associated with traditional VAEs. This dynamic approach allows for more adaptive and context-aware latent variable representations, thereby potentially capturing complex data structures more effectively. Unlike autoregressive prior models such as PixelCNN, which, despite their ability to model dependencies across data dimensions, suffer from significant computational inefficiencies and lack flexibility in handling diverse datasets. In our work, we propose to use a generative quantum-compatible approach to help alleviate the shortcomings of autoregressive prior model in VQ-VAEs. Restricted Boltzmann Machines (RBMs) are a viable alternative prior model that can learn prior distributions in a faster and more flexible manner. In this research endeavor, we meticulously curate a state-of-the-art dataset leveraging satellite MODIS data in conjunction with VIIRS fire masks, derived from Fire Radiative Power (FRP), thereby encapsulating diverse wildfire scenarios and environmental contexts. We developed a conditional VQ-VAE architecture with the RBM prior model that is trained in a supervised manner for segmenting wildfire masks. This innovative approach synergistically harnesses deep learning capabilities, enabling the generation of segmentation maps characterized by heightened precision, granularity, and contextual relevance. Furthermore, replacing the autoregressive prior learning method proposed by the original VQ-VAE with a prior density approximation via quantum-compatible RBM facilitates expedited inference processes, augments flexibility in prior sampling, optimizes computational efficiency and establishes a groundbreaking benchmark in wildfire segmentation methodologies.

quantum machine learning↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A new data-driven map predicts substantial undocumented peatland areas in Amazonia

Tropical peatlands are among the most carbon-dense terrestrial ecosystems yet recorded. Collectively, they comprise a large but highly uncertain reservoir of the global carbon cycle, with wide-ranging estimates of their global area (441 025–1700 000 km 2 ) and below-ground carbon storage (105–288 Pg C). Substantial gaps remain in our understanding of peatland distribution in some key regions, including most of tropical South America. Here we compile 2413 ground reference points in and around Amazonian peatlands and use them alongside a stack of remote sensing products in a random forest model to generate the first field-data-driven model of peatland distribution across the Amazon basin. Our model predicts a total Amazonian peatland extent of 251 015 km 2 (95th percentile confidence interval: 128 671–373 359), greater than that of the Congo basin, but around 30% smaller than a recent model-derived estimate of peatland area across Amazonia. The model performs relatively well against point observations but spatial gaps in the ground reference dataset mean that model uncertainty remains high, particularly in parts of Brazil and Bolivia. For example, we predict significant peatland areas in northern Peru with relatively high confidence, while peatland areas in the Rio Negro basin and adjacent south-western Orinoco basin which have previously been predicted to hold Campinarana or white sand forests, are predicted with greater uncertainty. Similarly, we predict large areas of peatlands in Bolivia, surprisingly given the strong climatic seasonality found over most of the country. Very little field data exists with which to quantitatively assess the accuracy of our map in these regions. Data gaps such as these should be a high priority for new field sampling. This new map can facilitate future research into the vulnerability of peatlands to climate change and anthropogenic impacts, which is likely to vary spatially across the Amazon basin.

54 ENVIRONMENTAL SCIENCES↗

Activities of the Remote Sensing Information Sciences Research Group

Topics on the analysis and processing of remotely sensed data in the areas of vegetation analysis and modelling, georeferenced information systems, machine assisted information extraction from image data, and artificial intelligence are investigated. Discussions on support field data and specific applications of the proposed technologies are also included.

John E Estes↗

US Participation in the GOME and SCIAMACHY Projects

This report summarizes research done under NASA Grant NAGW-2541 through September 30, 1997. The research performed under this grant includes development and maintenance of scientific software for the GOME retrieval algorithms, consultation on operational software development for GOME, sensitivity and instrument studies to define GOME and SCIAMACHY instruments, consultation on optical and detector issues for both GOME and SCIAMACHY, consultation and development for SCIAMACHY near-real-time (NRT) and off-line (OL) data products, and development of infrared line-by-line atmospheric modeling and retrieval capability for SCIAMACHY. The European Space Agency selected the SAO to participate in GOME validation and science studies, part of the overall ERS AO. This provided access to all GOME data; The SAO activities that are carried out as a result of selection by ESA were funded by the present grant. The Global Ozone Monitoring Experiment was successfully launched on the ERS- 2 satellite on April 20, 1995, and remains working in normal fashion. SCIAMACHY is currently scheduled for launch in early 2000. The first two European ozone monitoring instruments (OMI), to fly on the q series of operational meteorological satellites being planned by Eumetsat, have been selected to be GOME-type instruments (the first, in fact, will be the refurbished GOME flight spare). K. Chance is the U.S. member of the OMI Users Advisory Group.

Chance, K. V.↗

Testing the Pairs-Reflection Model with X-Ray Spectral Variability and X-Ray Properties of Complete Samples of Radio-Selected BL Lacertae Objects

This grant was awarded to Dr. C. Megan Urry of the Space Telescope Science Institute in response to two successful ADP proposals to use archival Ginga and Rosat X-ray data for 'Testing the Pairs-Reflection model with X-Ray Spectral Variability' (in collaboration with Paola Grandi, now at the University of Rome) and 'X-Ray Properties of Complete Samples of Radio-Selected BL Lacertae Objects' (in collaboration with then-graduate student Rita Sambruna, now a post-doc at Goddard Space Flight Center). In addition, post-docs Joseph Pesce and Elena Pian, and graduate student Matthew O'Dowd, have worked on several aspects of these projects. The grant was originally awarded on 3/01/94; this report covers the full period, through May 1997. We have completed our project on the X-ray properties of radio-selected BL Lacs.

Urry, C. Megan↗

Spectroscopic comparison of effects of electron radiation on mechanical properties of two polyimides

The differences in the radiation durabilities of two polyimide materials, Du Pont Kapton and General Electric Ultem, are compared. An explanation of the basic mechanisms which occur during exposure to electron radiation from analyses of infrared (IR) and electron paramagnetic resonance (EPR) spectroscopic data for each material is provided. The molecular model for Kapton was, in part, established from earlier modeling for Ultem (pp. 1293-1298 of IEEE Transactions on Nuclear Science, December 1984). Techniques for understanding the durability of one complex polymer based on the understanding of a different and equally complex polymer are demonstrated. The spectroscopic data showed that the primary radiation-generated change in the tensile properties of Ultem (a large reduction in tensile elongation) was due to crosslinking, which followed the capture by phenyl radicals of hydrogen atoms removed from gem-dimethyl groups. In contrast, the tensile properties of Kapton remained unchanged because radical-radical recombination, a self-mending process, took place.

Long, Edward R., Jr.↗

New constraint on the Np 237 ( n , γ ) Np 238 integral cross section using the Godiva-IV critical assembly

Accurate knowledge of the 237 Np(n, γ) 238 Np cross section at fast neutron energies is important for applied nuclear science. The presently available experimental data has large disagreements in the fast neutron region. Perform a model-independent measurement of the 237 Np(n, γ) 238 Np integral cross section using a well characterized fast neutron source and compare the result with previous measurements and current nuclear data evaluations. Provide an integral measurement that can be used as a benchmark for current evaluations. Multiple samples of 237 Np were irradiated in the Godiva-IV critical assembly. Following the irradiation, the samples placed in a γ-ray counting setup and the γ-rays emitted from the decay of 238 Np were measured over a time period of approximately 7 days. Multiple γ-ray decay branches of 238 Np were observed. The observed activity of 238 Np was used to calculate the amount of 238 Np produced during the irradiation via the 237 Np(n, γ) 238 Np reaction and an integral cross section of 342(11) mb was measured for the Godiva-IV neutron spectrum. Further, the 238 Np half-life has been measured with a result of 50.31(5) hours. The 237 Np(n, γ) 238 Np integral cross section measured in this work is in agreement with overlapping 1σ error bands to ENDF/B-VIII.0. However, the measured value is 3σ away from the calculated integral cross section using JENDL-5. This measurement offers a reliable benchmark for future 237 Np(n, γ) 238 Np cross section evaluations.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Utilizing Mars Global Reference Atmospheric Model (Mars-GRAM 2005) to Evaluate Entry Probe Mission Sites

The Mars Global Reference Atmospheric Model (Mars-GRAM 2005) is an engineering-level atmospheric model widely used for diverse mission applications. An overview is presented of Mars-GRAM 2005 and its new features. The "auxiliary profile" option is one new feature of Mars-GRAM 2005. This option uses an input file of temperature and density versus altitude to replace the mean atmospheric values from Mars-GRAM's conventional (General Circulation Model) climatology. Any source of data or alternate model output can be used to generate an auxiliary profile. Auxiliary profiles for this study were produced from mesoscale model output (Southwest Research Institute's Mars Regional Atmospheric Modeling System (MRAMS) model and Oregon State University's Mars mesoscale model (MMM5) model) and a global Thermal Emission Spectrometer (TES) database. The global TES database has been specifically generated for purposes of making Mars-GRAM auxiliary profiles. This data base contains averages and standard deviations of temperature, density, and thermal wind components, averaged over 5-by-5 degree latitude-longitude bins and 15 degree Ls bins, for each of three Mars years of TES nadir data. The Mars Science Laboratory (MSL) sites are used as a sample of how Mars-GRAM' could be a valuable tool for planning of future Mars entry probe missions. Results are presented using auxiliary profiles produced from the mesoscale model output and TES observed data for candidate MSL landing sites. Input parameters rpscale (for density perturbations) and rwscale (for wind perturbations) can be used to "recalibrate" Mars-GRAM perturbation magnitudes to better replicate observed or mesoscale model variability.

Justh, Hilary L.↗

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science [Slides]

Machine intelligence has the potential to revolutionize materials science, enabling autonomous synthesis, self-driving characterization, and accelerated modeling. However, despite the promise, successful implementation of these methods in day-to-day research remains a challenge. This talk will delve into the reasons behind this, exploring how truly intelligent experiments are hindered by opaque experiment control, a lack of domain-specific models, and human-centric design. Through a focus on the characterization of next-generation microelectronics and energy storage materials, I will share insights from both successful and failed attempts to implement machine intelligence. We will then explore the next steps necessary to unlock the full potential of machine intelligence in materials science, creating a future where intelligent systems work seamlessly alongside researchers to drive innovation and discovery.

36 MATERIALS SCIENCE↗

CHUWD-H v1.0: a comprehensive historical hourly weather database for U.S. urban energy system modeling

Reliable and continuous meteorological data are crucial for modeling the responses of energy systems and their components to weather and climate conditions, particularly in densely populated urban areas. However, existing long-term datasets often suffer from spatial and temporal gaps and inconsistencies, posing great challenges for detailed urban energy system modeling and cross-city comparison under realistic weather conditions. Here we introduce the Historical Comprehensive Hourly Urban Weather Database (CHUWD-H) v1.0, a 23-year (1998-2020) gap-free and quality-controlled hourly weather dataset covering 550 weather station locations across all urban areas in the contiguous United States. CHUWD-H v1.0 synthesizes hourly weather observations from stations with outputs from a physics-based solar radiation model and a reanalysis dataset through a multi-step gap filling approach. A 10-fold Monte Carlo cross-validation suggests that the accuracy of this gap filling approach surpasses that of conventional gap filling methods. Designed primarily for urban energy system modeling, CHUWD-H v1.0 should also support historical urban meteorological and climate studies, including the validation and evaluation of urban climate modeling.

54 ENVIRONMENTAL SCIENCES↗

Hacking Kilometer-Scale Models: A Participative Model for Climate Information

In May 2025, nearly 700 participants from all around the world coalesced at 10 regional nodes and a few satellite nodes to take part in a global hackathon of kilometer-scale (horizontal grid spacing < 10 km) regional and global Earth system models. Exciting science is emerging from these efforts, ranging across novel model analysis, new ways of integrating with satellite data, and emulation with machine learning. New technologies were trialed that enable the community to work in new and complementary ways to democratize access to global information at a local scale from a set of the world’s highest-resolution climate models. The hackathon demonstrated how exascale data can be organized to be accessible to anyone. Fundamentally, the community could apply these techniques and technologies to move toward more participative models for coproduction and delivery of diverse sources of climate information for climate scientists and citizens alike.

Climate models↗

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Data Management as a Cluster Middleware Centerpiece

Through earth and space modeling and the ongoing launches of satellites to gather data, NASA has become one of the largest producers of data in the world. These large data sets necessitated the creation of a Data Management System (DMS) to assist both the users and the administrators of the data. Halcyon Systems Inc. was contracted by the NASA Center for Computational Sciences (NCCS) to produce a Data Management System. The prototype of the DMS was produced by Halcyon Systems Inc. (Halcyon) for the Global Modeling and Assimilation Office (GMAO). The system, which was implemented and deployed within a relatively short period of time, has proven to be highly reliable and deployable. Following the prototype deployment, Halcyon was contacted by the NCCS to produce a production DMS version for their user community. The system is composed of several existing open source or government-sponsored components such as the San Diego Supercomputer Center s (SDSC) Storage Resource Broker (SRB), the Distributed Oceanographic Data System (DODS), and other components. Since Data Management is one of the foremost problems in cluster computing, the final package not only extends its capabilities as a Data Management System, but also to a cluster management system. This Cluster/Data Management System (CDMS) can be envisioned as the integration of existing packages.

Zero, Jose↗

Machine Learning Lifecycle for Earth Science Application: A Practical Insight into Production Deployment

Earth science domain presents unique sets of problems that are increasingly being solved using data driven approaches. The availability of big Earth science data offers immense potential for Machine learning (ML) as evident from numerous research publications lately. However, many of these publications are not ending up as production applications mainly because the data scientists who develop the ML models are now expected to complete the ML lifecycle by deploying and scaling the models in production. We introduce ML lifecycle to the Earth science community including the opportunities and challenges that lie ahead in each phase of the lifecycle. We demonstrate the lifecycle using an Earth science problem that we used ML to address and transitioned to production.

Maskey, Manil↗