Search NASA⌕ Search

SEARCH · Search NASA

Results for “data platform”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗

Portable Software Environment for Ultrahigh-Resolution ELM Development on GPUs

This paper presents our endeavors in developing the large-scale, ultra-high-resolution E3SM Land Model (uELM), specifically designed for exascale computers furnished with accelerators such as Nvidia GPUs. The uELM is a sophisticated code that substantially relies on High-Performance Computing (HPC) environments, necessitating particular machine and software configurations. To facilitate community-based uELM developments employing GPUs, we have created a portable, standalone software environment preconfigured with uELM input datasets, simulation cases, and source code. This environment, utilizing Docker, encompasses all essential code, libraries, and system software for uELM development on GPUs. It also features a functional unit test framework and an offline model testbed for comprehensive numerical experiments. From a technical perspective, the paper discusses GPU-ready container generations, uELM code management, and input data distribution across computational platforms. Lastly, the paper demonstrates the use of environment for functional unit testing, end-to-end simulation on CPUs and GPUs, and collaborative code development.

E3SM Land Model↗

RTN-045: Guidelines for User Tutorials

This document defines the guidelines, principles, and formats for user-facing tutorials that demonstrate how to use the Rubin Science Platform (RSP) to analyze data from the Legacy Survey of Space and Time (LSST). All Rubin staff and the broader science community should use these guidelines when contributing to the sets of Jupyter Notebook or documentation-based tutorials maintained by the Rubin Community Science team (CST).

79 ASTRONOMY AND ASTROPHYSICS↗

Pixel-Registered Multimodal Synchrotron XRF and FTIR Microscopies Reveal Salinity Stress Response Mechanisms in Pistachio

Background: Salinity is a major abiotic stress that negatively affects nearly all plant species at all stages of growth. Drought and poor-quality irrigation cause high soil salinity and salt accumulation via evaporation, reducing crop productivity. Despite its critical importance, the spatial localization of salt ions and associated biochemical changes within plants experiencing high salinity remains largely unknown. In this study, we developed a multimodal imaging pipeline to understand the impact of salinity on the pistachio rootstock UCB-1 (Pistacia atlantica x Pistacia integerrima). We directly link biochemical fingerprints in stem tissue architecture with salt ion localization to provide insights into the strategies pistachio uses to tolerate salinity. Results: We observed that Pistacia spp. exposed to high salt conditions accumulated Ca, Si, Cl, Al and Mg as hotspots within the pith, compared to the control (of which only Ca and Al co-locate). In contrast, there was a decrease in K between the control and salinity treatment. Hotspots of amide I and II were present in the cortex and pith of the salinity treated sample. Additionally, the salinity treatment resulted in an increased abundance of pectin and carbohydrates within the pith compared to the control, and the abundance of esters/carboxylic acid was greater in the salinity treatment. Conclusions: We determined that Cl and K, S and P, and biochemical components polysaccharide and pectin, esters and carboxylic acid, amide I and cellulose are the strongest drivers of salinity- treatment induced variability. In the cortex and phloem/xylem, a negative K-Ca correlation decreases in the salinity treatment. Several hotspots of elements and amide I (proteins) appear under salinity treatment, particularly in the cortex, suggesting an increase in the production of stress-related proteins (in response to high Cl) and/or structural proteins (i.e. Ca). Together, these results indicate that pistachio responds to salinity through ion compartmentalization coupled with a targeted biochemical adjustment, rather than a broadscale tissue-wide response. Overall, these novel, spatially resolved pixel-registered multimodal imaging data provide an enabling platform to understand the mechanisms of salinity tolerance in Pistacia spp and can be broadly applied to studying stress-related phenotype response in various plant tissues.

FTIR spectromicroscopy↗

Can protein expression be ‘solved’?

Recombinant protein expression is central to biotechnology’s application in academic exploration as well as human health, climate applications and the bioeconomy in general. However, not all proteins can be expressed in all organisms, and the field lacks a predictive model of soluble protein overexpression that could replace laborious experimental trial-and-error. Here, we discuss the state of the field and identify the lack of large, high-fidelity datasets as the primary bottleneck to progress. We review possible assays that could be used for data collection to identify a path toward an extensible experimental platform for collecting soluble recombinant protein overexpression data across organisms. We suggest that the resulting dataset should be used to train increasingly generalizable predictive models of protein expression to answer the question: “How can predictive protein expression be solved?”.

59 BASIC BIOLOGICAL SCIENCES↗

Distributed Wind Monitoring Best Practices

Accessible performance and operational data have been identified as a key enabler for distributed wind energy industry advancement. While utility-scale wind turbines benefit from reliable and continuous supervisory control and data acquisition (SCADA)-based monitoring platforms, monitoring of the U.S. fleet of distributed wind (DW) turbines has been more inconsistent, unreliable, and sometime difficult to access. Without fleet monitoring data, the industry will never understand and thus work to improve turbine under-performance and reliability issues. For the DW industry to scale up, attract investors, and boost credibility, fleetwide monitoring must be robust and reliable, select data must be made accessible to stakeholders, and the data must be in a format useful to users. To help move the industry toward a more standardized, accessible stream of monitoring data, this distributed wind monitoring best practices report attempts to cover topics including key monitoring channels, hardware, communication strategies, and accessibility. Strategic engagement with DW original equipment manufacturers (OEMs), service providers, lab and university researchers, testing organization, certification bodies, end users and solar photovoltaic (PV) monitoring experts has enabled a better understanding of the current state-of-the-art of monitoring and aided in articulating this set of best practices that will guide OEMs toward harmonized monitoring strategies, aimed at a future goal of achieving accessible performance and operational data for the entire fleet of U.S. distributed wind turbines.

17 WIND ENERGY↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

A semantics-driven framework to enable demand flexibility control applications in real buildings

Decarbonising and digitalising the energy sector requires scalable and interoperable Demand Flexibility (DF) applications. Semantic models are promising technologies for achieving these goals, but existing studies focused on DF applications exhibit limitations. These include dependence on bespoke ontologies, lack of computational methods to generate semantic models, ineffective temporal data management and absence of platforms that use these models to easily develop, configure and deploy controls in real buildings. This paper introduces a semantics-driven framework to enable DF control applications in real buildings. The framework supports the generation of semantic models that adhere to Brick and SAREF while using metadata from Building Information Models (BIM) and Building Automation Systems (BAS). The work also introduces a web platform that leverages these models and an actor and microservices architecture to streamline the development, configuration and deployment of DF controls. The paper demonstrates the framework through a case study, illustrating its ability to integrate diverse data sources, execute DF actuation in a real building, and promote modularity for easy reuse, extension, and customisation of applications. The paper also discusses the alignment between Brick and SAREF, the value of leveraging BIM data sources, and the framework's benefits over existing approaches, demonstrating a 75% reduction in effort for developing, configuring, and deploying building controls.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Oak Ridge National Laboratory EAGLE-I TM : Modeling Electric Utility County Customers for Situational Awareness

During natural hazard events (hurricanes, wildfires, earthquakes, etc.) and recent man-made events (e.g., cyber attacks), the exchange of near real-time, spatially refined data within the response community is critical. The EAGLE-I$^{TM}$ platform is one tool that facilitates this data for decision makers within the energy sector. While much information can be collected and integrated into the system directly, other pertinent data must be augmented by other derived data products to enhance the information and allow for a consistent evaluation of on-the-ground conditions. One such data set that requires the addition of other derived data is the electric utility customer outage data that is aggregated to the county level within the EAGLE-I application. Without a county customer data set, outages can only be compared on total counts, which gives greater importance to higher population outages. Including an electric utility customer data set at the county level allows for these outage counts to be converted to percent outages and brings a consistent classification of outages and equal importance to all outages. To achieve this, several available data sets were combined and spatial disaggregation techniques were employed to model customer estimates at the county scale. This paper presents the approach to produce this data for the United States and lessons learned from working with these disparate data sets. Data validation is provided, where possible, and limitations of the model and possible improvements are discussed.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Vera C. Rubin Observatory Data Preview 1

We present Rubin Data Preview 1 (DP1), the first data from the National Science Foundation–Department of Energy Vera C. Rubin Observatory, comprising raw and calibrated single-epoch images, coadds, difference images, detection catalogs, and ancillary data products. DP1 is based on 1792 optical–near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera (LSSTComCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile in late 2024. DP1 covers ∼15 deg 2 distributed across seven roughly equal-sized noncontiguous fields, each independently observed in six broad photometric bands, ugrizy. The median FWHM of the point-spread function across all bands is approximately 1"14, with the sharpest images reaching about 0." 58. The 5σ point-source depths for coadded images in the deepest field, the Extended Chandra Deep Field South, are u = 24.55, g = 26.18, r = 25.96, i = 25.71, z = 25.07, and y = 23.1. Other fields are no more than 2.2 mag shallower in any band, where they have nonzero coverage. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band in coadds, and 431 solar system objects, of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and is available to Vera C. Rubin Observatory data rights holders via the Rubin Science Platform, a cloud-based environment for the analysis of petascale astronomical data. While small compared to future LSST releases, its high quality and diversity of data support a broad range of early science investigations ahead of full operations in 2026.

Ground-based astronomy↗

RTN-095: The Vera C. Rubin Observatory Data Preview 1

We present Rubin Data Preview 1 (DP1), the first release of data from the NSF-DOE Vera C. Rubin Observatory, consisting of raw and calibrated single-epoch images, coadds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of ~15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. The median image quality across all bands, measured by the FWHM of the point-spread function, is approximately 1.13 arcseconds, with the sharpest images reaching about 0.65 arcseconds. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band, and 431 solar system objects, of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and available to Rubin data rights holders via the Rubin Science Platform, a cloud-based environment for the analysis of petascale astronomical data. While small compared to future LSST releases, its high quality and diversity of data support a broad range of early science investigations across all four LSST themes, providing a valuable opportunity to engage with Rubin data ahead of the start of full operations in late 2025.

79 ASTRONOMY AND ASTROPHYSICS↗

A digital twin platform for building performance monitoring and optimization: Performance simulation and case studies

Advancements in sensor technology, data analytics, affordable compute, and communication infrastructure have paved the way for Digital Twin technology in optimizing building operations and controls. This study presents the development of an open and interoperable web-based Digital Twin platform for integrating diverse data streams and facilitating effective user interactions. The platform utilizes modern technologies for the web framework and time-series data management, ensuring scalability and responsiveness. The backend supports seamless integration of diverse data sources and emulators, incorporating data from building sensors and meters, external weather Application Programming Interfaces, and advanced EnergyPlus simulation models of the building and its energy systems including the Distributed Energy Resources that are formulated in Functional Mockup Units. A simulation case study was conducted with FlexLab, a test facility on Lawrence Berkeley National Laboratory campus. The case study includes normal operations, Distributed Energy Resource integration, and power outage scenarios, to illustrate the Digital Twin’s ability to provide critical insights into energy performance and thermal resilience. The results demonstrated the platform’s potential as a decision-support tool for optimizing building energy performance and enhancing resilience against extreme weather events. Future work will focus on deploying the Digital Twin platform to a real building for field validation, extending its capabilities to cover more scenarios such as bidirectional Electric Vehicle interactions, and enhancing user engagement.

EnergyPlus↗

An analysis of alcohol compression ignition strategies on different engine platforms

Strategies that enable ignition of the low cetane alcohol fuels (methanol and ethanol) include intake heating (direct or via hot residual trapping), increased compression ratio (CR), multi-injection strategies, ignition enhancers, and active prechambers. The literature consistently shows alcohol compression ignition (CI) is more efficient than diesel at high loads irrespective of ignition strategy but shows contradictory results at low loads. To understand this, this work combines data from four different engine platforms ranging from a displaced volume of 0.4 to 2.1 L/cyl. using different alcohol CI ignition strategies. Using modeling to support experimental analysis, alcohol-specific efficiency penalties are highlighted to inform alcohol combustion system development. The alcohols’ lower energy density means more fuel must be injected near top dead center to achieve a given load, increasing the relative penalty of sensible/latent heating. Pilot injections help reduce this penalty. Intake heating incurs both a heat transfer penalty and a thermodynamic penalty due to a reduced ratio of specific heats during compression. A strategy that combines elevated CR, large pilot injections, and ignition enhancers can avoid this penalty. Alternatively, a prechamber strategy can achieve this without changing the CR, nor needing ignition enhancers in the fuel. At low loads, diesel burns in a partially premixed mode characterized by high burn rates and low heat losses, reducing alcohol-specific combustion benefits. Alcohol CI on the light-duty platform showed abnormally high heat transfer over diesel compared to other platforms, favoring a more premixed/partially premixed strategy despite a combustion efficiency tradeoff.

33 ADVANCED PROPULSION SYSTEMS↗

Integrating Data Centers and Grid Technologies at Scale

This presentation focuses on the challenge of integrating AI-driven data centers with the power grid at scale. It examines the AI data center capacity challenge and the role of new Medium Voltage Direct Current (MVDC) and other grid-enhancing technologies in enabling efficient and reliable power delivery. The session will highlight the National Laboratory of the Rockies' ARIES capabilities and planning tools, along with collaborative examples involving Verrus, Compass, and Schneider through the Agora test bed for grid-friendly data center evaluations, and ON. Energy for UPS evaluation. It will showcase the NLR Stable Grid Platform for studying oscillations caused by large-scale data centers, along with planning tools to assess grid security and reliability. Additionally, the presentation covers reconductoring strategies to increase grid capacity and explores innovative data center architectures, including the Advanced DC Architectures with Power-electronic Transformers (ADAPT) platform, which enables testing of complete DC architectures for data centers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Shock equation of state experiments in MgO up to 1.5 TPa and the effects of optical depth on temperature determination

Laser-driven shock compression enables an experimental study of phase transitions at unprecedented pressures and temperatures. One example is the shock Hugoniot of magnesium oxide (MgO), which crosses the B1–B2-liquid triple point at 400–600 GPa, 10 000–13 000 K (0.86–1.12 eV). MgO is a major component within the mantles of terrestrial planets and has long been a focus of high-pressure research. Here, we combine time-resolved velocimetry and pyrometry measurements with a decaying shock platform to obtain pressure–temperature data on MgO from 300 to 1500 GPa and 9000 to 50 000 K. Pressure–temperature–density Hugoniot data are reported at 1500 GPa. These data represent the near-instantaneous response of an MgO [100] single crystal to shock compression. We report on a prominent temperature anomaly between 400 and 460 GPa, in general agreement with previous shock studies, and draw comparison with equation-of-state models. We provide a detailed analysis of the decaying shock compression platform, including a treatment of a pressure-dependent optical depth near the shock front. We show that if the optical depth of the shocked material is larger than 1 μm, treating the shock front as an optically thick gray body will lead to a noticeable overestimation of the shock temperature.

36 MATERIALS SCIENCE↗

Modular Autonomous Experimentation for Biological Applications (Full Report)

The Modular Autonomous Research System (MARS) was developed to address the pressing need for faster, more reliable, and more adaptable scientific discovery. Traditional experimentation is limited by manual labor, long cycle times, and fragmented data streams, which constrain the ability to explore complex chemical and materials design spaces. To overcome these limitations, we created an integrated, modular platform that combines laboratory robotics, diverse measurement instruments, and a central data infrastructure with artificial intelligence–driven decision-making. The system links liquid handling robots, robotic arms, and optical plate readers into a closed loop where experiments are executed automatically, data is analyzed in real time, and subsequent experimental conditions are adaptively chosen to maximize information gain. Over the course of the project, MARS was validated on two primary test cases—spectroscopic metal–ligand binding assays and peptide-directed mineralization—which highlighted the system’s ability to handle uncertainty and variability in experimental measurements. To further demonstrate modularity and extensibility, we also established additional testbeds in electrochemistry for catalyst discovery and electrolyte formulation for advanced batteries. The results show that MARS can reliably conduct autonomous campaigns with minimal human intervention, adapt to distinct scientific domains, and provide a scalable model for future self-driving laboratories. This work establishes new capabilities for modular, uncertainty-aware automation and directly supports the need for advanced, data-driven research platforms capable of accelerating discovery across a wide range of scientific and national security missions.

59 BASIC BIOLOGICAL SCIENCES↗

Quantum Time Dynamics Mediated by the Yang–Baxter Equation and Artificial Neural Networks

Quantum computing shows great potential, but errors pose a significant challenge. This study explores new strategies for mitigating quantum errors using artificial neural networks (ANNs) and the Yang–Baxter equation (YBE). Unlike traditional error mitigation methods, which are computationally intensive, we investigate artificial error mitigation. We developed a novel method that combines ANNs for noise mitigation combined with the YBE to generate noisy data. This approach effectively reduces noise in quantum simulations, enhancing the accuracy of the results. The YBE rigorously preserves quantum correlations and symmetries in spin chain simulations in certain classes of integrable lattice models, enabling effective compression of quantum circuits while retaining linear scalability with the number of qubits. This compression facilitates both full and partial implementations, allowing the generation of noisy quantum data on hardware alongside noiseless simulations using classical platforms. By introducing controlled noise through the YBE, we enhance the data set for error mitigation. We train an ANN model on partial data from quantum simulations, demonstrating its effectiveness in mitigating errors in time-evolving quantum states, providing a scalable framework to enhance quantum computation fidelity, particularly in noisy intermediate-scale quantum (NISQ) systems. We demonstrate the efficacy of this approach by performing quantum time dynamics simulations using the Heisenberg XY Hamiltonian on real quantum devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗