Search NASASearch

SEARCH · Search NASA

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform

Spurious solar-wind effects on acceleration noise in LISA Pathfinder

Spurious solar-wind effects are a potential noise source in future Laser Interferometer Space Antenna (LISA) measurements. One noise coupling mechanism is constrained by estimating solar-wind effects on acceleration noise in LISA Pathfinder (LPF). While LISA is designed for drag-free differential measurement, predicting the realistic impact both bounds the operational environment and assesses whether LISA could provide serendipitous space-weather observations. Data from NASA's Advanced Composition Explorer (ACE), situated at the L1 Lagrange point, serves as a reliable source of solar-wind data. The data sets are compared over the 114 d time period from 1 March 2016 to 23 June 2016. This period gives the longest readily-available open data set, without interference from other commissioning activities. To evaluate space weather effects, the data from both satellites are formatted, gap-filled/interpolated, and fast-Fourier transformed for amplitude spectral density and coherence comparisons. Solar wind effects are not seen in a coherence plot between LPF and ACE; modest coherence in the planned LISA observational frequency band can be attributed to chance. This result indicates that measurable correlation due to solar-wind acceleration noise over 3 month timescales will be a negligible noise source. LISA is unlikely to inform solar wind measurements routinely. Another source of noise from the Sun, solar radiation pressure, is estimated to impart greater acceleration noise, but has yet to be analyzed.

79 ASTRONOMY AND ASTROPHYSICS

Open Source Intelligence for Cybersecurity Events via Twitter Data

Open-Source Intelligence (OSINT) is largely regarded as a necessary component for cybersecurity intelligence gathering to secure network systems. With the advancement of artificial intelligence (AI) and increasing usage of social media, like Twitter, we have a unique opportunity to obtain and aggregate information from social media. In this study, we propose an AI-based scheme capable of automatically pulling information from Twitter, filtering out security-irrelevant tweets, performing natural language analysis to correlate the tweets about each cybersecurity event (e.g., a malware campaign), and validating the information. This scheme has many applications, such as providing a means for security operators to gain insight into ongoing events and helping them prioritize vulnerabilities to deal with. To give examples of the possible uses, we present three case studies demonstrating the event discovery and investigation processes. We also examine the potential of OSINT for identifying the network protocols associated with specific events, which can aid in the mitigation procedures by informing operators if the vulnerability is exploitable given their system’s network configurations.

Dale, Dakota

The Analysis Description Language Ecosystem: Latest developments and physics applications

We present latest developments in Analysis Description Language (ADL), a declarative domain-specific language describing the physics algorithm of a HEP data analysis decoupled from software frameworks. Analyses written in ADL can be integrated into any framework for various tasks. ADL is a multipurpose construct with uses ranging from analysis design to preservation, reinterpretation, queries, visualisation, combination, etc. The most advanced infrastructure to execute ADL on events is the CutLang runtime interpreter. Recent technical developments include an automated interface with different data types, generation of the abstract syntax tree, a visualization tool that that auto-converts analysis flows to graphs, incorporation of trained machine learning models and a Jupyter-based plotting tool. We also report physics implications including a large scale LHC analysis implementation and validation effort for beyond the standard model reinterpretation purposes and studies with ATLAS and CMS open data.

Sekmen, Sezen [Kyungpook National Univ., Daegu (Ko

ν-point energy correletors with F AST EEC: Small-x physics from LHC jets

In recent years, energy correlators have emerged as a powerful tool for studying jet substructure, with promising applications such as probing the hadronization transition, analyzing the quark-gluon plasma, and improving the precision of top quark mass measurements. The projected N-point correlator measures correlations between N final-state particles by tracking the largest separation between them, showing a scaling behavior related to DGLAP splitting functions. These correlators can be analytically continued in N, commonly referred to as ν-correlators, allowing access to non-integer moments of the splitting functions. Of particular interest is the ν → 0 limit, where the small momentum fraction behavior of the splitting functions requires resummation. Originally, the computational complexity of evaluating ν-correlators for M particles scaled as 2 2M , making it impractical for real-world analyses. However, by using recursion, we reduce this to M 2M , and through the FastEEC method of dynamically resolving subjets, M is replaced by the number of subjets. This breakthrough enables, for the first time, the computation of ν-correlators for LHC data. In practice, limiting the number of subjets to 16 is sufficient to achieve percent-level precision, which we validate using known integer-ν results and convergence tests for non-integer ν. We have implemented this in an update to FastEEC and conducted an initial study of power-law scaling in the perturbative regime as a function of ν, using CMS Open Data on jets. The results agree with DGLAP evolution, except at small ν, where the anomalous dimension saturates to a value that matches the BFKL anomalous dimension.

Energy correlators

Conformal collider physics meets LHC data

The remarkably high energies of the Large Hadron Collider (LHC) have allowed for the first measurements of the shapes and scalings of multipoint correlators of energy flow operators, ⟨ Ψ | E ( n → 1 ) E ( n → 2 ) ⋯ E ( n → k ) | Ψ ⟩ , providing new insights into the Lorentzian dynamics of quantum chromodynamics (QCD). In this letter, we use recent advances in effective field theory to derive a rigorous factorization theorem for the light-ray density matrix, ρ = | Ψ ⟩ ⟨ Ψ | , inside high transverse momentum jets at the LHC. Using the light-ray operator product expansion, the scaling behavior of multipoint correlators can be computed from the expectation value of the twist-2 spin- J light-ray operators, O [ J ] , in this state, Tr [ ρ O [ J ] ] . We compute the light-ray density matrix at next-to-leading order, and combine this with results for the next-to-leading logarithmic scaling behavior of the correlators up to six-points, comparing with CMS open data. This theoretical accuracy allows us to resolve the quantum scaling dimensions of QCD light-ray operators inside jets at the LHC. Our factorization theorem for the light-ray density matrix at the LHC completes the link between recent developments in the study of energy correlators and LHC phenomenology, opening the door to a wide variety of precision jet substructure studies. Published by the American Physical Society 2025

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

The NASA ACTIVATE Mission

The NASA Aerosol Cloud Meteorology Interactions over the Western Atlantic Experiment (ACTIVATE) conducted 162 joint flights with two aircraft over the northwest Atlantic to study aerosol–cloud interactions (ACIs), which represent the largest uncertainty in estimating total anthropogenic radiative forcing. The combination of a high-flying King Air and low-flying HU-25 Falcon, equipped with remote sensing and in situ instruments, characterized trace gases, aerosol particles, clouds, and meteorological variables with data collected nearly simultaneously below, within, and above marine boundary layer (MBL) clouds. Flights spanning warm and cold seasons across 3 years (2020–22) provided a broad range of conditions associated with aerosol particles, cloud properties (including particle size and phase), and meteorology, ideally suited for robust ACI calculations and assessing how well models simulate a wide range of MBL clouds from stratiform to cumulus. ACTIVATE data suggest that drivers of cloud droplet number concentration N d , including aerosol particles and MBL dynamics, vary between winter and summer months with a stronger potential to convert aerosol particles into cloud droplets in winter. Models of varying complexity not only highlight some skills in simulating winter and summer cloud types but also identify challenges that still need to be addressed such as treatment of turbulence, wet scavenging, and mesoscale organization. Remote sensing advances range from new retrieval methods for N d , cloud phase classification, vertically resolved aerosol and cloud condensation nuclei number concentration, and ocean surface wind speed. This work describes these scientific and technological advances along with efforts in outreach and open data science.

aerosol indirect effect

Using Visual Systems Mapping to Improve Transparency and Comparability of Life Cycle Assessment Baseline Scenarios

Visual systems mapping is a systems engineering approach used to represent complex processes and interactions. This study evaluates its application for documenting assumptions in life cycle assessment (LCA) baseline scenarios. In LCA, the baseline or reference case represents the business as usual system against which changes in impacts (e.g., emissions) are assessed. These baseline assumptions are particularly influential in biomass LCAs, yet they often vary across studies due to regional context, system boundaries, and simplifying assumptions that are not consistently or transparently documented. As a result, key feedbacks, omitted processes, and boundary choices may remain unclear, limiting comparability across studies and weakening their usefulness for decision-making. This study examines whether visual systems mapping can improve the transparency and comparability of biomass LCA baseline scenarios. A case study of five published biomass-related LCAs were reviewed, and their baseline scenarios were translated into visual system maps to identify included processes, omitted components, and underlying assumptions. The analysis demonstrates that visual systems mapping can make baseline assumptions more explicit, highlight excluded dynamics, and improve documentation of system boundaries. Based on these findings, the study recommends the use of visual systems mapping alongside open data repositories and reproducible workflows to support greater transparency, reproducibility, and comparability in LCAs. These improvements can strengthen the role of LCAs in informing decisions related to sustainable biomass systems.

Davis, Maggie [ORNL] (ORCID:0000000181319328)

Integration of a grey-box refrigerated case model in EnergyPlus via Python plugin

Commercial buildings, in particular grocery stores (due mainly to their large refrigeration load), provide opportunities for energy cost reductions. Grocery stores could offer substantial load flexibility to the power grid through participation in demand response programs because of their usage patterns and relatively high energy intensity. This load flexibility could come from modifying the control of heating, ventilation, and air conditioning (HVAC) systems, refrigeration systems, or both. Although estimation of the HVAC system’s load flexibility potential is relatively targeted in the literature, estimating load flexibility of refrigeration systems is nascent and has been a challenge, in part because of the lack of proper simulation tools that capture the dynamics in the refrigeration cases. The existing refrigerated case model within EnergyPlus, a whole building energy simulation program, assumes a constant case temperature throughout the simulation period and does not explicitly model the cycling of the compressor serving the refrigerated case. In addition, it does not encompass modeling of temperatures of the product inside the refrigerated case. This difference between modeled and actual operation can be a barrier to the development of demand control algorithm and accurate analysis of load flexibility potential. In this paper, we present a grey-box model for modeling refrigerated cases in grocery stores, which include medium temperature and low temperature. Four cases are modeled; two are low-temperature closed cases and two are medium-temperature cases with one closed and one open. Data from an experimental facility are used to train and test the models. Results demonstrate the efficacy of the grey-box models in predicting the temperatures. This model is integrated into EnergyPlus to capture the dynamic effects of case temperature on the environment and enhance the calculation of sensible and latent heat exchange with the environment (case credits). These enhancements can be leveraged more broadly to model advanced refrigeration controls such as defrost, develop and test unique algorithms that could affect refrigeration interactions with HVAC, and refine store design for any commercial building with refrigeration.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Multimodal super-resolution: discovering hidden physics and its application to fusion plasmas

Understanding complex physical systems often requires integrating data from multiple diagnostics, each with limited resolution or coverage. We present a machine learning framework that reconstructs synthetic high-temporal-resolution data for a target diagnostic using information from other diagnostics, without direct target measurements during the inference. This multimodal super-resolution technique improves diagnostic robustness and enables monitoring even in case of measurement failures or degradation. Applied to fusion plasmas, our method targets edge-localized modes (ELMs), which can damage plasma-facing materials. By reconstructing super-resolution Thomson Scattering data from complementary diagnostics, we uncover fine-scale plasma dynamics and validate the role of resonant magnetic perturbations (RMPs) in ELM suppression through magnetic island formation. The approach provides new observation supporting the plasma profile flattening due to these islands. Our results demonstrate the framework’s ability to generate high-fidelity synthetic diagnostics, offering a powerful tool for ELM control development in future reactors like ITER. The approach is broadly transferable to other domains facing sparse, incomplete, or degraded diagnostic data, opening new avenues for discovery.

Jalalvand, Azarakhsh [Princeton Univ., NJ (United

Geographic, sectoral, and constituent characteristics of US off‐site manufacturing wastewater disposal

This study employs recent US Environmental Protection Agency off-site manufacturing wastewater disposal data and a socioenvironmental assessment tool for the United States to understand the characteristics of such disposal in terms of geography, major contributing sectors and pollutants, and environmental impacts. Environmental impact analyses of manufacturing are insufficient without encompassing the extended physical boundary of waste management. Our analysis reveals that off-site manufacturing wastewater disposal occurs disproportionately in populations residing in census tracts already negatively impacted by environmental hazards (impacted populations, or IPs) with 44% of transfers and 55% of Risk Screening Environmental Indicators (RSEI) Hazard going to these areas, compared to their national share of 37%. Disposal hazard is concentrated in a small number of populations, with a Gini coefficient of 0.99 for RSEI Hazard. Four manufacturing sub-sectors are significant generators: Chemicals, Fabricated Metals, Primary Metals, and Transportation Equipment, with Chemicals off-site wastewater disposal largely occurring in IPs. For individual contaminants, chromium compounds and chromium represent more than 85% of the hazard but less than 10% of transfers. We explore transfer distances and waste generation and disposal hotspots, finding that the Midwest hosts a disproportionate share of off-site wastewater disposal. Further, RSEI Hazard steeply rises at shorter distances and plateaus over distances >500 miles, revealing opportunities to reduce hazard by reducing 20–500-mile transfers. Our findings strongly support targeted mitigation strategies like process substitutions, control technologies, on-site recycling and treatment, and minimizing transfer distances. This article met the requirements for a gold-gold JIE data openness badge described at http://jie.click/badges.

Fuchs, Heidi [Lawrence Berkeley National Laborator

Sparks in the dark

This study presents a novel method for the definition of signal regions in searches for new physics at collider experiments. By leveraging multi-dimensional histograms with precise arithmetic and utilizing the SparkDensityTree library, it is possible to identify high-density regions within the available phase space, potentially improving sensitivity to very small signals. Inspired by a search for dark mesons at the ATLAS experiment, CMS open data is used for this proof-of-concept intentionally targeting an already excluded signal. Signal regions are defined based on density estimates of signal and background. These preliminary regions align well with the physical properties of the signal while effectively rejecting background events.

Gudnadottir, Olga Sunneborn [Uppsala Univ. (Sweden

Pumped Storage Hydropower Resource and Deployment Potential

There is growing interest in deploying new pumped storage hydropower (PSH) to meet grid needs for flexibility, reliability, and resiliency. This presentation describes how NLR resource assessment, cost modeling, and capacity expansion modeling are used to identify technical and economic PSH deployment potential and support industry decision-making on PSH investments. NLR's open data and tools demonstrate the vast technical potential of PSH on the order of 80 TW, and modeling shows economic PSH deployment under an attractive cost-value proposition.

13 HYDRO ENERGY

Aspen Open Jets: unlocking LHC data for foundation models in particle physics

Foundation models are deep learning models pre-trained on large amounts of data which are capable of generalizing to multiple datasets and/or downstream tasks. This work demonstrates how data collected by the CMS experiment at the Large Hadron Collider can be useful in pre-training foundation models for HEP. Specifically, we introduce the AspenOpenJets (AOJs) dataset, consisting of approximately 178 M high p T jets derived from CMS 2016 Open Data. We show how pre-training the OmniJet-α foundation model on AOJs improves performance on generative tasks with significant domain shift: generating boosted top and QCD jets from the simulated JetClass dataset. In addition to demonstrating the power of pre-training of a jet-based foundation model on actual proton–proton collision data, we provide the ML-ready derived AOJs dataset for further public use.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES

Recommendations for Best Practices for Data Preservation and Open Science in HEP

These recommendations are the result of reflections by scientists and experts who are, or have been, involved in the preservation of high-energy physics data. The work has been done under the umbrella of the Data Lifecycle panel of the International Committee of Future Accelerators (ICFA), drawing on the expertise of a wide range of stakeholders. A key indicator of success in the data preservation efforts is the long-term usability of the data. Experience shows that achieving this requires providing a rich set of information in various forms, which can only be effectively collected and preserved during the period of active data use. The recommendations are intended to be actionable by the indicated actors and specific to the particle physics domain. They cover a wide range of actions, many of which are interdependent. These dependencies are indicated within the recommendations and can be used as a road map to guide implementation efforts. These recommendations are best accessed and viewed through the web application, see https://icfa-data-best-practices.app.cern.ch/

Campana, Simone [CERN]