Search NASA⌕ Search

SEARCH · Search NASA

Results for “open access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Spin-Controllable Dynamics in Defect-Engineered Carbon Nanotubes as Single Photon Emitters: Data-Driven Modeling and Computations

Quantum technologies, such as quantum computing and sensing, require efficient single-photon emission (SPE) sources that operate at room temperature in telecom wavelengths. While several materials can serve as SPE sources, no single platform meets all the criteria for efficiency, ambient operation, and scalability. Single-walled carbon nanotubes (SWCNTs) with covalently attached molecules offer a promising solution. Their SPE can be easily tuned via modifications of the SWCNT's diameter, chirality, and bonded molecules, enabling emission across near-IR to telecom wavelengths at ambient conditions. However, to fully realize the potential of SWCNTs and unlock their quantum capabilities, a deeper understanding of how structural defects from molecular adducts affect their emission and competing photoexcited processes is essential. To address this gap in our knowledge, this project combined quantum chemistry calculations with data-driven methods of cheminformatics (QSAR) and machine learning (ML). The developed computational approaches have provided several design strategies for covalent functionalization of SWCNTs to improve their optical response. The collaboration with Los Alamos National Lab (LANL) enabled direct comparison of computational and experimental data, facilitating method validation. This partnership was enhanced through access to LANL's Center for Integrated Nanotechnologies (CINT) utilizing User Facility Program and summer internships, which provided three NDSU graduate students with hands-on experience at LANL. The outcomes of this project included (1) Advancing the current stage of computational methods in accurate modeling of non-adiabatic spin-dependent photoexcited dynamics and its applicability to nanosystems consisting of thousands of atoms, realized as open-access codes linked to existing DFT-based software; (2) Establishing the relationship between the structure of adducts and SWCNTs and intrinsic excitonic and spin properties of defect states for guiding novel synthetic strategies and experimental probes of chemically functionalized SWCNTs as near-IR emitting materials; (3) Generating virtual libraries of hypothetical functionalized SWCNTs for virtual screening of their chemical structures and optical properties, leveraging new functionalities of SWCNTs; (4) Offering a unique experience for NDSU graduate students that prepared them for future scientific careers related to materials modeling and big data processing. These results were summarized in 12 published journal papers and 3 recently submitted papers. One of a key finding is that the position of defect sites on the SWCNT surface primarily drives the emission redshift (up to 100 meV), while the polarity of the defect-inducing molecules has a much smaller effect (~10 meV). However, the electron-donating or withdrawing properties of a molecule influence selecting reactivity of defect sites. These insights important for optimizing synthetic protocols for desired emissions in SWCNTs. We also revealed that the interaction between two defects at various positions on the SWCNT enhances the redshift and optical activity of states, favoring strong near-IR emission. This suggests that manipulations in defect concentrations is a promising strategy for controlling efficient emission. Mostly important, the defect position was found controllable by the spin states of photoexcited intermediates: Excited aromatic molecules form ortho defects with SWCNTs at their singlet states in the presence of oxygen, while oxygen-free conditions favor para defects via the triplet-state mechanism. Additionally, a heat-activated [2+2] cycloaddition reaction facilitates divalent defect formation with fewer bonding positions that narrows emission bands. These groundbreaking findings have been experimentally validated and significantly advance our understanding of defect chemistry in SWCNTs. Using a novel encoding technique and 3D-MoRSE descriptors, we developed highly accurate ML/QSAR models to predict both the 3D structure and optical properties of SWCNTs with chemical defects. This model enabled the creation of a virtual library of 125,556 structures, providing new insights into the relationship between SWCNT-defect structure and emission.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

The Pierre Auger Observatory open data

The Pierre Auger Collaboration has embraced the concept of open access to their research data since its foundation, with the aim of giving access to the widest possible community. A gradual process of release began as early as 2007 when 1% of the cosmic-ray data was made public, along with 100% of the space-weather information. In February 2021, a portal was released containing 10% of cosmic-ray data collected by the Pierre Auger Observatory from 2004 to 2018, during the first phase of operation of the Observatory. The Open Data Portal includes detailed documentation about the detection and reconstruction procedures, analysis codes that can be easily used and modified and, additionally, visualization tools. Since then, the Portal has been updated and extended. In 2023, a catalog of the highest-energy cosmic-ray events examined in depth has been included. A specific section dedicated to educational use has been developed with the expectation that these data will be explored by a wide and diverse community, including professional and citizen scientists, and used for educational and outreach initiatives. This paper describes the context, the spirit, and the technical implementation of the release of data by the largest cosmic-ray detector ever built and anticipates its future developments.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The U.S. Agrivoltaic Shading Tool: A National-Scale Interface for Modeling Light and Shade Patterns in Ten Common Agrivoltaic Configurations

Agrivoltaic systems are dual-use configurations that co-locate agriculture and photovoltaic (PV) infrastructure and require careful design to balance crop performance and energy generation. A critical element of agrivoltaic design is the spatial and temporal distribution of irradiance and shade within and around PV arrays. To support research, planning, and stakeholder decision-making, we introduce the U.S. Agrivoltaic Shading Tool, a novel web-based application that delivers high-resolution irradiance and photosynthetically active radiation (PAR) modeling for ten standardized PV configurations across the conterminous United States. The tool leverages the National Laboratory of the Rockies (NLR) System Advisor Model (SAM) to perform detailed irradiance simulations, using meteorological data from the National Solar Radiation Database (NSRDB). Outputs include seasonal, monthly, weekly, and diurnal patterns of available sunlight, amount of shade, irradiance, and PAR at ground level within agrivoltaic system footprints. For a user's selected location, these results are visualized through interactive visualizations, heatmaps, and time-series plots, designed to be accessible to both technical and non-technical users. In addition to facilitating rapid spatial exploration of agrivoltaic light environments, the tool will offer seamless integration with the InSPIRE Agrivoltaics Design and Analysis Model (ADAM). This optional workflow will allow users to port selected site and configuration parameters into a more advanced modeling environment for further customization of structural layouts, crop-system compatibility, power generation, and technoeconomic performance. Finally, to promote open science, the entire dataset will be hosted and available for open access through the OpenEI platform. By standardizing and disseminating high-quality irradiance data and design tools, the U.S. Agrivoltaic Shading Tool supports a wide range of users, including researchers, landowners, energy developers, and policymakers, in evaluating the agronomic and energetic feasibility of agrivoltaic systems across the United States.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Building a Bilingual Google Earth Engine Dashboard to Increase Accessibility to Long-term Time Series Remote Sensing Data for Monitoring Saline System Changes in Chile’s Atacama Desert

Saline systems, consisting of salt flats, ponds, and marshes, provide vital water resources to wildlife and communities in northern Chile’s Atacama Desert, one of the driest regions in the world. Mining is extensive in the Atacama, which contains 30% of the world’s lithium reserves and is abundant in potassium and boron. The groundwater that feeds into salt marshes and ponds is extracted in large volumes for mining operations, limiting the availability of water for ecosystems. However, identifying long-term and large-scale environmental impacts from local lithium mining on the saline systems is limited by region inaccessibility and terrain variability. Open access satellite imagery and cloud computing technology has made studying Atacama saline systems feasible and allowed for collaboration across different agencies and countries. The NASA DEVELOP Program partnered with Chile’s la Universidad de La Serena and Servicio Nacional de Geología y Minería (SERNAGEOMIN) to create the Saline Analysis Tool (SalT) in Google Earth Engine (GEE). SalT is used to analyze the extent and distribution of remote saline systems in the Atacama from 1986 to the present day. The tool filters Landsat 5 Thematic Mapper (TM) and Landsat 8 Operational Land Imager (OLI) data from GEE’s data catalog and creates a single composite image per year for analysis. Additional output analyses include land cover classification, Normalized Difference Vegetation Index (NDVI) and Normalized Difference Water Index (NDWI) raster images that can be displayed on the map interface or exported. The tool can also generate time-lapse videos and charts displaying NDVI, NDWI, and land cover over time. A key feature of the tool is the use of a bilingual graphical user interface to make analysis accessible and customizable to different users’ needs—SalT provides options to select an analysis area, analysis time period, and outputs to display or export. The tool also incorporates new Earth observations as they are added to GEE’s catalog. The ability to easily visualize and analyze long-term remote sensing imagery will enable SERNAGEOMIN and la Universidad de la Serena to continually monitor changes in these saline systems and inform future land management policy.

NASA DEVELOP↗

Evaluating SAR Radiometric Terrain Correction products: Optimal products for applied users

Operational applications for Synthetic Aperture Radar (SAR) are under development around the world, driven by the free-and-open access of SAR C-band observations that Sentinel-1 of Copernicus has been providing since 2014. Groups like SERVIR, a joint initiative between NASA and USAID, are at the forefront of remote sensing applied uses, and have made many significant contributions to lower the barrier to access, process, and apply SAR for ecosystem services. A takeaway from the SERVIR experience in using SAR is the need to use the appropriate SAR polarimetric product. Radiometric Terrain Correction (RTC) is a key entry-level product for multiple applications that range from ecosystems to hazards. Many software packages exist to create RTC products from SLC or GRD-type Level-1 SAR data, some of which were released only recently, e.g. Interferometric SAR Computing Environment (ISCE) added an RTC module in April 2020. In addition, new versions of open source softwares are expected to address known issues from previous versions, such as Sentinel-1 Toolbox from the European Space Agency (SNAP-7). Despite the growing availability of RTC software solutions, little work has been done to identify differences between RTC products from different softwares. And to address the question, which open-source software produces the most accurate RTC product? This work evaluates Sentinel-1 RTC products created with three different softwares and approaches, including SNAP-7, ISCE-2, and a pseudo RTC product derived from GEE. The GAMMA-derived RTC product, a known optimal RTC and implemented by Alaska Satellite Facility (ASF), is used as a reference. Time series stacks over ten different sites representing varied terrain and ecosystems are evaluated. Products are evaluated for geolocation quality, absolute radiometric calibration, and for the fidelity of the radiometric terrain flattening. The results provide direct guidance and recommendations about the quality of the RTC products obtained from open source methods. This understanding is key to develop operational applications that rely on SAR Sentinel-1 data that need affordable and scalable solutions.

Africa Flores-Anderson↗

Dissemination of Global Flood Severity and Surface Water Mapping using Remote Sensing Data to Global Stakeholders

Flooding is a natural event that occurs frequently with high severity worldwide, responsible for significant societal and economic impacts. Disaster managers face significant challenges managing essential information for preparedness, response, and recovery efforts. The development of an open access, global flood alerting system for effective identification of flood impacted areas, classification of potential impacts, and the formulation of effective emergency response measures requires the incorporation of a wide variety of flood models and remote sensing data sources from multiple platforms. NASA is currently funding projects focused on flood forecasting, post-event flood mapping, flood depth estimation and pre-event flood severity estimation using Earth observation (EO) datasets and derived flood products. A new initiative in the Disasters Program is underway to disseminate flood products from different hydrologic models and sensors to global stakeholders via Pacific Disaster Center’s DisasterAWARE®, NASA’s Disasters Mapping Portal and potentially other mechanisms. This initiative focuses on improving response capacity and use of EO products in near real-time by a broader community for resource planning in case of extreme events. As part of this initiative, we have deployed Model of Models (MoM) – an open-source ensemble approach, that integrates outputs from hydrologic models and EO data from optical imagery to assess flood severity daily at sub-watershed level globally. The MoM output is integrated with the incident event system of DisasterAWARE to generate flood severity risk and flood impact boundaries, which are disseminated via the DisasterAWARE platform to different stakeholders globally for decision-making and response efforts. The next step will focus on using MoM outputs to estimate flood depth and extent mapping using high-resolution Synthetic Aperture Radar imagery, impact assessment using optical imagery and population datasets, and damage estimation using critical infrastructure datasets, which would be disseminated via DisasterAWARE to decision-makers, emergency managers and first responders around the world.

flood↗

A Comprehensive Calibration Framework for the Northwest River Forecast Center

We present a comprehensive framework developed by the Northwest River Forecast Center for calibrating hydrologically diverse basins. The framework includes models for snow, soil moisture, routing, channel loss, and consumptive use. Data inputs include a wide range of open-access datasets for meteorology, land use, topography, and land cover. The framework uses conceptual hydrologic models to handle basins with various hydrologic regimes including rain-driven and snowmelt-dominated basins. We also develop a flexible automatic calibration system that can handle numerous unobservable model parameters in a computationally efficient manner. A single-basin automatic calibration run can typically be completed on a modern laptop in under 10 min. We found that model performance metrics for this new approach match the quality of the NWRFC's previous labor-intensive manual calibrations. The model performance also rivals that of a state-of-the-art deep learning model at a fraction of the computational cost. This framework presents a new standard for the quality of calibrations possible with lumped conceptual hydrologic models, combining careful data curation, an objective calibration framework, and expert local knowledge. In addition, we have made software packages available for the entire suite of National Weather Service River Forecast System models, including SAC-SMA, SNOW-17, and Lag-K. These modern interfaces are intended to increase accessibility and facilitate future research.

Forecasting↗

Spaceflight Biospecimen Sharing in Support of Science Discovery and Exploration

For decades, NASA and international partners have flown non-human biological experiments in space to understand the effects of spaceflight and address potential biological hazards. Sending organisms into space is a costly endeavor which makes space-flown biological specimens a valuable resource. To enable maximum scientific return, samples not required by the Principal Investigators are harvested and collected mostly by NASA’s Space Biology Biospecimen Sharing Program. These specimens are collected according to well-established SOPs that maintain quality and integrity. The specimens are then preserved, archived, and made available to the international scientific community through NASA’s Institutional Scientific Collection (ISC) at Ames Research Center (ARC). The ISC-ARC biospecimens and descriptive metadata are findable and accessible for request through the Life Sciences Data Archive (LSDA). The NASA ISC-ARC currently stores over 32,000 specimens from Shuttle, International Space Station, and ground-based investigations (spaceflight analog experiments involving either hindlimb unloading, centrifugation, or partial weight-bearing study designs). Tissues are predominantly from mice and rats, though samples are also available from bacteria and quail. The specimens include tissues from many physiological systems including musculoskeletal, neurosensory, reproductive, respiratory, circulatory, and digestive. Tissues are stored at -80°C, -20°C, +4°C, or ambient and preserved in various fixatives. Descriptive metadata is available for all samples. Historically, these tissues have been used for a wide range of analyses, including histology, genomics, and transcriptomics. Plans are underway to expand the ISC-ARC beyond the mostly-rodent contents, to include a space-relevant microbial culture collection including bacteria, fungi, and yeast. This expansion of the ISC-ARC will now involve identifying and standardizing best practices for microbial curations. To ensure safe long-term storage of microbial isolates, a microbiology laboratory will be dedicated for identification, cell culture, and lyophilization. Awarding of tissue to public science investigators has resulted in 33 publications since 2011, with 48 requests being submitted since 2016. Of note, NASA GeneLab has been awarded ISC-ARC biospecimens in the past few years. GeneLab processes the biospecimens to generate various levels of ‘omics’ data, which are published on GeneLab’s open access online platform for bioinformatics analysis and visualization. This has helped a systems biology community grow around the processed-biospecimens’ datasets, resulting in many new publications and insights. Websites: https://www.nasa.gov/ames/research/space-biosciences/isc-bsp ; https://lsda.jsc.nasa.gov/Biospecimen

Ryan T. Scott↗

GeneLab: A Systems Biology Platform for Omics Analysis

NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). The GeneLab Data Systems (GLDS) version 4.0 will be available on October 1st 2019, and will provide the latest in terms of professional state-of-the-art bioinformatics platform for the space biology and radiation community to upload their data into an omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. Started in 2015 as a repository designed to archive omics data from space experiments, GeneLab has expanded its scope to all ionizing radiation omics experiments conducted on the ground and has put considerable effort in providing carefully characterized radiation metadata on all dataset. GeneLab is also providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome) to help users explore important questions: 1) Which genes or proteins are expressed differently in space for various living organisms? 2) What specific DNA mutations or epigenetic changes happen in space or after exposure to ionizing radiation? and 3) How does genetics affect these responses? Processed data available on GeneLab are derived by standard data analysis workflows vetted by hundreds of scientists who volunteered to join one of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). In this presentation, we will discuss how to bridge the gap between irradiation studies performed on earth and biological experiments conducted in space since the early 1990's. We will discuss how radiation dosimetry was estimated for datasets derived from samples collected during the Space Shuttle era or on the International Space Station. Finally, we will address future strategies regarding dose monitoring in future missions into space, inter-agency efforts to unify data under one umbrella, and knowledge dissemination across the radiation research community and the space biology community.

open-science↗

NASA GeneLab Space Omics Database: Expanding from Space to Ionizing Radiation Data on the Ground

NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or ground experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). The GeneLab Data Systems (GLDS) version 4.0 will be available on October 1st 2019, and will provide a state-of-the-art bioinformatics platform for the space biology and radiation communities to upload their data into an omics data commons, to process their data with vetted standard workflows and to compare with existing analyses. Started in 2015 as a repository designed to archive omics data from space experiments, GeneLab has expanded its scope to all ionizing radiation omics experiments conducted on the ground and has put considerable effort in providing carefully characterized radiation metadata on all datasets. GeneLab is also providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome) to help users explore important questions: 1) Which genes or proteins are expressed differently in space for various living organisms? 2) What specific DNA mutations or epigenetic changes happen in space or after exposure to ionizing radiation? and 3) How does genetics affect these responses? Processed data available on GeneLab are derived by standard data analysis workflows vetted by hundreds of scientists who volunteered to join one of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). In this presentation, we will discuss how to bridge the gap between irradiation studies performed on earth and biological experiments conducted in space since the early 1990's. We will discuss how radiation dosimetry was estimated for datasets derived from samples collected during the Space Shuttle era on the International Space Station and on other orbiting platforms. Finally, we will address future strategies regarding dose monitoring in future missions into space, inter-agency efforts to unify data under one umbrella, and knowledge dissemination across the radiation research community and the space biology community.

open-science↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Changes to virus taxonomy, the international code of virus classification and nomenclature, and the ICTV statutes ratified by the International Committee on Taxonomy of Viruses (2025)

Abstract The 56th meeting of the Executive Committee (EC) of the International Committee on Taxonomy of Viruses (ICTV) was held in Bari, Italy, in July/August, 2024, and 115 submitted taxonomy proposals were reviewed. A total of 112 were subsequently ratified by the ICTV membership. An additional 9 error correction proposals were also approved in August 2025. This article lists the taxonomy proposals that have now been incorporated into release 40 version v2 of the Master Species List ( https://ictv.global/msl ), the Virus Metadata Resource ( https://ictv.global/vmr ), and associated ICTV databases. In addition to the assignments of 1,563 new virus species, 243genera, 55 families, 11 orders, and 8 classes, there were substantial additions to higher taxonomic ranks. These include the creation of a new realm ( Singelaviria ), which is based on the recognition of a separate evolutionary origin for the hallmark capsid genes of members of the kingdom Helvetiavirae. These express capsid proteins forming a single jelly-roll fold that is structurally and evolutionarily distinct from those of members of the family Bamfordvirae , assigned to the realm Varidnaviria . Furthermore, the realm Varidnaviria underwent a major reorganization, including the addition of a new kingdom, Abadenavirae . Another notable change was the classification of the vertebrate-infecting single-stranded DNA anellovirids into a new phylum Commensaviricota (kingdom Shotokuvirae , realm Monodnaviria ). Archaeal viruses infecting the hyperthermophilic Archaeoglobi were assigned to a new phylum Calorviricota , in the kingdom Trapavirae (realm Monodnaviria ), whereas RNA viruses infecting hyperthermophilic bacteria were classified into a new phylum Artimaviricota (realm Riboviria ). In recognition of his extensive and valuable contributions to virus taxonomic developments in Study Groups and over the period of his EC membership, Stuart Siddell was honoured as a new life member of the ICTV. The ICTV has created a new strategy for disseminating information on taxonomy advances through annual open-access publication of citeable taxonomy proposal summaries from each ICTV Subcommittee. A collective total of 354 co-authors of the seven summaries were drawn from members of each Subcommittee, the EC, and a very large number of contributors from the wider virology community.

Simmonds, Peter (ORCID:0000000279644700)↗

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database↗

A new database website for nuclear level densities

We introduce a new open-access, web-based database (http://nld.ascsn.net), Current Archive of Nuclear Density of Levels (CANDL), that hosts experimental nuclear level density (NLD) datasets from a variety of techniques and energy ranges. Built using the Dash framework in Python, the database is designed to be interactive and user-friendly, allowing researchers to search, visualize, fit, and export NLD data with minimal effort. This resource includes data extracted from evaporation spectra, Oslo method variants, and other experimental techniques that cover excitation energies beyond the neutron resonance region. The database supports on-the-fly fitting with two widely-used phenomenological models—the Constant Temperature (CT) model and the Back-Shifted Fermi Gas (BSFG) model—selected for their simplicity and computational efficiency. Future versions aim to include additional datasets and model types, as well as easy-to-use interfaces to data science techniques. Here, this platform offers a vital tool for the nuclear physics, astrophysics, medicine, and reactor design communities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

TPCpp-10M: Simulated proton-proton collisions in a time projection chamber for AI foundation models

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community. This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability. The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of $\sqrt{s}$ = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory. This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.

Data Analysis, Statistics and Probability (physics↗

A cost–benefit framework to evaluate capacity upgrade options in overhead line transmission planning

This paper presents the methodology behind the new Reconductoring Economic and Financial Analysis (REFA) tool, an open-access software, used by transmission utilities to evaluate transmission capacity enhancement options. The proposed methodology is intended to be used in a new planning stage, after the capacity expansion and prior to the individual transmission project engineering, allowing capacity upgrade options (reconductoring, rebuild or voltage upgrade), and respective conductor selection, to be compared under the same economic basis. Furthermore, the REFA tool implements a methodology to rank project options and conductor types based on economic criteria, considering an approximation of the ampacity and sag constraints. Results, using 5 real transmission lines in the US, show that least-cost combinations of project and conductor types can be very diverse, which emphasizes the need for the proposed methodology and tool.

Advanced conductors↗

Forest aboveground biomass estimation through integration of sentinel-2 and PALSAR-2 time series: assessing models trained on GEDI and field inventory benchmarks

Accurate and spatially explicit forest Aboveground Biomass (AGB) mapping through remote sensing is critical for quantifying terrestrial carbon stocks and informing effective forest management strategies. However, AGB estimation in dense forests with complex terrain remains challenging due to satellite sensor signal saturation problem (saturation issue occurs in high biomass forests), structural complexity, and limited ground truth for calibration. This study presents a novel framework that integrates multi-temporal Sentinel-2 optical imagery, ALOS PALSAR-2 Synthetic Aperture Radar (SAR) data, and topographic variables with explainable Machine Learning to map AGB across mountainous forests within subtropical and temperate oceanic climate zones of Mexico. We evaluate the effects of temporal granularity and sensor synergy by comparing multiple temporal inputs and sensor configurations (Sentinel-2, PALSAR-2, and their fusion), and assess model performance using two reference datasets: NASA GEDI LiDAR-derived biomass and Mexico’s National Forest and Soil Inventory (INFyS). Our results showed that models trained on INFyS consistently outperformed those trained on GEDI, highlighting limitations in GEDI’s reliability in biomass estimates within this study region. Furthermore, the integration of Sentinel-2 and PALSAR-2 provided improved predictions compared to single-sensor models, particularly when combined with temporally explicit yearly statistics. The best-performing model, which was trained on INFyS data, and considered both Sentinel-2 and PALSAR-2 yearly statistics, as well as topographic variables, achieved an R2 of 0.64, RMSE of 51.10 Mg/ha, and relative RMSE (rRMSE) of 58.69%. Explainable ML analysis identified Sentinel-2 spectral indices and topographic features as key predictors, while PALSAR-2 metrics provided complementary information, partially mitigating saturation effects in high-biomass areas. Specifically, integrating both sensors substantially improved AGB estimation in high biomass forest (≥200 Mg/ha), yielding 98% gains over optical-only model, with resulting estimates exceeding GEDI L4B by 29% and ESA-CCI-BIOMASS by 174%. Terrain-stratified analysis indicated close agreement with GEDI in low-slope areas, with increasing divergence as slope steepness increased, while estimates remained consistently higher than ESA-CCI-BIOMASS across all slope classes. The proposed approach advances multi-sensor fusion and temporal feature engineering for AGB mapping using open-access satellite datasets, providing a scalable and reproducible framework for annual biomass monitoring in topographically complex mountainous forests. The resulting 25 m resolution biomass product has the potential to provide spatially detailed information for forest monitoring and may support applications in carbon accounting and forest management.

54 ENVIRONMENTAL SCIENCES↗

Quantifying the economic costs of power outages owing to extreme events: A systematic review

Quantifying the economic cost of long-duration power outages is crucial to justifying investments in resiliency and reliability improvements. However, extensive study on the subject complicates the identification of power outage costs and determining the most suitable approach to quantify them for an individual, specific facility, particularly in the context of extreme events. Here, this research provides a systematic review of economic studies estimating the impact of environmental disasters at the microeconomic and macroeconomic levels. Of 326 articles, evaluating the costs of power outages in extreme events, this work identified 22 studies that attempted to quantify the economic costs. These findings indicate that quantifying power outage costs lacks standardization, posing challenges for comparing different studies. Most analyses aiming to quantify these costs for utilities, sectors, and the overall economy rely on outdated survey data, which offer generalized rather than specific cost estimations. The costs of power outages exhibit a significant dependence on factors such as the sector involved, the type of customer affected, and the outage duration. To quantify industry costs, the research in this study suggests that using the National Renewable Energy Laboratory's online, open-access Customer Damage Function Calculator is the best option for individual-level assessments of industries, hospitals, offices, education centers, and similar facilities. However, the Interruption Cost Estimate Calculator can estimate outage costs across industrial, commercial, and residential sectors for macroeconomic outcomes. Finally, this article discusses the relative strengths of these methods and tools and the potential directions for future research.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗