Search NASASearch

SEARCH · Search NASA

Results for “Open Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

OPEN-Augmented Reality GUI for Bioenergy Crop Phenotyping and Precision Agriculture (Donald Danforth Plant Science Center Final Scientific Technical Report)

The project led by the Donald Danforth Plant Science Center, in collaboration with Arizona State University, George Washington University, and Saint Louis University, has made significant strides in advancing the phenotypic analysis of bioenergy crops through the development of an innovative AI processing pipeline. This initiative was primarily funded by ARPA-E, with additional cost-sharing provided by the participating institutions. The project successfully utilized a variety of sensors—3D scanners, thermal, RGB, and hyperspectral—to refine algorithms for data-driven trait signature identification and improve the classification and visualization of plant traits. The developed AI processing pipeline is capable of handling the complex, multidimensional data characteristic of dynamic agricultural environments. 1) Contributions to understanding: The research has advanced the field of plant phenomics by showcasing the synergistic use of various sensor data to enhance the precision of trait analysis in bioenergy crops. Through the integration of 3D scanners, thermal, RGB, and hyperspectral sensors, the project has developed robust data-driven trait signature algorithms and visualization techniques. These innovations have facilitated detailed monitoring and management of plant traits, providing vital insights into plant growth dynamics and stress responses. Further, the project has broadened our understanding of how machine learning can be effectively applied in multi-sensor environments to refine trait analysis. By leveraging diverse datasets, the research has not only improved the accuracy of phenotypic assessments but also established a versatile methodological framework that can be extended beyond agriculture to other fields requiring detailed phenotypic analysis. 2) Technical effectiveness and economic feasibility: The AI processing pipeline developed in this project demonstrated significant technical effectiveness, achieving high throughput analysis of extensive phenotypic data and meeting targeted accuracies. This system exemplified the capability of advanced machine learning technologies to efficiently manage and analyze large, complex datasets. Economically, the implementation of the project-developed pipelines may offer substantial cost savings across multiple sectors. It enhances data analysis processes and significantly reduces the need for manual data interpretation, thereby decreasing both the time and resources required. 3) Public benefit: The project has significantly broadened the scope of agricultural methodologies to enhance phenotypic analysis, with potential applications in various sectors beyond agriculture. Additionally, the initiative fostered an enriching educational and collaborative environment, significantly enhancing the technical skills of participants. It also made substantial contributions to the scientific community by providing open-access data sets and tools, encouraging ongoing research and development across various disciplines. Overall, the project not only met its scientific goals but also showcased the extensive utility of integrating advanced machine learning and sensor data analysis technologies. These advancements have proven instrumental in driving forward both theoretical research and practical applications, setting a strong foundation for future explorations and innovations in data-driven science.

60 APPLIED LIFE SCIENCES

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product

An MLCommons Scientific Benchmarks Ontology

Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper introduces an ontology for scientific benchmarking developed through a unified, community-driven effort that extends the MLCommons ecosystem to cover physics, chemistry, materials science, biology, climate science, and more. Building on prior initiatives such as XAI-BENCH, FastML Science Benchmarks, PDEBench, and the SciMLBench framework, our effort consolidates a large set of disparate benchmarks and frameworks into a single taxonomy of scientific, application, and system-level benchmarks. New benchmarks can be added through an open submission workflow coordinated by the MLCommons Science Working Group and evaluated against a six-category rating rubric that promotes and identifies high-quality benchmarks, enabling stakeholders to select benchmarks that meet their specific needs. The architecture is extensible, supporting future scientific and AI/ML motifs, and we discuss methods for identifying emerging computing patterns for unique scientific workloads. The MLCommons Science Benchmarks Ontology provides a standardized, scalable foundation for reproducible, cross-domain benchmarking in scientific machine learning. A companion webpage for this work has also been developed as the effort evolves: https://mlcommons-science.github.io/benchmark/

Hawks, Ben [Fermilab] (ORCID:0000000157000288)

An Open Benchmark of One Million High-Fidelity Cislunar Trajectories

Cislunar space spans from geosynchronous altitudes to beyond the Moon and will underpin future exploration, science, and security operations. We describe and release an open dataset of one million numerically propagated cislunar trajectories generated with the open-source Space Situational Awareness Python package (SSAPy). The model includes high-degree Earth/Moon gravity, solar gravity, and Earth/Sun radiation pressure; other planetary gravities are omitted by design for computational efficiency. Initial conditions uniformly sample commonly used osculating-element ranges, and each trajectory is propagated for up to six years under a single, fixed start epoch. The dataset is intended as a reusable benchmark for method development (e.g., space domain awareness, navigation, and machine-learning pipelines), a reference library for statistical studies of orbit families, and a starting point for community-driven extensions (e.g., alternative epochs). We report empirically observed stability trends (e.g., a band near ~5 GEO and persistence of some co-orbital classes including L4/L5 librators) as dataset descriptors rather than new dynamical results. The chief contribution is the scale, fidelity, organization (CSV/HDF5 with full state time series and metadata), and open availability, which together lower the barrier to comparative and data-driven studies in the cislunar regime.

79 ASTRONOMY AND ASTROPHYSICS

Can protein expression be ‘solved’?

Recombinant protein expression is central to biotechnology’s application in academic exploration as well as human health, climate applications and the bioeconomy in general. However, not all proteins can be expressed in all organisms, and the field lacks a predictive model of soluble protein overexpression that could replace laborious experimental trial-and-error. Here, we discuss the state of the field and identify the lack of large, high-fidelity datasets as the primary bottleneck to progress. We review possible assays that could be used for data collection to identify a path toward an extensible experimental platform for collecting soluble recombinant protein overexpression data across organisms. We suggest that the resulting dataset should be used to train increasingly generalizable predictive models of protein expression to answer the question: “How can predictive protein expression be solved?”.

59 BASIC BIOLOGICAL SCIENCES

Design and Analysis of the Open-Surface Slow Li Flow Divertor and Comparison to the Fast Li Flow Divertor

In the ongoing U.S. project, “Liquid Metal Plasma Facing Components,” sponsored by the U.S. Department of Energy, efforts have been taken to develop two open-surface divertor designs for the Fusion Nuclear Science Facility using liquid lithium (Li) as a heat and particle flux removal media. The main focus of this study is the design and analysis of a slow (~1 mm/s) and thin (<1 mm) open-surface Li flow divertor with a Li-cooled substrate, which is then compared with an earlier design of a fast (up to 10 m/s) and thick (~0.5 cm) Li flow divertor with the substrate cooled with helium. The slow Li flow divertor design is based on the original LiWall concept developed at the Princeton Plasma Physics Laboratory. Such a thin and slow Li layer can remove the particle flux by reducing the recycling flux, while the heat flux is removed mainly through the heat sink located beneath. In the present study, the heat sink is provided through a Li cooling flow inside the substrate of reduced activation ferritic/martensitic steel. By performing a multiphysics analysis with COMSOL that included liquid-metal magnetohydrodynamics (MHD), heat transfer, and structural mechanics, the impact of various factors on the divertor heat removal capability, such as Li flow velocity, MHD effects, and inlet velocity boundary condition, were examined. Based on comparisons of the two divertor designs, it was shown that the fast-flow divertor significantly outperformed the slow-flow design, whose heat removal capability was limited to ~1 to 2 MW/m2.

Jiang, Yuchen [ORNL] (ORCID:0000000283280284)

Polyhalohydrins: Investigating Vicinal Functionalities by Ring-Opening of Epoxides on Polyolefins

Halohydrins are functional groups that are underexplored in polymer science. This study synthesized and examined structure–property relationships of halohydrin-functionalized polymers derived from polycyclooctene (PCOE), which was obtained via ring-opening metathesis polymerization (ROMP). The polyhalohydrins, including polychlorohydrin, polybromohydrin, and polyiodohydrin, were produced through a two-step process involving epoxidation of PCOE followed by epoxide ring-opening using hydrochloric, hydrobromic, or hydroiodic acid. These stereoirregular and regioirregular polymers are amorphous as measured by differential scanning calorimetry and X-ray scattering. Lap joint shear testing revealed an enhanced adhesive performance in polychlorohydrin and polybromohydrin, with the former demonstrating nearly three times the ultimate shear stress compared to a model polyethylene, hydrogenated PCOE. Additionally, contact angle measurements and surface free energy analysis showed an increase in hydrophilicity and polarity from iodine- to bromine- to chlorine-functionalized polyhalohydrins, aligning with trends in adhesion strength. Furthermore, these results underscore the potential of halohydrin functionalization as a versatile approach for tuning surface properties, offering new opportunities for polymer-to-polymer transformations.

Functional groups

Neuromorphic heat transport effects in a molecular junction

Understanding energy transport at the nanoscale is an open and fundamental challenge in the molecular sciences with direct implications for the design of new electronics, computing devices, and materials. While nanoscale energy transport under steady-state conditions has been studied extensively, there is much less known about energy transport under time-dependent driving forces, particularly in the far-from-equilibrium regime. In this work, we use nonequilibrium molecular dynamics simulations and stochastic thermodynamics to investigate energy transport in a well-studied nanoscale system—a molecular junction—subjected to a time-periodic temperature gradient. The primary observation is that molecular junctions can exhibit heat transport hysteresis, a phenomenon in which the heat flux through a system depends not only on the instantaneous value of a time-dependent temperature bias but also on the temporal history of that bias. The presented findings illustrate that molecular junctions can exhibit the specific memory effect—heat transport hysteresis—that is essential for the design of thermal neuromorphic computers. This work elucidates a potential pathway toward the realization of such devices.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Tractometry of the Human Connectome Project: resources and insights

The Human Connectome Project (HCP) has become a keystone dataset in human neuroscience, with a plethora of important applications in advancing brain imaging methods and an understanding of the human brain. We focused on tractometry of HCP diffusion-weighted MRI (dMRI) data. We used an open-source software library (pyAFQ; https://yeatmanlab.github.io/pyAFQ) to perform probabilistic tractography and delineate the major white matter pathways in the HCP subjects that have a complete dMRI acquisition (n = 1,041). We used diffusion kurtosis imaging (DKI) to model white matter microstructure in each voxel of the white matter, and extracted tract profiles of DKI-derived tissue properties along the length of the tracts. We explored the empirical properties of the data: first, we assessed the heritability of DKI tissue properties using the known genetic linkage of the large number of twin pairs sampled in HCP. Second, we tested the ability of tractometry to serve as the basis for predictive models of individual characteristics (e.g., age, crystallized/fluid intelligence, reading ability, etc.), compared to local connectome features. To facilitate the exploration of the dataset we created a new web-based visualization tool and use this tool to visualize the data in the HCP tractometry dataset. Finally, we used the HCP dataset as a test-bed for a new technological innovation: the TRX file-format for representation of dMRI-based streamlines. We released the processing outputs and tract profiles as a publicly available data resource through the AWS Open Data program's Open Neurodata repository. We found heritability as high as 0.9 for DKI-based metrics in some brain pathways. We also found that tractometry extracts as much useful information about individual differences as the local connectome method. We released a new web-based visualization tool for tractometry—“Tractoscope” (https://nrdg.github.io/tractoscope). We found that the TRX files require considerably less disk space-a crucial attribute for large datasets like HCP. In addition, TRX incorporates a specification for grouping streamlines, further simplifying tractometry analysis.

59 BASIC BIOLOGICAL SCIENCES

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE

Three-dimensional modeling of hyphal fusion, branching, and nutrient transport in filamentous fungi

Fungi exhibit behaviors distinct from other microbes. Filamentous fungi grow by extending complex networks of branched filaments collectively referred to as the mycelium. These networks can expand over large distances and traverse low-nutrient areas by translocating nutrients through the filament network. This spatial characteristic makes filamentous fungi crucial for soil ecosystems, supporting stable microbial communities and promoting plant growth. However, simulating these behaviors is complex. The elongated nature of fungal compartments results in different mechanical interactions compared to the commonly modeled spherical bacteria. These detailed hyphal mechanics require specialized consideration and are often excluded from conventional fungal simulation packages. Additionally, the extensive fungal networks in nature demand computationally intensive simulations, necessitating high-performance algorithms. Therefore, realistic fungi simulations require specialized software. Here, we introduce a fungal modeling expansion to the high-performance biological modelling and interface exchange (bmx) software suite. bmx leverages adaptive mesh refinement in AMReX for chemical diffusion and incorporates a full mechanical model for bacterial cells, accelerated by GPUs. By extending bmx to model filamentous particles, we demonstrate the formation of complex filament networks through interactions like hyphal branching and fusion (anastomosis). We show that the networks produced match real-world fungal structures through various metrics. This work supports computational studies of fungal growth dynamics and can be adapted to investigate the growth of other filamentous structures in biology or materials science. The expanded-BMX package is open-sourced and is available online.

Cell mechanics

Development of PNNL’s Plutonium Metallography Capability

The goal of this project was to develop plutonium metallography capabilities and establish workflows for characterization at PNNL’s Radiological Processing Facility (RPL). Developing and expanding on this capability opens more opportunities for PNNL to better support science through collaborations with other labs, plants, and sites withing the Department of Energy Complex and through programs within the National Nuclear Security Administration. This project established metallography equipment in an air glovebox, polishing protocols, as well as radiological and facility controls and procedures so that gram quantities of plutonium metal could be analyzed on multiple characterization tools. Data was collected and analyzed on two different delta phase plutonium - Gallium samples and a calciothermically reduced alpha plutonium metal using a combination of optical and electron microscopy, powder X-ray diffraction, and atom probe tomography. This report highlights the metallography capabilities, some preliminary data collections, and plans at PNNL to support plutonium material science.

36 MATERIALS SCIENCE

Muon g-2 Experiment

The Muon g-2 experiment probes the magnetic properties of the muon—a heavy sibling of the electron—in the search for new physics. Muon g-2 (pronounced gee-minus-two) is an international collaboration between Fermilab and dozens of labs and universities in seven countries. The Muon g-2 experiment studies the precession or “wobble” of muons when placed in a magnetic field. Based on what we already know about muons and other particles, scientists can predict with great precision the value of the muons’ wobble. If the experiment comes up with something different, it means that our current understanding of physics is incomplete, and it may indicate the presence of additional particles or hidden subatomic forces. It would open the door to exciting new realms of science.

Fermilab, Fermilab

A Microservices Architecture Toolkit for Interconnected Science Ecosystems

Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for inte- grating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architec- ture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability develop- ment, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.

Brim, Michael

Ptychography at all wavelengths

Ptychography is a computational imaging technique that operates across multiple wavelength regimes, from electron (picometres) to X-ray (~0.1 nm), extreme ultraviolet (~10 nm) and visible light (micrometres). By reconstructing both amplitude and phase from diffraction patterns, ptychography enables high-resolution, quantitative imaging without conventional limitations imposed by lens-based optics. Ptychography has enabled advances across a range of scales: achieving deep-sub-angstrom resolution with electron microscopy, becoming an indispensable tool at X-ray synchrotron facilities worldwide and overcoming the trade-offs between resolution and field-of-view in optical imaging. This Primer provides a unified treatment of ptychography across these wavelength regimes. First, we discuss theoretical foundations, reconstruction algorithms, experimental considerations and wavelength-specific challenges. We then give examples of raw and processed data from various configurations and wavelengths. Next, we highlight key applications of ptychography in life sciences, materials science and industry. We also discuss data standards, open-source software implementations and best practices for ensuring reproducibility across different wavelength regimes. Finally, we consider limitations and future opportunities for ptychography. Together with accompanying datasets and code implementations, this Primer aims to serve newcomers and experienced practitioners in the field, facilitating broader adoption of ptychography across different disciplines.

47 OTHER INSTRUMENTATION

poppler-science

The “Poppler-science” software is a fork of the existing open-source Poppler project (https://poppler.freedesktop.org/) for converting PDF files to text. Modifications to the Poppler source code include (a) per-glyph optical character recognition (for correcting the non-standard font glyph remapping that is common in the scientific literature), (b) inference of text markup for commonly used scientific formatting (like superscripts and subscripts), (c) table and figure recognition, and (d) improved ordering of text output for complex scientific manuscripts (e.g., multi-column text, figure and table captions, etc.).

Gans, Jason [Los Alamos National Laboratory]

Toward the Neutrino Discovery Platform: An Auditable, Uncertainty-Bearing Toolchain for MINERvA Open-Data Cross-Section Analysis

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure