Search NASASearch

SEARCH · Search NASA

Results for “Open Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Brain‐age prediction: Systematic evaluation of site effects, and sample age range and size

Abstract Structural neuroimaging data have been used to compute an estimate of the biological age of the brain (brain‐age) which has been associated with other biologically and behaviorally meaningful measures of brain development and aging. The ongoing research interest in brain‐age has highlighted the need for robust and publicly available brain‐age models pre‐trained on data from large samples of healthy individuals. To address this need we have previously released a developmental brain‐age model. Here we expand this work to develop, empirically validate, and disseminate a pre‐trained brain‐age model to cover most of the human lifespan. To achieve this, we selected the best‐performing model after systematically examining the impact of seven site harmonization strategies, age range, and sample size on brain‐age prediction in a discovery sample of brain morphometric measures from 35,683 healthy individuals (age range: 5–90 years; 53.59% female). The pre‐trained models were tested for cross‐dataset generalizability in an independent sample comprising 2101 healthy individuals (age range: 8–80 years; 55.35% female) and for longitudinal consistency in a further sample comprising 377 healthy individuals (age range: 9–25 years; 49.87% female). This empirical examination yielded the following findings: (1) the accuracy of age prediction from morphometry data was higher when no site harmonization was applied; (2) dividing the discovery sample into two age‐bins (5–40 and 40–90 years) provided a better balance between model accuracy and explained age variance than other alternatives; (3) model accuracy for brain‐age prediction plateaued at a sample size exceeding 1600 participants. These findings have been incorporated into CentileBrain ( https://centilebrain.org/#/brainAGE2 ), an open‐science, web‐based platform for individualized neuroimaging metrics.

60 APPLIED LIFE SCIENCES

An affordable platform for automated synthesis and electrochemical characterization

In recent years, self-driving laboratories (SDLs) have emerged as a powerful tool to expedite various areas of chemical research. For optimal functionality, these laboratories must be adaptable, readily modifying configurations to meet researchers' specific needs. Despite these advances, much of chemistry still depends on proprietary equipment from specialized vendors, which can be restrictive and difficult to customize for diverse lab setups. Moreover, ensuring reproducibility requires full disclosure of equipment details. In this work, we introduce an automated system featuring a cost-effective, self-designed potentiostat and a straightforward synthesis platform. We provide complete transparency by disclosing the electronic schematics of the potentiostat and the software used in the system. Our aim is to reduce the barriers to entry for SDLs and promote the principles of open science.

Pablo-García, Sergio

Energy density driven ultrafast electronic excitations in a cuprate superconductor

Controlling nonequilibrium dynamics in quantum materials requires ultrafast probes with spectral selectivity. We report femtosecond reflectivity measurements on the cuprate superconductor Bi 2⁢ Sr 2⁢ CaCu 2 ⁢O 8+𝛿 using free-electron laser extreme-ultraviolet (23.5–177 eV) and near-infrared (1.5 eV) pump pulses. EUV pulses access deep electronic states, while NIR light excites valence-band transitions. Despite these distinct channels, both schemes produce nearly identical dynamics: above 𝑇 𝑐 , excitations relax through fast (100–300 fs) and slower (1–5 ps) channels; below 𝑇 𝑐 , a delayed component signals quasiparticle recombination and condensate recovery. We find that when electronic excitations are involved, the ultrafast response is governed mainly by absorbed energy rather than by the microscopic nature of the excitation. In contrast, bosonic driving in the THz or midinfrared produces qualitatively different dynamics. By demonstrating that EUV excitation of a correlated superconductor yields macroscopic dynamics converging with those from optical pumping, this work defines a new experimental paradigm: FEL pulses at core-level energies provide a powerful means to probe and control nonequilibrium electronic states in quantum materials on their intrinsic femtosecond timescales. This establishes FEL-based EUV pumping as a new capability for ultrafast materials science, opening routes toward soft x-ray and attosecond studies of correlated dynamics.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono

2024 NMDC Ambassador Training Materials [Slides]

The NMDC is a sustainable data discovery platform that promotes open science and shared-ownership across a broad and diverse community of researchers, funders, publishers, societies, and other collaborators. The NMDC aims to enable multi-omic microbiome research to accelerate scientific discovery. The NMDC is a Department of Energy funded program that is a collaboration between 3 National Laboratories: Lawrence Berkeley National Laboratory (LBNL), Los Alamos National Laboratory (LANL), and Pacific Northwest National Laboratory (PNNL).

54 ENVIRONMENTAL SCIENCES

A Data-Driven Approach to Real-World Degradation of Backsheets

The objectives of this project are as follows: • The population behavior of fielded modules in various conditions of use • Predictions of materials in specific climatic zones • Understanding of a module’s local environment in the field on its degradation It aims to understand how and what backsheet materials of photovoltaic modules degrade in the real world field. • Field Survey Protocol This project is started from the protocol, because all the data, information, domain knowledge are from experience of the real world field surveys, which is based on the protocol. The Protocol is explored from the experience of the field survey observations. With the increasing of the field surveys, it is refined for three versions, which are Task 1.0 (Section 3.1), Task 6.0 (Section 3.6), Task 10.0 (Section 3.10) respectively. It includes a document and a training video, which is able to direct other teams to follow the same procedure with the sites surveyed during this project. The documents include the detailed information but is not limited to the terminology definition, instruments SOP, preparation items for the surveys, form for the data collections, the order of the information collection. The final version of the protocol can be found at Appendix A, see Section 5. Additional, It can also be found at Open Science Framework (OSF), see Section 3.18 for detail. • Written Waiver Request Before staring the field surveys, the request of the waiver for international surveys is completed, because the limitation of climate zone in the United States, see Appendix B in Section 6 for the request documents. Unfortunately, only 1 international site from Taiwan, China can be finished, due to the COVID-19. • Field Survey According to the protocol we built in Section 3.1, 3.6 and 3.10, 41 sites have been surveyed across seven different climate zones (Cfa, Csa, Csb, BSk, Dfa, Dfb, Am). A variety of materials, including Polyethylene Naphthalate (PEN), Polyethylene Terephthalate (PET), Polyvinyl Fluoride (PVF), Polyvinylidene Fluoride (PVDF), Acrylic PVDF, Fluoroethylene Vinyl Ether (FEVE), and Glass, were identified. These sites are located in various states including California, South Carolina, New Mexico, Maryland, Ohio, Tennessee, Florida, Massachusetts, Illinois, Minnesota, Oregon, Colorado, and Taiwan, Republic of China. The ages of the sites ranged from 2 - 38 years in service and the field size varied from 1 MW - 25 MW. All requirements for the modeling have been satisfied. Some observations like ’Edge Effect’ for the rows and Junction box heating will also be a useful knowledge to build the model. Section 3.7 provides detailed information on the sites visited during this reporting period.

14 SOLAR ENERGY

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING

DoE as a “Digital Innovation” Sponsor of the WCRP OSC2023 (Final Report)

The WCRP Open Science Conference (https://wcrp-osc2023.org/) was a once-in-a-decade opportunity to jointly explore the transformative actions urgently needed to ensure a sustainable future. Held in Kigali, Rwanda on October 23 -27, 2023, it showcased advances in climate science, helped identify gaps and opportunities, and provided a forum for communities to jointly develop future activities. Scientists, practitioners, politicians, policy makers, intergovernmental agencies and NGOs showcased their work, learned from each other, and explored new ways to work together.

54 ENVIRONMENTAL SCIENCES

Multi-system analysis of offshore geologic carbon storage: a review of open-source data science solutions

Geologic carbon storage projects are maturing worldwide and the footprint of deployment in the offshore is expanding. At present, there are ten projects in operation or that have been completed, more than 50 in construction and development, and dozens of characterization studies completed or underway. Offshore geologic carbon storage offers potential benefits over onshore geologic carbon storage. These offshore projects are generally remote in location, distant from population centers, and avoid complicated pore space rights while having abundant prospective storage potential. Some offshore fields targeted for carbon storage have comparatively fewer prior borehole penetrations except for areas that have been explored for petroleum production, minimizing potential issues such as pressure interference and infrastructure impacts. Yet offshore geologic carbon storage projects face distinctive technical and economic challenges, such as seafloor geohazards (e.g., seabed instability), expensive maritime transport, and meteorological-oceanographic conditions that can damage infrastructure and impact operations. Analytical capabilities and improved computational speeds have advanced engineering, earth and energy sciences in the wake of the arrival of modern data science over the last decade. These advancements have created an opportunity for integrated, multi-systems modeling approaches utilizing artificial intelligence and machine learning that are no longer limited by computational issues. Analytical tools developed alongside this advancement in data science can be leveraged to calibrate the potential advantages and challenges of carbon storage operations in the offshore. New methods and approaches that incorporate data science to analyze multiple aspects of engineered and natural systems can provide insights that complement the characterization and onsite engineering that traditional commercial and operational software addresses. These new methods and approaches can potentially improve the outcome of energy operations and carbon storage. Providing multi-system, science-driven data analytics enhances the knowledge base that offshore developers, operators, and regulatory bodies may draw from to improve offshore site selection and operational efficiency. Here, we provide a brief synopsis of geologic carbon storage efforts to date, an overview of the engineered and natural systems involved in offshore geologic carbon storage, and a review of publicly available, open-source, offshore and/or carbon storage related data- and science-driven tools developed by 2010 or later that are suitable for screening and assessing regions for offshore geologic carbon storage.

artificial intelligence

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES

gRASPA

GPU Monte Carlo Simulation Code with a taste of RASPA We present enhancements in Monte Carlo simulation speed and functionality within an open-source code, gRASPA, which uses graphical processing units (GPUs) to achieve significant performance improvements compared to serial, CPU implementations of Monte Carlo. The code supports a wide range of Monte Carlo simulations, including canonical ensemble (NVT), grand canonical, NVT Gibbs, Widom test particle insertions, and continuous-fractional component Monte Carlo. Implementation of grand canonical transition matrix Monte Carlo (GC-TMMC) and a novel feature to allow different moves for the different components of metal-organic framework (MOF) structures exemplify the capabilities of gRASPA for precise free energy calculations and enhanced adsorption studies, respectively. The introduction of a High-Throughput Computing (HTC) mode permits many Monte Carlo simulations on a single GPU device for accelerated materials discovery. The code can incorporate machine learning (ML) potentials. The open-source nature of gRASPA promotes reproducibility and openness in science, and users may add features to the code and optimize it for their own purposes. The code is written in CUDA/C++ and SYCL/C++ to support different GPU vendors. The gRASPA code is publicly available at https://github.com/snurr-group/gRASPA.

Li, Zhao [Purdue/Northwestern/Notre Dame Universit

Unlocking the benefits of transparent and reusable science for climate-risk management

People around the world seek climate-risk information to guide their decisions. For instance, projections about future flood risk inform where households choose to live, how lenders manage credit risks, and which communities receive federal funding. Yet data limitations and fundamental validation challenges raise important concerns about the reliability of such projections. The principles of transparency and reusability help address these concerns by enabling scrutiny of assumptions and methods, development of foundational data and tools, and consistent application of evaluation standards. While there is ongoing debate about how much transparency commercial climate-risk services should provide, many expect non-commercial actors to lead the way on operationalizing transparency and reusability to fulfill their knowledge-building role in the climate-risk ecosystem. However, despite prominent success stories, we find a substantial gap between principles and practice: only four percent of the most-cited peer-reviewed climate-risk studies in recent years fully share their data and code despite this being a widely accepted minimum standard for transparency. We highlight low-cost measures that non-commercial researchers can take now to improve transparency and reusability. We also emphasize that transformative progress requires substantial investment, cross-sector collaboration, and careful consideration of tradeoffs, data rights, and multiple perspectives on equity. We hope this perspective accelerates both immediate actions and longer-term conversations to improve the ability of science to effectively support timely, evidence-based, and sound climate-risk management.

Open Science

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

MSD CoP Webinar: "Advances in MSD-LIVE to Support the MSD Community of Practice"

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Advances in MSD-LIVE to Support the MSD Community of Practice Presenters: Casey Burleyson and Zoe Guillen (Pacific Northwest National Laboratory) Abstract: The MultiSector Dynamics Living, Intuitive, Value-adding, Environment (MSD-LIVE; msdlive.org) is a cloud-based data management system and advanced computing platform that enables MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and workflows within the MSD Community of Practice. Recently, several high-profile datasets have attracted many new users to MSD-LIVE. This webinar has two goals: 1) To refamiliarize the MSD community and new users with the components of the platform (e.g., the data repository, model training notebooks, and data dashboards) and to highlight examples of how these components are advancing MSD science and 2) To demonstrate new features in v3 of the platform, released in late 2025. The main new feature in v3 is the ability to interactively explore data in MSD-LIVE without downloading it. MSD-LIVE users can now click a button in our data repository and launch a blank Jupyter notebook with access to the underlying data on AWS. Users can use the notebook to write analysis, visualization, or subsetting routines that process the data directly on the AWS cloud. We also added a GitHub integration feature that allows users to share analysis or visualization code they develop with the community of MSD-LIVE users. The webinar will wrap up with a look at what's coming next for MSD-LIVE in 2026. Moderator: Patrick M. Reed (MSD CoP Facilitation Team) This webinar was held on: May 12th, 2026 from 1-2 PM EST.

Open Science

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES