Search NASA⌕ Search

SEARCH · Search NASA

Results for “data access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Material Needs and Measurement Challenges for Advanced Semiconductor Packaging: Understanding the Soft Side of Science

This Perspective builds upon insights from the National Institute of Standards and Technology (NIST)-organized workshop, “Materials and Metrology Needs for Advanced Semiconductor Packaging Strategies,” held at the 35th annual Electronics Packaging Symposium in Binghamton, NY, on September 5, 2024. It outlines critical challenges and opportunities related to polymer-based “soft” materials in advanced semiconductor packaging, with emphasis on polymer science, measurement science (metrology), and the strategic development of Research-Grade Test Materials (RGTMs). These efforts, led by the NIST CHIPS team, aim to advance the fundamental understanding of structure-property-processing relationships, promote standardized guidelines and innovative methods for material characterization, and accelerate the development, qualification, and adoption of next-generation packaging materials. The Perspective also distills key insights from the panel discussion with industry experts, emphasizing the need for close collaboration among materials scientists, process engineers, and metrology experts to enable a holistic strategy, further highlighting the importance of cross-sector partnerships among industry, academia, and government to address pressing challenges in packaging materials and processes.

97 MATHEMATICS AND COMPUTING↗

Proximal remote sensing: an essential tool for bridging the gap between high‐resolution ecosystem monitoring and global ecology

Summary A new proliferation of optical instruments that can be attached to towers over or within ecosystems, or ‘proximal’ remote sensing, enables a comprehensive characterization of terrestrial ecosystem structure, function, and fluxes of energy, water, and carbon. Proximal remote sensing can bridge the gap between individual plants, site‐level eddy‐covariance fluxes, and airborne and spaceborne remote sensing by providing continuous data at a high‐spatiotemporal resolution. Here, we review recent advances in proximal remote sensing for improving our mechanistic understanding of plant and ecosystem processes, model development, and validation of current and upcoming satellite missions. We provide current best practices for data availability and metadata for proximal remote sensing: spectral reflectance, solar‐induced fluorescence, thermal infrared radiation, microwave backscatter, and LiDAR. Our paper outlines the steps necessary for making these data streams more widespread, accessible, interoperable, and information‐rich, enabling us to address key ecological questions unanswerable from space‐based observations alone and, ultimately, to demonstrate the feasibility of these technologies to address critical questions in local and global ecology.

Plant Sciences↗

becquerel (bq) v0.7.0

Becquerel is a Python package for analyzing nuclear spectroscopic measurements. The core functionalities are reading and writing different spectrum file types, fitting spectral features, rebinning spectrum counts to different bin edges, performing detector calibrations and interpreting measurement results. It also includes tools for visualizing radiation spectra and fits of different spectral features, as well as convenient access to tabulated nuclear data both from remote servers and local caches. It relies heavily on the standard scientific Python stack of numpy, scipy, matplotlib, pandas, and numba. It is intended to be general-purpose enough that it can be useful to anyone from an undergraduate taking a laboratory course to the advanced researcher.

Bandstra, Mark [Lawrence Berkeley National Laborat↗

Managing negative values is reservoir inflow computation: A case study

Reservoir inflow is conventionally estimated using the water balance method, which involves the reservoir release and the change in storage during the period considered. As a result, the estimated inflow may sometimes be negative as the errors involved in each input variable build-up to the output. In our study, the fleet data was provided by the Tennessee Valley Authority (TVA) for their Norris Hydropower facility. Unlike the flow release data, which was readily accessible, the change in storage had to be calculated using the reservoir elevation and volume relationship. The original inflow estimates produced a wide range of negative values with large outliers, making it difficult to visualize the current trends. This paper describes a methodology to remove the negative values encountered during the inflow computation, and the results were analyzed by correlating with the nearby streamflow gaging stations.

Shibu, Asha↗

Standardizing Scientific Metadata at Los Alamos National Laboratory

This document is intended for publication in Descriptive Notes, the blog of the Description Section of the Society of American Archivists. This paper focuses on cataloging processes for scientific datasets and the creation of scientific metadata to ensure accessibility of LANL researchers' data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

DER Cybersecurity Standards: Assessment and Gap Analysis

The purpose of this report is to share the comprehensive gap analysis of existing cybersecurity standards applicable to Distributed Energy Resources (DERs) within the electric power sector. This analysis aims to identify critical deficiencies in current standards, assess their alignment with industry needs, and provide actionable recommendations for enhancing cybersecurity measures. The scope encompasses various DER technologies, including solar, wind, energy storage, and hydrogen fuel cells, and emphasizes the significance of establishing robust cybersecurity frameworks and standards to safeguard these increasingly integrated systems. The report provides valuable insights for stakeholders in the DER ecosystem, including manufacturers, utilities, and regulators. It underscores the importance of continued development and refinement of cybersecurity standards to keep up with the technical advances in DERs and associated cybersecurity challenges. The analysis evaluated IEC, IEEE, ISA, ISO, and UL standards relevant to DER cybersecurity. Standards were assessed on their coverage of key requirements including data availability, integrity, confidentiality, access control, authentication, encryption, and system hardening. For each standard, the analysis assessed its alignment with current industry practices, regulatory compliance, effectiveness in addressing known risks, coverage of emerging risks, and how it promotes interoperability. The evaluation also considered potential integration challenges and barriers to adoption.

97 MATHEMATICS AND COMPUTING↗

Data Format and Descriptions for the Alabama Carbon Storage: Data Sharing and Engagement Project

The Alabama Carbon Storage: Data Sharing and Engagement (ACS-DSE) project seeks to develop publicly accessible geologic carbon storage models and data across the southern Gulf Coastal Plain of Alabama. The public online platform developed for this project will include geologic, geophysical, infrastructure, and other relevant datasets and geologic models of the study area. Datasets, model surfaces (e.g. structural contour maps, isolith maps, porosity maps), and infrastructure data (e.g. offshore pipelines, field boundaries) will be downloadable in commonly used file formats. The anticipated primary geologic datasets are well headers, formation tops, average reservoir properties, and core analyses; these will be available as commaseparated values (CSV) text files and MS Excel workbooks. Geophysical logs will be available in Log ASCII Standard (LAS) file format. Modeled surfaces, such as structure contour maps, will be available in ArcGIS formats and text files. Infrastructure data will be available as ArcGIS shapefiles. This document provides information on the data sources and attributes of the datasets.

01 COAL, LIGNITE, AND PEAT↗

Web-accessible sorption database (Interim Progress Report)

This progress report (Level 4 Milestone Number M4SF-26LL010204053) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Crystalline Host Rock Properties & Processes - LLNL Number SF-26LL01020405. This research is focused on the development of a web-accessible database of sorption data for minerals and rocks (bentonite backfill and crystalline host rock) that are likely to be present in a US hosted repository, including elevated temperature data that are expected in a DPC DGR scenario.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability

Overview Code and data repository for NIL Manuscript. Documentation includes sequence processing examples and data analysis. Supplemental sequence processing and R statistical analysis for publication, which compares the microbiome of teosinte-B73 Near Isogenic Lines. Sample Data Amplicon sequence data for 16S rRNA genes, the fungal ITS2 region, and nitrogen-cycling functional genes are available through the NCBI Sequence Read Archive (SRA) under accession number PRJNA1042643(https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1042643). Raw metabolomic data are available on Metabolomics Workbench, Project ID: PR002654. This study is available at the NIH Common Fund's National Metabolomics Data Repository (NMDR) website, the Metabolomics Workbench, https://www.metabolomicsworkbench.org where it has been assigned Study ID ST004211. The data can be accessed directly via its Project DOI: http://dx.doi.org/10.21228/M8KV8T.

Near Isogeneic Lines↗

sdt (Solar Data Tools) [SWR-25-130]

Solar Data Tools (sdt) is an open-source Python library for analyzing PV power (and irradiance) time-series data. It was developed to enable analysis of unlabeled PV data, i.e. with no model, no meteorological data, and no performance index required, by taking a statistical signal processing approach in the algorithms used in the package’s main data processing pipeline. Solar Data Tools empowers PV system fleet owners or operators to analyze system performance a hundred times faster even when they only have access to the most basic data stream—power output of the system.

Meyers-Im, Bennet [National Laboratory of the Rock↗

Unlocking nighttime mobility: Land use and accessibility in public transit for night commuters

Night commuters are integral to urban transportation systems. Essential services such as healthcare and manufacturing rely on workers who travel at night, and reliable mobility options are crucial for them. A gap exists in understanding how land use and accessibility influence public transportation use among night commuters. This study addresses this gap by using public data to explore land use and accessibility factors that affect night commuters' public transportation use in New York State. We investigated (1) the demographic characteristics of night commuters; (2) the influence of land use and accessibility on nighttime public transportation use; and (3) potential improvements to increase public transportation use and their impact. We combined data from the National Household Travel Survey with the Smart Location Database to link home locations with land use characteristics. Using logistic regression, we found that although females are generally less likely to be night commuters, they are more likely to use public transportation. Longer commute distances are associated with higher use of public transportation. Increasing job density along fixed-guideway transit routes and improving overall job accessibility via public transportation significantly enhances public transportation use among night commuters. In conclusion, this research provides actionable insights for public transportation agencies and urban planners to support night commuters, improving access and encouraging nighttime employment.

Job accessibility↗

SARS-CoV-2 wastewater variant surveillance: pandemic response leveraging FDA’s GenomeTrakr network

ABSTRACT Wastewater surveillance has emerged as a crucial public health tool for population-level pathogen surveillance. Supported by funding from the American Rescue Plan Act of 2021, the FDA‘s genomic epidemiology program, GenomeTrakr, was leveraged to sequence SARS-CoV-2 from wastewater sites across the United States. This initiative required the evaluation, optimization, development, and publication of new methods and analytical tools spanning sample collection through variant analyses. Version-controlled protocols for each step of the process were developed and published on protocols.io. A custom data analysis tool and a publicly accessible dashboard were built to facilitate real-time visualization of the collected data, focusing on the relative abundance of SARS-CoV-2 variants and sub-lineages across different samples and sites throughout the project. From September 2021 through June 2023, a total of 3,389 wastewater samples were collected, with 2,517 undergoing sequencing and submission to NCBI under the umbrella BioProject,PRJNA757291. Sequence data were released with explicit quality control (QC) tags on all sequence records, communicating our confidence in the quality of data. Variant analysis revealed wide circulation of Delta in the fall of 2021 and captured the sweep of Omicron and subsequent diversification of this lineage through the end of the sampling period. This project successfully achieved two important goals for the FDA’s GenomeTrakr program: first, contributing timely genomic data for the SARS-CoV-2 pandemic response, and second, establishing both capacity and best practices for culture-independent, population-level environmental surveillance for other pathogens of interest to the FDA. IMPORTANCE This paper serves two primary objectives. First, it summarizes the genomic and contextual data collected during a Covid-19 pandemic response project, which utilized the FDA’s laboratory network, traditionally employed for sequencing foodborne pathogens, for sequencing SARS-CoV-2 from wastewater samples. Second, it outlines best practices for gathering and organizing population-level next generation sequencing (NGS) data collected for culture-free, surveillance of pathogens sourced from environmental samples.

Microbiology↗

Multi-Scale 3D Imaging for Machine Learning Property Upscaling: Mt. Simon Sandstone Case Study

Petrographic properties of principal target reservoirs for carbon sequestration, such as the Mt. Simon Sandstone, are relevant to broad interest groups. The Mt. Simon Sandstone is a deep, saline, regionally extensive Cambrian sandstone, overlain by low permeability sealing formations, making it one of the viable geologic carbon storage reservoirs in the Midwestern US. Its thickness (exceeding 2400 ft in some localities), depth, and lateral extent, combined with high porosity and permeability make it a high-priority target of multiple ongoing geologic carbon sequestration efforts in the United States of America. The National Energy Technology Laboratory in Morgantown, West Virginia, has been engaged in characterization efforts of the Mt. Simon for over a decade, with a strong focus on Computed Tomographic data acquisition. Data generated during this period has been hitherto not accessible to the public. This archival effort focused on preservation of historical CT data and associated metadata, and facilitating their accessibility, culminating with the publication of the entire dataset on NETL’s Energy Data eXchange (EDX) and the associated Gill et. al (2024) paper.

Gill, Magdalena K.↗

Characterization of throughput on the AXI DMA bus for burst data transfer over Ethernet

cThe Xilinx AXI Direct Memory Access (AXI DMA) module is an efficient solution for medium-speed data transfer in Xilinx SoC FPGAs, supporting data rates greater than 1000 Gbps even in very suboptimal operating modes. It facilitates direct transfer of AXI stream data into processor memory without constant software intervention, which reduces overhead and ensures consistent data logging. By utilizing the FPGA's available memory, large circular buffers (1-5 GiB) are used to buffer data and accommodate network limitations, enabling high-rate data bursts. In this study, we measured the performance of AXI DMA under conditions simulating its lowest practical data transfer speeds. The Arbitrary Length Data Sender was used to transmit AXI stream packets at 32-bit width and 100 MHz frequency, a narrow width and slow speed. Results show that the AXI DMA can transfer up to 3192.76 Mbps with large packet sizes but experiences reduced performance for smaller packets, as low as 2.6 Mbps for 4-byte packets. For Ethernet-limited applications, packet sizes between 8,000 and 16,000 bytes provided optimal transfer speeds of 874 to 1600 Mbps. These findings suggest that the AXI DMA is not the limiting factor in systems where packet sizes exceed 8,000 bytes.

43 PARTICLE ACCELERATORS↗

Shared and Ownership Mobility Technologies in the US: Data Availability and Usage Trends

This report supports the vision for a more sustainable transportation future by summarizing and analyzing the latest data on new mobility technologies, including ridesharing, shared and privately owned bikes, e-bikes, and scooters that have emerged over the past two decades. Having access to accurate and current data that is representative of new mobility systems and individual usage of these systems across different parts of the country is critical for researchers, city and regional planning professionals, and current and potential industry technology developers to better understand and forecast usage trends both nationwide as well as across different existing and potential future markets across the country. Building on the previous study published in 2022, this report incorporates the latest available market and usage data on new mobility technologies and compares usage by Chicago and New York City demographic characteristics. Moreover, this report includes recent developments and insights on privately owned micromobility technologies. Our analysis found that more downtown areas in Chicago show high per capita usage for all three modes than in the previous study, likely due to the full launch of shared e-scooter systems citywide in 2022. Notably, the majority of high shared mobility usage is concentrated in high-income, densely populated downtown areas in Chicago, which also have good public transit access. In contrast, TNC and bikeshare usage hotspots in central Manhattan are more widely distributed, though also appear to be shaped by the geography of the public transit system. Analysis of privately owned micromobility shows that the greatest energy savings occurred when e-bikes replaced single-occupancy vehicle (SOV) trips (i.e., gasoline-powered cars driven alone). Based on the literature review and analysis results, we also make recommendations for supporting the development of both shared and privately owned micromobility programs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗