Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Importance and Incorporation of User Feedback in Data Stewardship

Since August 1994, The National Aeronautics and Space Administration's (NASA's) Earth Observation System Data and Information System (EOSDIS) has been serving a global community of over 3 million users with Earth science data in a variety of disciplines. During the entire life of EOSDIS, various mechanisms for user feedback have been extremely important and valuable to its evolution and proven performance. Some inputs from user groups have resulted in fundamental changes in the architecture, design and operations of EOSDIS, while others have provided ideas for incremental changes. The purpose of this paper is to share this experience and the benefits that have resulted from the user feedback.In early to mid-1990s, the EOSDIS Advisory Panel (a.k.a. Data Panel) provided significant inputs for the architecture and design of EOSDIS resulting in NASA's establishment of a set of Distributed Active Archive Centers (DAACs) and development of a "working prototype with operating elements" called Version 0 EOSDIS, which went into operation in August 1994. The Data Panel also participated in many of the requirement and design reviews and influenced the design of EOSDIS through their comments.In 1995, the U.S. National Research Council's Committee on Global Change Research conducted a review of the U.S. Global Change Research Program and NASA's Mission to Planet Earth/EOS, including the plans for EOSDIS. One of this committee's recommendations was that the "Responsibility for product generation and publication and for user services should be transferred to a federation of partners selected through a competitive process open to all". In response, NASA initiated an experiment with a "self-governing" federation called the Working Prototype ESIP (WP-ESIP) Federation. This federation, with support from NASA, NOAA and USGS, has now grown into the Earth Science Information Partners (ESIP) with over 140 member organizations.The EOSDIS DAACs' User Working Groups (UWGs) represent broad user communities served by the respective DAACs. As regular users of the DAACs and experts in their scientific disciplines, the UWG members provide valuable inputs for planning and prioritizing the services as well as addition of new datasets for the benefit of the community.

Research Integrity↗

Genelab: Scientific Partnerships and an Open-Access Database to Maximize Usage of Omics Data from Space Biology Experiments

NASA's mission includes expanding our understanding of biological systems to improve life on Earth and to enable long-duration human exploration of space. The GeneLab Data System (GLDS) is NASA's premier open-access omics data platform for biological experiments. GLDS houses standards-compliant, high-throughput sequencing and other omics data from spaceflight-relevant experiments. The GeneLab project at NASA-Ames Research Center is developing the database, and also partnering with spaceflight projects through sharing or augmentation of experiment samples to expand omics analyses on precious spaceflight samples. The partnerships ensure that the maximum amount of data is garnered from spaceflight experiments and made publically available as rapidly as possible via the GLDS. GLDS Version 1.0, went online in April 2015. Software updates and new data releases occur at least quarterly. As of October 2016, the GLDS contains 80 datasets and has search and download capabilities. Version 2.0 is slated for release in September of 2017 and will have expanded, integrated search capabilities leveraging other public omics databases (NCBI GEO, PRIDE, MG-RAST). Future versions in this multi-phase project will provide a collaborative platform for omics data analysis. Data from experiments that explore the biological effects of the spaceflight environment on a wide variety of model organisms are housed in the GLDS including data from rodents, invertebrates, plants and microbes. Human datasets are currently limited to those with anonymized data (e.g., from cultured cell lines). GeneLab ensures prompt release and open access to high-throughput genomics, transcriptomics, proteomics, and metabolomics data from spaceflight and ground-based simulations of microgravity, radiation or other space environment factors. The data are meticulously curated to assure that accurate experimental and sample processing metadata are included with each data set. GLDS download volumes indicate strong interest of the scientific community in these data. To date GeneLab has partnered with multiple experiments including two plant (Arabidopsis thaliana) experiments, two mice experiments, and several microbe experiments. GeneLab optimized protocols in the rodent partnerships for maximum yield of RNA, DNA and protein from tissues harvested and preserved during the SpaceX-4 mission, as well as from tissues from mice that were frozen intact during spaceflight and later dissected on the ground. Analysis of GeneLab data will contribute fundamental knowledge of how the space environment affects biological systems, and as well as yield terrestrial benefits resulting from mitigation strategies to prevent effects observed during exposure to space environments.

bioinformatics↗

GeneLab: Scientific Partnerships and an Open-Access Database to Maximize Usage of Omics Data from Space Biology Experiments

NASA's mission includes expanding our understanding of biological systems to improve life on Earth and to enable long-duration human exploration of space. The GeneLab Data System (GLDS) is NASAs premier open-access omics data platform for biological experiments. GLDS houses standards-compliant, high-throughput sequencing and other omics data from spaceflight-relevant experiments. The GeneLab project at NASA-Ames Research Center is developing the database, and also partnering with spaceflight projects through sharing or augmentation of experiment samples to expand omics analyses on precious spaceflight samples. The partnerships ensure that the maximum amount of data is garnered from spaceflight experiments and made publically available as rapidly as possible via the GLDS. GLDS Version 1.0, went online in April 2015. Software updates and new data releases occur at least quarterly. As of October 2016, the GLDS contains 80 datasets and has search and download capabilities. Version 2.0 is slated for release in September of 2017 and will have expanded, integrated search capabilities leveraging other public omics databases (NCBI GEO, PRIDE, MG-RAST). Future versions in this multi-phase project will provide a collaborative platform for omics data analysis. Data from experiments that explore the biological effects of the spaceflight environment on a wide variety of model organisms are housed in the GLDS including data from rodents, invertebrates, plants and microbes. Human datasets are currently limited to those with anonymized data (e.g., from cultured cell lines). GeneLab ensures prompt release and open access to high-throughput genomics, transcriptomics, proteomics, and metabolomics data from spaceflight and ground-based simulations of microgravity, radiation or other space environment factors. The data are meticulously curated to assure that accurate experimental and sample processing metadata are included with each data set. GLDS download volumes indicate strong interest of the scientific community in these data. To date GeneLab has partnered with multiple experiments including two plant (Arabidopsis thaliana) experiments, two mice experiments, and several microbe experiments. GeneLab optimized protocols in the rodent partnerships for maximum yield of RNA, DNA and protein from tissues harvested and preserved during the SpaceX-4 mission, as well as from tissues from mice that were frozen intact during spaceflight and later dissected on the ground. Analysis of GeneLab data will contribute fundamental knowledge of how the space environment affects biological systems, and as well as yield terrestrial benefits resulting from mitigation strategies to prevent effects observed during exposure to space environments.

spaceflight↗

Meteoric 10Be Flux Calibration Data for the East River Watershed, Colorado, USA

This data package contains tabular and geospatial data used to quantify and model meteoric beryllium-10 fluxes in the East River watershed, Colorado, USA. The tabular component includes calibration-site data from five glacial moraine sites and includes environmental variables used to evaluate spatial controls on meteoric 10Be delivery, including elevation, mean annual precipitation (MAP), mean snow depth, and mean snow water equivalent (SWE). These site-level data were used to compare observed fluxes with environmental gradients across the watershed and to evaluate the effects of erosion correction on flux estimates. The package also includes supporting slope and curvature values used to assess topographic inputs to the erosion analysis. A second component of the data package contains updated manuscript tables and regression outputs used to summarize the relationships between meteoric 10Be flux and environmental predictors. These tables include meteoric 10Be sample information and AMS results, site-level environmental values, site-level meteoric 10Be inventory and flux values, watershed-averaged predicted fluxes, soil bulk density measurements, fine-fraction values, soil pH measurements, and regression statistics including slope, intercept, coefficient of determination, and p-value. The regression products include both standard linear regressions and regressions in which the intercept is constrained to pass through zero, and they support the analyses presented in the companion manuscript. Together, these tabular files provide the numerical basis for the manuscript tables and the regression-based interpretation of meteoric 10Be flux variability in a snow-dominated mountain watershed. The geospatial component of the package consists of GeoTIFF raster files used to generate the map products presented in Figures 2 and 6 of the companion manuscript. These rasters represent watershed-scale spatial layers for environmental variables and regression-based predictions of meteoric 10Be flux. This dataset contains comma-separated values files (.csv), Microsoft Excel files (.xlsx), GeoTIFF raster files (.tif), and upporting metadata files, including CSV data dictionaries and readme text files (.csv, .txt). The tabular files can be opened with standard spreadsheet software, and the raster files can be viewed and analyzed in GIS software such as ArcGIS Pro or QGIS. Together, these files document the numerical and spatial datasets used to calibrate and predict meteoric 10Be delivery in the East River watershed.

East River↗

Quantum Hardware-Enabled Molecular Dynamics via Transfer Learning

The ability to perform ab initio molecular dynamics simulations using potential energy surfaces provided by quantum computers would open the door to virtually exact dynamics for a variety of chemical and biochemical systems, with impacts on catalysis and biophysics. Nonetheless, performing molecular dynamics on surfaces produced by quantum hardware has been hampered by the noisy energies typically produced by quantum computers and challenges associated with computing gradients and scaling to large systems interest. A recent set of advances in machine learning, known as transfer learning, provides a new path forward for molecular dynamics simulations on quantum hardware. Transfer learning offers a workaround, where one first trains models on larger, less accurate classical datasets and then refines them on smaller, more accurate quantum datasets. We explore this approach by training machine learning models to predict a molecule's potential energy based on its geometric structure using Behler-Parrinello neural networks. When successfully trained, the model enables energy gradient predictions necessary for dynamic simulations. To reduce the quantum resources needed, the model is initially trained with data derived from classical density functional theory and subsequently refined with a smaller dataset obtained from a variational quantum eigensolver optimization of the unitary coupled cluster ansatz. We show that this approach significantly reduces the size of the needed quantum training dataset while capturing the high accuracies needed within quantum chemistry simulations. The success of this two-step training method opens more opportunities to apply machine learning models on quantum data, a significant stride towards efficient quantum-classical hybrid computational models.

quantum computing↗

IM3 Data Center Driven Grid Stress Dataset for the U.S. Western Interconnection

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from an open-source grid operations modeling framework - GO), for the Integrated Multisector, Multiscale Modeling (IM3) project, under varying levels of data center demand growth between 2025 and 2035 in the U.S. Western Interconnection. The scenarios and sensitivity experiments are combinations of different data center demand growth rates and energy, weather, population and economic pathways. Data center demand growth projections were sourced from the Electric Power Research Institute (EPRI). The data center demand growth projection names are: Low (3.71% annual data center demand growth) Moderate (5% annual data center demand growth) High (10% annual data center demand growth) Higher (15% annual data center demand growth) Energy, weather, population and economic pathways are informed by two Shared Socioeconomic Pathways (SSP3 and SSP5) and two Representative Concentration Pathways (RCP4.5 and RCP8.5) following the hotter general circulation model (GCM) forcing group from a set of perturbed thermodynamics simulations. The resulting pathway names are: rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 The main scenarios and sensitivity experiments are detailed below. Reference scenario: The projected grid stress and reliability results for the U.S. Western Interconnection from a previous study. This scenario does not consider data center demand growth explicitly. Data center scenario: Building on the reference scenario, this scenario considers various data center growth rates and how they impact the U.S. Western Interconnection. Data center loads are modeled as flat 8760-hr profiles. This scenario does not consider new generation and transmission capacities specifically designed to meet the new data center demands. The related folder is named "flat". Delayed generator retirements sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different levels of natural gas and nuclear generator retirement delays. The resulting scenario names are: (1) postponing 100% nuclear retirements; (2) postponing 100% nuclear and 25% natural gas retirements; (3) postponing only 50% natural gas retirements; (4) postponing 100% nuclear and 50% natural gas retirements; (5) postponing 100% nuclear and 75% natural gas retirements; and (6) postponing 100% nuclear and 100% natural gas retirements. The related folder names are: no_gen_retire_0_gas, no_gen_retire_25_gas, no_gen_retire_50_gas, no_gen_retire_50_gas_only, no_gen_retire_75_gas, and no_gen_retire_100_gas. Demand response through curtailment sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different participation and compensation levels of data center demand response. The resulting scenario names are: (1) 5% demand available for curtailment with 750 $/MWh compensation; (2) 5% demand available for curtailment with 500 $/MWh compensation; (3) 5% demand available for curtailment with 250 $/MWh compensation; (4) 15% demand available for curtailment with 750 $/MWh compensation; (5) 15% demand available for curtailment with 500 $/MWh compensation; and (6) 15% demand available for curtailment with 250 $/MWh compensation. The related folder names are: dr_cost_250_drup_0_drdown_5, dr_cost_250_drup_0_drdown_15, dr_cost_500_drup_0_drdown_5, dr_cost_500_drup_0_drdown_15, dr_cost_750_drup_0_drdown_5, and dr_cost_750_drup_0_drdown_15. Combination of delayed generator retirements and demand response through curtailment sensitivity experiment: The impact of combining postponing 100% nuclear and 25% natural gas retirements with 5% demand available for curtailment with 750 $/MWh compensation is simulated. The related folder is named "dr_cost_750_drup_0_drdown_5_nuc_100_gas_25". Please refer to the README file for a detailed description of the dataset including individual files and references.

Artificial Intelligence↗

Soil Temperature Sensor Data, May - December 2024, Five sites in Knoxville, Tennessee

This dataset contains surface soil temperature measurements from five urban parks in Knoxville: Cumberland Estates Park (CE), Socially Equal Energy Efficient Development (SD), West View Park (WV), Victor Ashe Park (VA), and West Hills Park (WH). The dataset comprises 29 Excel (.xlsx) and Comma-Separated Values (.csv) files documenting soil temperature recorded by six HOBO Pendant MX Water Temperature Data Loggers at each site, except for CE, which has five sensors. Each logger was installed at a depth of 10 inches and positioned approximately 3 to 6 feet from the weather station at each site. To open both csv and xlsx files, users can utilize spreadsheet applications such as Microsoft Excel, Google Sheets, LibreOffice Calc, and Apple Numbers. This dataset is part of a broader study examining the effects of soil moisture and plant evapotranspiration on ambient temperature and relative humidity across multiple urban parks in Knoxville, Tennessee.

EARTH SCIENCE > LAND SURFACE > SOILS↗

OPEN ALPHADIFFRACT

Open-source release of the AlphaDiffract data generation and training system. Includes only the public Materials Project dataset retrievers.AlphaDiffract is a deep learning framework that achieves state-of-the-art performance in predicting the crystal system, space group, and lattice parameters directly from PXRD patterns. AlphaDiffract utilizes a 1D adaptation of the ConvNeXt architecture, a modern convolutional neural network that integrates key design principles from transformers, coupledwith dedicated prediction heads for each crystallographic property.

Prince, Michael [Argonne National Laboratory (ANL)↗

Rocket Launch Detection with Smartphone Audio and Transfer Learning

Rocket launches generate infrasound signatures that have been detected at great distances. Due to the sparsity of the networks that have made these detections, however, most signals are detected tens of minutes to hours after the rocket launch. In this work, a method of near-real-time detection of rocket launches using data from a network of smartphones located 10–70 km from launch sites is presented. A machine learning model is trained and tested on the open-access Aggregated Smartphone Timeseries of Rocket-generated Acoustics (ASTRA), Smartphone High-explosive Audio Recordings Dataset (SHAReD), and ESC-50 datasets, resulting in a final accuracy of 97% and a false positive rate of <1%. The performance and behavior of the model are summarized, and its suitability for persistent monitoring applications is discussed.

acoustics↗

Provenance in Data Interoperability for Multi-Sensor Intercomparison

As our inventory of Earth science data sets grows, the ability to compare, merge and fuse multiple datasets grows in importance. This requires a deeper data interoperability than we have now. Efforts such as Open Geospatial Consortium and OPeNDAP (Open-source Project for a Network Data Access Protocol) have broken down format barriers to interoperability; the next challenge is the semantic aspects of the data. Consider the issues when satellite data are merged, cross-calibrated, validated, inter-compared and fused. We must match up data sets that are related, yet different in significant ways: the phenomenon being measured, measurement technique, location in space-time or quality of the measurements. If subtle distinctions between similar measurements are not clear to the user, results can be meaningless or lead to an incorrect interpretation of the data. Most of these distinctions trace to how the data came to be: sensors, processing and quality assessment. For example, monthly averages of satellite-based aerosol measurements often show significant discrepancies, which might be due to differences in spatio- temporal aggregation, sampling issues, sensor biases, algorithm differences or calibration issues. Provenance information must be captured in a semantic framework that allows data inter-use tools to incorporate it and aid in the intervention of comparison or merged products. Semantic web technology allows us to encode our knowledge of measurement characteristics, phenomena measured, space-time representation, and data quality attributes in a well-structured, machine-readable ontology and rulesets. An analysis tool can use this knowledge to show users the provenance-related distrintions between two variables, advising on options for further data processing and analysis. An additional problem for workflows distributed across heterogeneous systems is retrieval and transport of provenance. Provenance may be either embedded within the data payload, or transmitted from server to client in an out-of-band mechanism. The out of band mechanism is more flexible in the richness of provenance information that can be accomodated, but it relies on a persistent framework and can be difficult for legacy clients to use. We are prototyping the embedded model, incorporating provenance within metadata objects in the data payload. Thus, it always remains with the data. The downside is a limit to the size of provenance metadata that we can include, an issue that will eventually need resolution to encompass the richness of provenance information required for daata intercomparison and merging.

Lynnes, Chris↗

High-Resolution IR Absorption Spectroscopy of Polycyclic Aromatic Hydrocarbons in the 3-micrometers Region: Role of Periphery

In this work we report on high-resolution IR absorption studies that provide a detailed view on how the peripheral structure of irregular polycyclic aromatic hydrocarbons (PAHs) affects the shape and position of their 3-micrometers absorption band. To this purpose we present mass-selected, high-resolution absorption spectra of cold and isolated phenanthrene, pyrene, benz[a]antracene, chrysene, triphenylene, and perylene molecules in the 2950-3150 per cm range. The experimental spectra are compared with standard harmonic calculations, and anharmonic calculations using a modified version of the SPECTRO program that incorporates a Fermi resonance treatment utilizing intensity redistribution. We show that the 3-micrometers region is dominated by the effects of anharmonicity, resulting in many more bands than would have been expected in a purely harmonic approximation. Importantly, we find that anharmonic spectra as calculated by SPECTRO are in good agreement with the experimental spectra. Together with previously reported high-resolution spectra of linear acenes, the present spectra provide us with an extensive dataset of spectra of PAHs with a varying number of aromatic rings, with geometries that range from open to highly-condensed structures, and featuring CH groups in all possible edge configurations. We discuss the astrophysical implications of the comparison of these spectra on the interpretation of the appearance of the aromatic infrared 3-micrometers band, and on features such as the two-component emission character of this band and the 3-micrometers emission plateau.

techniques: spectroscopic↗

A Bayesian Learning Approach to Wireless Outdoor Heatmap Construction using Deep Gaussian Process

We present a novel Bayesian learning approach to outdoor radio heatmap construction utilizing deep Gaussian process (GP). The proposed approach employs a two-layer hierarchy which consists of two cascaded Gaussian processes that are capable of modeling more complex input-output relations than standard single-layer Gaussian processes. Since deriving the exact model likelihood is challenging, a lower bound is optimized instead so that gradient descent-based methods can be performed to find out the optimal model parameters. Typically, inducing points are used in GPs to facilitate low-rank approximation of covariance (kernel) matrices for computation speedup. However, the inaccuracy induced by inducing points can accumulate when stacking multiple layers of GP which may hinder the performance of deep GP. Moreover, since inducing points need to be learned, having them at all layers of deep GP also incurs computational burden. To overcome the above challenges, in contrast to the canonical deep GP model, we use a modified architecture where a full standard GP resides in the first layer and inducing points are only introduced for the second layer. This modified architecture strikes a balance between model accuracy and training complexity. In the proposed model, the noise parameter of the first GP layer is also eliminated to improve the training efficiency as the noise parameter at the output of the second layer suffices to model the uncertainty in the output. The proposed approach is evaluated on real-world datasets, in the form of location-Received Signal Strength (RSS) pairs, collected from the Platform for Open Wireless Data-driven Experimental Research (POWDER) located at the campus of the University of Utah. Experiment results show that the proposed approach can achieve smaller prediction errors on various training and testing data configurations than DNN-based and GP-based methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

AI Foundation Models for Science: An Open Collaborative Initiative

Foundation Models (FMs), AI models designed to replace task-specific models, are increasingly being recognized for their versatility across numerous downstream applications. These models, trained using self-supervised techniques on any type of sequence data, circumvent the need for large annotated datasets, a major bottleneck in traditional AI model development. FMs can be applied to downstream tasks using few-shot learning and fine-tuning, significantly reducing the need for large labeled training datasets and computational resources. However, the development of FMs requires substantial resources, including access to data and compute power, expertise in the latest models, and specialized scientific knowledge for systematic evaluation. It is challenging for a single group to possess all these capabilities. To address this, NASA IMPACT has initiated an open collaborative effort, leveraging partnerships with the private sector and other groups within and outside NASA, to jointly build FMs. The overarching goal is to develop a consistent and collaborative approach to building FMs for high-value science datasets. This initiative has fostered collaboration within NASA and with external partners, including IBM Research, Clark University, DOE’s ORNL, ESA, and USGS. The effort focuses on identifying key datasets with a wide range of downstream applications, pretraining and building FMs using modified transformer architectures, evaluating compute infrastructure needs, and sharing models, pretraining and fine-tuning code, and data with the community. Furthermore, it aims to train the Earth science community to fine-tune these models for various downstream applications. Our initial effort resulted in the creation of a 100 million parameter HLS Geospatial Model within six months, which was released on HuggingFace. We are now expanding our scope to include data from weather and climate models and investigating multimodal models. We invite those interested in participating in this effort to join us by sharing their use cases, expertise, or data.

Rahul Ramachandran↗

Enhancing ChatPORT with CUDA-to-SYCL Kernel Translation Capability

Large Language Models (LLMs) have shown strong capabilities in general code translation. However, code translation involving parallel programming models remains largely unexplored. This work enhances the capabilities of code LLMs in CUDA-to-SYCL kernel translation with parameter-efficient fine-tuning. The resultant fine-tuned LLM, called ChatPORT, is an effort to provide high-fidelity translations from one programming model to another. We describe the preparation of datasets from heterogeneous computing benchmarks for model fine-tuning and testing, the parameter-efficient fine-tuning of 19 open-source code models ranging in size from 0.5 to 34 billion parameters and evaluate the correctness rates of the SYCL kernels by the fine-tuned models. The experimental results show that most code models fail to translate CUDA codes to SYCL correctly. However, fine-tuning these models using a small set of CUDA and SYCL kernels can enhance the capabilities of these models in kernel translation. Depending on the sizes of the models, the correctness rate ranges from 19.9% to 81.7% for a test dataset of 62 CUDA kernels.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Vegetation classification map and covariates associated with NEON AOP survey, East River, CO 2018

This package includes geospatial data layers developed to investigate how environmental gradients—specifically topography and near-surface soil properties—drive the spatial arrangement of dominant plant communities in mountainous watersheds. The geospatial products, which support the analysis of these ecological relationships, are derived from airborne hyperspectral and LiDAR datasets acquired by the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP), in conjunction with an extensive ground field campaign conducted in summer 2018. This work is part of the DOE Watershed Function Science Focus Area (SFA) and features geospatial datasets developed based on observations and ground data collected at East River, Colorado, in collaboration with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey in June 2018. Classification Map: - Classification Map (PNG, GeoTIFF): Derived from hyperspectral and LiDAR airborne data using a machine learning approach. - Class Code Mapper (CSV): Associates pixel values with corresponding vegetation/non-vegetation classes. - Classification Reference Data (CSV): Reference data used in the machine learning procedure. LiDAR-Derived Products: - Topographical Metrics (GeoTIFFs): Elevation, slope, curvature, TWI, TPI, solar insolation, and canopy height model (CHM), smoothed with a 5x5 pixel window. Vegetation Indices: - GeoTIFFs of NDVI, NDNI, NDWI: Vegetation indices derived from hyperspectral data. Urban Masks: - Urban Mask (GeoTIFF): Applied to the mapping to convert bare soil classes to urban classes. Software Compatibility: GeoTIFFs: Can be visualized with GIS software or libraries that support GeoTIFF images. CSV Files: Can be opened with any software that handles comma-separated values. The FLMD file provides details and links to the source datasets used to derive the products. The manuscript (in the Method session) provides details on how each product was derived. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Update on 2026-03-25: Since the original dataset publication date of 02/28/2020, this package has a new classification map derived by an improved methodology. This update also includes additional ground data that improved the representation of some of the communities. See the methods for further details on what has changed between versions.

2018 NEON and 2025 CHESS Campaigns↗

Harnessing large language models’ zero-shot and few-shot learning capabilities for regulatory research

Abstract Large language models (LLMs) are sophisticated AI-driven models trained on vast sources of natural language data. They are adept at generating responses that closely mimic human conversational patterns. One of the most notable examples is OpenAI's ChatGPT, which has been extensively used across diverse sectors. Despite their flexibility, a significant challenge arises as most users must transmit their data to the servers of companies operating these models. Utilizing ChatGPT or similar models online may inadvertently expose sensitive information to the risk of data breaches. Therefore, implementing LLMs that are open source and smaller in scale within a secure local network becomes a crucial step for organizations where ensuring data privacy and protection has the highest priority, such as regulatory agencies. As a feasibility evaluation, we implemented a series of open-source LLMs within a regulatory agency’s local network and assessed their performance on specific tasks involving extracting relevant clinical pharmacology information from regulatory drug labels. Our research shows that some models work well in the context of few- or zero-shot learning, achieving performance comparable, or even better than, neural network models that needed thousands of training samples. One of the models was selected to address a real-world issue of finding intrinsic factors that affect drugs' clinical exposure without any training or fine-tuning. In a dataset of over 700 000 sentences, the model showed a 78.5% accuracy rate. Our work pointed to the possibility of implementing open-source LLMs within a secure local network and using these models to perform various natural language processing tasks when large numbers of training examples are unavailable.

Biochemistry & Molecular Biology↗

Echo state network for coarsening dynamics of charge density waves

An echo state network (ESN) is a type of reservoir computer that uses a recurrent neural network with a sparsely connected hidden layer. Compared with other recurrent neural networks, one great advantage of ESN is the simplicity of its training process. Yet, despite the seemingly restricted learnable parameters, ESN has been shown to successfully capture the spatial-temporal dynamics of complex patterns. Here we build an ESN to model the coarsening dynamics of charge-density waves (CDWs) in a semiclassical Holstein model, which exhibits a checkerboard electron density modulation at half-filling stabilized by a commensurate lattice distortion. The inputs to the ESN are local CDW order parameters in a finite neighborhood centered around a given site, while the output is the predicted CDW order of the center site at the next time step. Special care is taken in the design of couplings between hidden layer and input nodes to ensure lattice symmetries are properly incorporated into the ESN model. Since the model predictions depend only on CDW configurations of a finite domain, the ESN is scalable and transferrable in the sense that a model trained on dataset from a small system can be directly applied to dynamical simulations on larger lattices. Furthermore, our work opens avenues for efficient dynamical modeling of pattern formations in functional electron materials.

2-dimensional systems↗

Leveraging Google Earth Engine User Interface for Semiautomated Wetland Classification in the Great Lakes Basin at 10 m With Optical and Radar Geospatial Datasets

As one of the world’s largest freshwater ecosystems,the Great Lakes Basin houses hundreds of thousands of acres of wetlands that support a variety of crucial ecological and environmental functions at the local, regional, and global levels.Monitoring these wetlands is critical to conservation and restoration efforts, however current methods that rely on field monitoring are labor-intensive, costly, and often outdated. In this study, we present a graphical user interface constructed in Google Earth Engine called the Wetland Extent Tool (WET),which allows semi-automatic wetland classification according to a user-input area of interest and date range. WET composites datasets and conducts multi source, moderate resolution processing utilizing Landsat 8 OLI, Sentinel-2 MSI, Sentinel-1 C-SAR, and Shuttle Radar Topography Mission (SRTM) datasets to classify wetlands in the entire Great Lakes Basin. We evaluated classification results of wetlands, uplands, and open water from May-September 2019, and tested whether SRTM elevation, slope,or the Dynamic Surface Water Extent produced the most accurate results in each Great Lake Basin in conjunction with optical indices and radar composites. We found that elevation produced the most accurate classification in Lake Erie, Michigan,and Ontario, while slope performed best in Lake Huron and Superior. Lake Erie, Michigan, Ontario, and Huron achieved high overall accuracy and identification of wetlands. WET leverages cloud-computing for multi source processing of moderate resolution remote sensing data, and employs a user interface in Google Earth Engine that wetland managers and conservationists can use to monitor wetland extent in the Great Lakes Basin in near real-time.

Vanessa L Valenti↗