MultiSolSegment: EL images and masks
This dataset was used to train MultiSolSegment, a multi-channel segmentation model for photovoltaic defect detection.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
This dataset was used to train MultiSolSegment, a multi-channel segmentation model for photovoltaic defect detection.
Nearly 30% of commercial building energy use is wasted due to equipment faults and HVAC controls problems. The result is increased emissions, compromised comfort and productivity, and less reliable coordination of building power needs with a clean grid. The energy impact alone represents $17 billion in potential savings. Today’s smart building software provides a robust solution to address these operational deficiencies. Energy management and information systems (EMIS) are saving up to 9% on average, with two-year paybacks. They are being incorporated into energy management processes, commissioning services, and utility programs. As effective as they are, two barriers prevent even deeper benefits; limited personnel to fix problems once they are identified, and the expense and time to manually implement changes in control systems. In partnership with the research community, the EMIS industry is developing new capabilities to overcome these barriers. Moving beyond siloed products for either fault detection and diagnostics, or optimal control, these new capabilities empower users to not only automatically identify faults, but also to push corrective action, and control improvements to their buildings. In this paper, several areas for enhancements are documented: ‘one-time’ correction of faults such as setpoints, schedules, and economizer lockouts; short-term active testing for automated proportional integral derivative (PID) loop tuning and functional testing; and continuous supervisory control for demand flexibility and year-round efficiency. Results are presented from a pair of partner implementations out of a dozen providers integrating these enhancements into their products, including field tests from across the country, and insights into operator acceptance and integration into operations and maintenance practices.
Mass spectra of 1,2-cyclohexanediol and deuterium labeled analogs data recorded and mechanistic rationalizations of degradation processes given
SCAN program uses scanning algorithm to locate tokens in line of input data. Tokens can be command words, numbers, data values, labels. Using SCAN subroutines, user extracts tokens from character strings in languages with simple or complex syntax. SCAN thoroughly tested and implemented in NASA's Descent Design System for Shuttle orbiter. SCAN useful for other programs requiring input scanning. SCAN written in FORTRAN 77.
Obtaining sufficient labelled training data is a persistent difficulty for speech recognition research. Although well transcribed data is expensive to produce, there is a constant stream of challenging speech data and poor transcription broadcast as closed-captioned television. We describe a reliable unsupervised method for identifying accurately transcribed sections of these broadcasts, and show how these segments can be used to train a recognition system. Starting from acoustic models trained on the Wall Street Journal database, a single iteration of our training method reduced the word error rate on an independent broadcast television news test set from 62.2% to 59.5%.
Exploration Class missions to Mars will require precautions against potential contamination by any native microorganisms that may be incidentally pathogenic to humans. While the results of NASA's Viking biology experiments of 1976 have been generally interpreted as inconclusive for surface organisms, the possibility of native surface life has never been ruled out and more recent studies suggest that the case for biological interpretation of the Viking Labeled Release data may now be stronger than it was when the experiments were originally conducted. It is possible that, prior to the first human landing on Mars, robotic craft and sample return missions will provide enough data to know with certainty whether or not future human landing sites harbor extant life forms. However, if native life is confirmed, it will be problematic to determine whether any of its species may present a medical risk to astronauts. Therefore, it will become necessary to assess empirically the risk that the planet contains pathogens based on terrestrial examples of pathogenicity and to take a reasonably cautious approach to bio-hazard protection. A survey of terrestrial pathogens was conducted with special emphasis on those pathogens whose evolution has not depended on the presence of animal hosts. The history of the development and implementation of Apollo anticontamination protocol and recent recommendations of the NRC Space Studies Board regarding Mars were reviewed. Organisms can emerge in nature in the absence of indigenous animal hosts and both infectious and non-infectious human pathogens are theoretically possible on Mars. The prospect of Martian surface life, together with the existence of a diversity of routes by which pathogenicity has emerged on Earth, suggests that the possibility of human pathogens on Mars, while low, is not zero. Since the discovery and study of Martian life can have long-term benefits for humanity, the risk that Martian life might include pathogens should not be an obstacle to human exploration. As a precaution, however, it is recommended that EVA suits be decontaminated when astronauts enter surface habitats when returning from field activity and that biosafety protocol approximating laboratory BSL 2 be developed for astronauts working in laboratories on the Martian surface. Quarantine of astronauts and Martian materials arriving on Earth should also be part of a human Mars mission and this and the surface biosafety program should be integral to human expeditions from the earliest stages of the mission planning.
For ten years the MODIS aerosol algorithm has been applied to measured MODIS radiances to produce a continuous set of aerosol products, over land and ocean. The MODIS aerosol products are widely used by the scientific and applied science communities for variety of purposes that span operational air quality forecasting in estimates o[ clear-sky direct radiative effects over ocean and aerosol-cloud interactions. The products undergo continual evaluation, including self-consistency checks and comparisons with highly accurate ground-based instruments. The result of these evaluation exercises is a quantitative understanding of the strengths and weaknesses of the retrieval, where and when the products are accurate and the situations where and when accuracy degrades. We intend 10 present results of the most recent critical evaluations including the first comparison of the over ocean products against the shipboard aerosol optical depth measurements of the Marine Aerosol Network (MAN), the demonstration of the lack of sensitivity to size parameter in the over land products and identification of residual problems and regional issues. While the current data set is undergoing evaluation, we are preparing for the next data processing, labeled Collection 6. Collection 6 will include transparent Quality Flags, a 3 km aerosol product and the 500m resolution cloud mask used within the aerosol n:bicvu|. These new products and adjustments to algorithm assumptions should provide users with more options and greater control, as they adapt the product for their own purposes.
No abstract available
We introduce a method convolutional neural networks to detect the presence of clouds in airborne camera images. We quantify the performance of this Cloud Detection Neural Network (CDNN) using human-labeled validation data where we report a 96% accuracy in detecting clouds in testing datasets for both Zenith- viewing and Forward-viewing models. We assess our performance by comparing the flight-averaged cloud fraction of zenith and forward CDNN retrievals, with that of the prototype hyperspectral total-diffuse Sunshine Pyranometer (SPN-S) instrument’s cloud optical depth data. Comparison of the CDNN with the SPN-S on time specific intervals resulted in 93% accuracy for the zenith- viewing CDNN and 84% for the forward- viewing CDNN. The comparison of the CDNNs with the SPN-S on flight-averaged cloud fraction resulted in an agreement of 0.15 for the Forward CDNN and 0.07 for the Zenith CDNN. We then quantify the ability of the CDNN to identify the presence of clouds above the aircraft using a forward- looking camera mounted inside the aircraft cockpit compared to the use of an All- Sky upward-looking camera that is mounted outside the fuselage on top of the aircraft. We present results from the CDNN based on airborne imagery from the NASA Aerosol Cloud Meteorology Interactions Over the Western Atlantic Experiment (ACTIVATE) and the Clouds, Aerosol and Monsoon Processes- Philippines Experiment (CAMP2Ex). For CAMP2Ex 53% of flight dates had above- aircraft cloud fraction above 50%, while for ACTIVATE 52% and 54% of flight dates observed above-aircraft cloud fraction above 50% for 2020 and 2021, respectively.
Explore the source record for details and available documents.
As particle accelerators grow in complexity, traditional control methods face increasing challenges in achieving optimal performance. This paper envisions a paradigm shift: a decentralized multi-agent framework for accelerator control, powered by Large Language Models (LLMs) and distributed among autonomous agents. We present a proposition of a self-improving decentralized system where intelligent agents handle high-level tasks and communication and each agent is specialized control individual accelerator components. This approach raises some questions: What are the future applications of AI in particle accelerators? How can we implement an autonomous complex system such as a particle accelerator where agents gradually improve through experience and human feedback? What are the implications of integrating a human-in-the-loop component for labeling operational data and providing expert guidance? We show two examples, where we demonstrate viability of such architecture.
Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.
A generic, U.S.-based analysis approach was evaluated with respect to corn and soybean identification in Argentina. Using crop separability expectations derived from the analysis of Argentina ancillary data and U.S. spectral data, the approach was applied to Argentina spectral data by an expert analyst. Eight classes were detected and labeled independent of ground data. A high correspondence between the labels and limited ground data was achieved. It was concluded that an approach of this type could be applied to Argentina without major difficulty.
This report summarizes the analyses and results produced by a five-member investigative team of Government, university, and industry experts, established by NASA HQ. The team examined data quality problems associated with high performance liquid chromatography (HPLC) analyses of pigment concentrations in seawater samples produced by the San Diego State University (SDSU) Center for Hydro-Optics and Remote Sensing (CHORS). This report shows CHORS did not validate the methods used before placing them into service to analyze field samples for NASA principal investigators (PIs), even though the HPLC literature contained easily accessible method validation procedures, and the importance of implementing them, more than a decade ago. In addition, there were so many sources of significant variance in the CHORS methodologies, that the HPLC system rarely operated within performance criteria capable of producing the requisite data quality. It is the recommendation of the investigative team to a) not correct the data, b) make all the data that was temporarily sequestered available for scientific use, and c) label the affected data with an appropriate warning, e.g., "These data are not validated and should not be used as the sole basis for a scientific result, conclusion, or hypothesis--independent corroborating evidence is required."
NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.
Machine learning (ML) is being increasingly utilized in Earth science research. Benefits of ML include efficiency, reduction of human error, and ability to extract hidden patterns within data. However, the mutual lack of each other’s domain knowledge by ML and Earth science stands as a barrier to timely and effective implementation. Earth science, in particular, faces challenges in generating sample data, compared to those of traditional ML problems such as face recognition or stock predictions, where data is abundant and not lacking in ground truth, which is necessary for labeling. Earth science data are more varying in formats, such as HDF5 and image resolutions, and are not standardized across instruments, even within a given Earth science discipline. Previous studies have been done to outline the specific challenges that Earth science faces with ML, while others have focused on using existing publications to mine information efficiently. Other resources such as Scikit-Learn have developed decision trees for choosing appropriate machine learning algorithms, but application within Earth science subjects becomes much more complex. For the current study, we propose a methodology and tool that aids in implementation of ML in Earth science using natural language processing (NLP). Our work comprises three main parts: (1) analyzing existing publications related to ML and Earth science, using natural language processing: (2) extracting from the publications information on ML models subjects in Earth Science: and (3) visualizing the extracted relationships as a network graph. The resulting network graph should aid the Earth science communities in applying optimal ML algorithms and guiding data preparation through visualization of similar studies. The network graph and analysis of document similarity will be the basis of our next step, which is to develop a decision tree for selecting optimal machine learning methodologies for specified Earth science applications.
The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production by conducting a statistical analysis of historical wind data over a five-year period (2020-2025) from a single 1.5MW turbine manufactured by General Electric (GE) located at NLR’s Flatirons Campus, to generate an experimental test profile that was deployed on a 1.25-MW proton exchange membrane type MC250 electrolyzer system manufactured by Nel Hydrogen . [1] While the electrolyzer balance-of-plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. The historical wind data provided several metrics, however, the analysis particularly focused on the measured power output by the wind turbine. The power output time series of data for each day was categorized by total energy generation and standard deviation, and the day that represented the highest combination of these two metrics was chosen – December 25th, 2022. This process was then repeated for a moving four-hour window within this day to identify the most statistically variable period. Finally, this four-hour period was scaled by 65% to match the 1.25 MW electrolyzer. The electrolysis system controls hydrogen production by varying DC current applied to the stack, from a maximum of 3000 A to a minimum safe operation of 300 A, or 10%. Because the current – voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical wind profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1 Hz frequency. For more details on the statistical analysis process, see the presentation labeled “ Public Reference Data for Megawatt-Scale Hydrogen Electrolysis” provided with each data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}_{scaling factor}-{electrolyzer ramp rate in amperes/second} For instance, “wind-GE1.5MW_0.65-400.zip” represents the hour-long experiment using historical data from the wind-GE1.5MW turbine, scaled to 65%, with the electrolyzer power supply set to a maximum ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and wind power input. A PDF file detailing the historical wind data statistical analysis used to generate the wind profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_historical_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis [1] nelhydrogen.com/product/mc-series-electrolyser .
Seismic sensors deployed near roadways effectively capture ground vibrations generated by passing vehicles. Although both traditional and machine‐learning algorithms have been utilized for analyzing such signals, independent validation of detected vehicle events remains limited. We applied two unsupervised machine‐learning algorithms, uniform manifold approximation and projection for dimension reduction, and hierarchical density‐based spatial clustering of applications with noise, to continuous seismic data collected along a road on the main campus of Oak Ridge National Laboratory. The algorithms identified seven distinct cluster labels across the entire dataset. By comparing these cluster labels with precipitation records from a nearby weather station and image‐derived labels from a local camera system, we identified one cluster associated with rainfall and another with vehicle activity. Our algorithms identified a greater number of vehicle‐related labels compared to the camera‐derived labels because seismic data are unaffected by poor lighting conditions. The arrival times of the newly detected vehicle signals corresponded well with the road’s speed limit, supporting our findings. Our algorithm outperformed the short‐term average/long‐term average method and k‐means clustering. Our results suggest that seismic data, when analyzed with machine‐learning algorithms, can complement existing vehicle monitoring systems, particularly under challenging environmental conditions.