Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Characterizing peak electricity demand for U.S. households: an assessment of end-use loads and demand factors

Understanding household peak electricity demand is critical to evaluate the technical need for electrical infrastructure upgrades. This study characterizes peak loads for existing and new equipment using metered data from a convenience sample of 11,940 U.S. dwellings from four sources, including 911 from two sources with end-use metering. After standardized data cleaning and labeling, we derived descriptive statistics for key metrics, such as maximum demand and demand factors, and developed predictive models relating 60- to 15-min demand for the National Electrical Code (NEC). Mean 15-min maximum demand was 9.7 kW (median 9.0 kW; IQR 7.0–11.5 kW, 95% CI 9.6–9.8 kW), indicating spare capacity in 98% of homes with hypothetical 100 A panels. Maximum demand increased with floor area and number of high-demand loads. Dwelling maximum demand was driven by higher-power, longer-duration heating appliances and vehicle charging, while most user-operated appliances contributed little. Demand factors are used to account for how most devices contribute less than their rated power to maximum demand. Existing load mean demand factors (28%; median 10%; IQR 0–58%; CI 28–29%) were higher than those for new loads (21%; median 7%; IQR 0–35%; CI 20–21%), because new loads changed the timing and magnitude of maximum demand. New high-demand loads had higher than average demand factors (40–60%). Whole dwelling demand factors support the NEC's 40% assumption, but they challenge its conservative 100% treatment of new HVAC. We propose a data-driven 50% demand factor for new equipment, which would align with metered data, improve affordability, and modernize electrical codes.

Appliances↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

recon3d

SAND2025-00533O recon3d is a software tool that provides automated 3D reconstruction and meshing capabilities. It processes labeled 3D image data from various sources, starting from image stacks, and calculates 3D feature distributions like size, shape, and location. The software also has tools for downscaling rectilinear grid data and creating tetrahedral meshes directly from image data. recon3d can be used by novice users via the command line with a properly formatted configuration file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Emery, John↗

Mono-mix strategy enables comparative proteomics of a cross-kingdom microbial symbiosis

Cross-kingdom microbial symbioses, such as those between algae and bacteria, are key players in biogeochemical cycles. The molecular changes during initiation and establishment of symbiosis are of great interest, but quantitatively monitoring such changes can be challenging, particularly when the microorganisms differ greatly in size or are intimately associated. Here, we analyze output from label-free, data-dependent acquisition (DDA) LC-MS/MS proteomics experiments investigating the well-studied interaction between the alga Chlamydomonas reinhardtii and the heterotrophic bacterium Mesorhizobium japonicum. We found that detection of bacterial proteins decreased in coculture by 50% proteome-wide due to the abundance of algal proteins. As a result, standard differential expression analysis led to numerous false-positive reports of significantly downregulated proteins, where it was not possible to distinguish meaningful biological responses to symbiosis from artifacts of the reduced protein detection in coculture relative to monoculture. We show that data normalization alone does not eliminate the impact of altered detection on differential expression analysis of the cross-kingdom symbiosis. We assessed two additional strategies to overcome this methodological artifact inherent to DDA proteomics. In the first, we combined algal and bacterial monocultures at a relative abundance that mimicked the coculture, creating a “mono-mix” control to which the coculture could be compared. This approach enabled comparable detection of bacterial proteins in the coculture and the monoculture control. In the second strategy, we enhanced detection of lowly abundant bacterial proteins by using sample fractionation upstream of LC-MS/MS analysis. When these simple approaches were combined, they allowed for meaningful comparisons of nearly 10,000 algal proteins and over 4,000 bacterial proteins in response to symbiosis by DDA. They successfully recovered expected changes in the bacterial proteome in response to algal coculture, including upregulation of sugar-binding proteins and transporters. They also revealed novel proteomic responses to coculture that guide hypotheses about algal-bacterial interactions.

Dupuis, Sunnyjoy [University of California, Berkel↗

The classification and mensuration subsystem

From an operational standpoint, the most significant item the classification and mensuration subsystem (CAMS) had to overcome in providing the acreage component of the wheat production estimates for LACIE was the scope (segment volume processing required). Peak processing requirements per day increased from 16 to 20 for phase 1 with 700 total segments, to 35 to 40 per day for phase 2 with 1700 total segments, to 75 to 80 per day for phase 3 with 3000 total segments. Key issues regarding interrelationships between man and machines were identified during phase 1 using first generation technology. Procedure 1, tested and evaluated during phase 2 and continued through the initial phase 3 processing period for winter wheat, showed the need for software modification, procedures development, and analyst training. CAMS operations are described with emphasis on the training backgrounds of the analysts, the available data, and the labeling logic.

Abotteen, K. M.↗

Development of land based radar polarimeter processor system

The processing subsystem of a land based radar polarimeter was designed and constructed. This subsystem is labeled the remote data acquisition and distribution system (RDADS). The radar polarimeter, an experimental remote sensor, incorporates the RDADS to control all operations of the sensor. The RDADS uses industrial standard components including an 8-bit microprocessor based single board computer, analog input/output boards, a dynamic random access memory board, and power supplis. A high-speed digital electronics board was specially designed and constructed to control range-gating for the radar. A complete system of software programs was developed to operate the RDADS. The software uses a powerful real time, multi-tasking, executive package as an operating system. The hardware and software used in the RDADS are detailed. Future system improvements are recommended.

Kronke, C. W.↗

Automated extraction of knowledge for model-based diagnostics

The concept of accessing computer aided design (CAD) design databases and extracting a process model automatically is investigated as a possible source for the generation of knowledge bases for model-based reasoning systems. The resulting system, referred to as automated knowledge generation (AKG), uses an object-oriented programming structure and constraint techniques as well as internal database of component descriptions to generate a frame-based structure that describes the model. The procedure has been designed to be general enough to be easily coupled to CAD systems that feature a database capable of providing label and connectivity data from the drawn system. The AKG system is capable of defining knowledge bases in formats required by various model-based reasoning tools.

Gonzalez, Avelino J.↗

Trick Simulation Environment 07

The Trick Simulation Environment is a generic simulation toolkit used for constructing and running simulations. This release includes a Monte Carlo analysis simulation framework and a data analysis package. It produces all auto documentation in XML. Also, the software is capable of inserting a malfunction at any point during the simulation. Trick 07 adds variable server output options and error messaging and is capable of using and manipulating wide characters for international support. Wide character strings are available as a fundamental type for variables processed by Trick. A Trick Monte Carlo simulation uses a statistically generated, or predetermined, set of inputs to iteratively drive the simulation. Also, there is a framework in place for optimization and solution finding where developers may iteratively modify the inputs per run based on some analysis of the outputs. The data analysis package is capable of reading data from external simulation packages such as MATLAB and Octave, as well as the common comma-separated values (CSV) format used by Excel, without the use of external converters. The file formats for MATLAB and Octave were obtained from their documentation sets, and Trick maintains generic file readers for each format. XML tags store the fields in the Trick header comments. For header files, XML tags for structures and enumerations, and the members within are stored in the auto documentation. For source code files, XML tags for each function and the calling arguments are stored in the auto documentation. When a simulation is built, a top level XML file, which includes all of the header and source code XML auto documentation files, is created in the simulation directory. Trick 07 provides an XML to TeX converter. The converter reads in header and source code XML documentation files and converts the data to TeX labels and tables suitable for inclusion in TeX documents. A malfunction insertion capability allows users to override the value of any simulation variable, or call a malfunction job, at any time during the simulation. Users may specify conditions, use the return value of a malfunction trigger job, or manually activate a malfunction. The malfunction action may consist of executing a block of input file statements in an action block, setting simulation variable values, call a malfunction job, or turn on/off simulation jobs.

Lin, Alexander S.↗

Prime agricultural land monitoring and assessment component of the California Integrated Remote Sensing System

The use of digital LANDSAT techniques for monitoring agricultural land use conversions was studied. Two study areas were investigated: one in Ventura County and the other in Fresno County (California). Ventura test site investigations included the use of three dates of LANDSAT data to improve classification performance beyond that previously obtained using single data techniques. The 9% improvement is considered highly significant. Also developed and demonstrated using Ventura County data is an automated cluster labeling procedure, considered a useful example of vertical data integration. Fresno County results for a single data LANDSAT classification paralleled those found in Ventura, demonstrating that the urban/rural fringe zone of most interest is a difficult environment to classify using LANDSAT data. A general raster to vector conversion program was developed to allow LANDSAT classification products to be transferred to an operational county level geographic information system in Fresno.

Estes, J. E.↗

Enhancement of the MODIS Snow and Ice Product Suite Utilizing Image Segmentation

A problem has been noticed with the current NODIS Snow and Ice Product in that fringes of certain snow fields are labeled as "cloud" whereas close inspection of the data indicates that the correct labeling is a non-cloud category such as snow or land. This occurs because the current MODIS Snow and Ice Product generation algorithm relies solely on the MODIS Cloud Mask Product for the labeling of image pixels as cloud. It is proposed here that information obtained from image segmentation can be used to determine when it is appropriate to override the cloud indication from the cloud mask product. Initial tests show that this approach can significantly reduce the cloud "fringing" in modified snow cover labeling. More comprehensive testing is required to determine whether or not this approach consistently improves the accuracy of the snow and ice product.

Tilton, James C.↗

Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers

Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Using Multiple Isotope-Labeled Infrared Spectra for the Structural Characterization of an Intrinsically Disordered Peptide

Intrinsically disordered proteins (IDPs) rapidly interconvert between conformers, requiring an ensemble description. This complicates their experimental characterization, and force field limitations pose challenges for their simulation. Here, in this work, we use isotope-labeled and unlabeled infrared (IR) spectra to reweight simulated ensembles of the elastin-like peptide GVGVPGVG, a paradigmatic disordered peptide. By comparing the results obtained with different spectra, we explicitly show that the weights are underdetermined by the ensemble averaged data. We identify which labels and frequency regions maximize structural information while minimizing sensitivity to simulation error and show that these regions report on whether the peptide makes specific interactions. Our work shows the importance of incorporating simulations and simulated spectra at the planning stages of isotope-labeled IR experiments and more generally provides a framework for interpreting IR data for IDPs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analysis of scanner data for crop inventories

Accomplishments for a machine-oriented small grains labeler T&E, and for Argentina ground data collection are reported. Features of the small grains labeler include temporal-spectral profiles, which characterize continuous patterns of crop spectral development, and crop calendar shift estimation, which adjusts for planting date differences of fields within a crop type. Corn and soybean classification technology development for area estimation for foreign commodity production forecasting is reported. Presentations supporting quarterly project management reviews and a quarterly technical interchange meeting are also included.

Horvath, R.↗

On the accuracy of pixel relaxation labeling

An analysis of pixel labeling by probabilistic relaxation techniques is presented to demonstrate that these labeling procedures degenerate to weighted averages in the vicinity of fixed points. A consequence of this is that undesired label conversions can occur, leading to a deterioration of labeling accuracy at a stage after an improvement has already been achieved. Means for overcoming the accuracy deterioration are suggested and used as the basis for a possible design strategy for using probabilistic relaxation procedures. The results obtained are illustrated using simple data sets in which labeling on individual pixels can be examined and also using Landsat imagery to show application to data typical of that encountered in remote sensing applications.

Richards, J. A.↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Digital autonomous terminal access communications

A significant problem for the Bus Monitor Unit is to identify the source of a given transmission. This problem arises from the fact that the label which identifies the source of the transmission as it is put into the bus is intercepted by the Digital Autonomous Terminal Access Communications (DATAC) terminal and removed from the transmission. Thus, a given subsystem will see only data associated with a label and never the identifying label itself. The Bus Monitor must identify the source of the transmission so as to be able to provide some type of error identification/location in the event that some problem with the data transmission occurs. Steps taken to alleviate this problem by modifications to the DATAC terminal are discussed.

Novacki, S.↗

Viking labeled release biology experiment - Interim results

All results of the labeled-release life-detection experiment conducted on Mars prior to conjunction are summarized. Tests at both landing sites provide remarkably similar evolution of radioactive gas upon addition of a radioactive nutrient to the Mars sample. The 'active' agent in the sample is stable to 18 C, but is substantially inactivated by heat treatment for 3 hours at 50 C and completely inactivated at 160 C, as would be anticipated if the active response were caused by microorganisms. Results from test and heat-sterilized control samples are compared with those obtained from terrestrial soils and a lunar sample. Possible nonbiological explanations of the Mars data are reviewed. Although such explanations of the labeled-release data depend on UV irradiation, the labeled-release response does not appear to depend on recent direct UV activation of surface material. Available facts do not yet permit a conclusion regarding the existence of life on Mars.

Levin, G. V.↗

Automated classification of scientific publications linked to GES DISC datasets

The data collections archived and distributedby the GES DISC NASA data center arewidely utilized for various Earth Science studies.As these collections are created, many researchworks are published regarding the collections, algorithms,validations and applications. SinceGES DISC collects these publications and providestheir citations for the users, it is helpful tocategorize them based on how they relate to the datasetsthey are associated with. Specifically,whether the publication that is linked to GES DISCdataset is using it for applicational research,or if it describes the algorithm for dataset creation,or the validation of the dataset, or providesthe general overview of the data collection. Currently,this process requires simple manuallabelling, and as such, may be possible to solve viaautomation. To approach this problem, wedeveloped machine learning classifiers to predictthe category a publication belongs to. We usedmanually labeled publications as training data forsupervised machine learning algorithms:Random Forest and Naive Bayes. We achieved classificationaccuracy that is substantially betterthan the baseline accuracy, thus greatly improvingthe efficiency of the publication internalanalysis.

Rohan Dayal↗