Search NASA⌕ Search

SEARCH · Search NASA

Results for “Labeled Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Heating effects on jack pine pyrogenic organic matter properties from a pyrocosm study in 2022

This dataset contains data associated with the preprint “Fire removes preexisting pyrogenic organic matter from the ecosystem through the mechanisms of both direct combustion and increasing mineralizability” (Luo et al., 2025b), which is the complementary study to the published paper “Reburning pyrogenic organic matter: a laboratory method for dosing dynamic heat fluxes from above” (Luo et al., 2025a). We designed a full-factorial experiment with different burial depths of jack pine (Pinus banksiana Lamb) pyrogenic organic matter (PyOM) (Surface, 1 cm, and 5 cm) and different heat-flux profiles (High, Low, and Control) to examine how subsequent fires affect the properties of preexisting PyOM. We measured total carbon (C), pH, dissolved organic carbon (DOC), dissolved inorganic carbon (DIC), and mineralized C (as CO₂-C, from a 12-week incubation).We found that high heat flux and/or surface placement resulted in substantial direct C losses through combustion. Intermediate heat exposure produced both combustion losses and increases in DOC and mineralizability, which may have complex long-term implications: an increased dissolved fraction of PyOM may promote downward transport into mineral soils and potentially contribute to deeper, longer-term C storage, but it may also make PyOM more susceptible to microbial decomposition. Under the lowest heat flux and deepest burial, most PyOM was retained, and changes in DOC and C mineralization were minimal. Finally, PyOM pH, an important chemical property, decreased under low-temperature heating but increased under higher temperatures.We uploaded pH data for all samples (“pH_of_all_samples.csv”); pH and temperature-related data (peak temperature and degree hours) for samples in High and Low heat-flux treatments (“pH_vs_peakT_and_degree_hours_only_for_heated_samples.csv”); total C data (“CN_pct_C_stock_C_loss_in_samples.csv”); DOC and DIC data (“doc_dic.csv”); and mineralized C (CO₂-C) data (“CO2-C_all_original.csv”). Additional details can be found in the Methods & Sampling section.All datasets uploaded to ESS-DIVE are clearly labeled, cleaned, and include both raw and derived data, ready for reuse in other analyses. All analysis code and raw datasets are also available on GitHub: https://github.com/MengmengLuo/Fire-removes-preexisting-pyrogenic-organic-matter-from-the-ecosystem.

54 ENVIRONMENTAL SCIENCES↗

Fluorescent Applications to Crystallization

By covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and tests with model proteins have shown that labeling u to 5 percent of the protein molecules does not affect the X-ray data quality obtained . The presence of the trace fluorescent label gives a number of advantages. Since the label is covalently attached to the protein molecules, it "tracks" the protein s response to the crystallization conditions. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a darker background. Non-protein structures, such as salt crystals, do not show up under fluorescent illumination. Crystals have the highest protein concentration and are readily observed against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. Preliminary tests, using model proteins, indicates that we can use high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that more rapid amorphous precipitation kinetics may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Experiments are now being carried out to test this approach using a wider range, of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗

Total Dissolved Nitrogen and Ammonia Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for total dissolved nitrogen (TDN) and ammonia concentrations for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. TDN was analyzed using a Shimadzu Total Nitrogen Module (TNM-1) combined with the TOC-VCSH analyzer (Shimadzu Corporation, Japan). TNM-1 is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. Ammonia was determined using a Lachat's QuikChem 8500 Series 2 Flow Injection Analysis System (LACHAT Instruments, QuckChem 8500 series 2, Automated Ion Analyzer, Loveland, Colorado). When ammonia in water samples is heated (60 degrees C) with salicylate and hypochlorite in an alkaline phosphate buffer, an emerald green color is produced which is proportional to the ammonia concentration. The color is intensified by the addition of nitroprusside. Ethylenediaminetetraacetic acid (EDTA) is added to the buffer to prevent the interference of metal ions (Ca, Mg, and Fe etc.). Ammonia-N is then determined by LACHAT flow injection and a colorimetric assay at an absorbance wavelength 660 nm. (Reference: LACHAT Instruments: QuickChem Method 90-107-06-3-A, Determination of Ammonia by Flow Injection Analysis (High Throughput, Salicylate Method/DCIC) (Multi Matrix method). Written by Lynn Egan (Application group), February 08, 2011.) All files are labeled by location and variable, and data reported are the mean values upon replicate measurements. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process as detailed in the methods. This data package contains (1) a zip file (tdn_ammonia_data_2015-2025.zip) containing a total of 299 files: 298 data files of ammonia and TDN data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the determination of Method Detection Limits (MDLs) for TDN data, which has been updated in 2026-08; and (5) PDF and docx files for the detemination of Method Detection Limits (MDLs) for Ammonia and the Interferences by LACHAT Flow Injection Analysis. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 105 locations containing TDN and Ammonia-N data. Update 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses and Determination of Method Detection Limit for Ammonia and the Interferences by LACHAT Flow Injection Analysis documents, which can be accessed as PDFs or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated total dissolved nitrogen and ammonia data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Units were listed incorrectly, but have been fixed to reflect correct units (ug/L). File level metadata (flmd) and data dictionary (dd) files were updated to reflect the updated versions of these files. Available data was added up until 2022-06-01. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-27. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for TDN were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for TDN data were added to this dataset.

54 ENVIRONMENTAL SCIENCES↗

Collaborative Supervised Learning for Sensor Networks

Collaboration methods for distributed machine-learning algorithms involve the specification of communication protocols for the learners, which can query other learners and/or broadcast their findings preemptively. Each learner incorporates information from its neighbors into its own training set, and they are thereby able to bootstrap each other to higher performance. Each learner resides at a different node in the sensor network and makes observations (collects data) independently of the other learners. After being seeded with an initial labeled training set, each learner proceeds to learn in an iterative fashion. New data is collected and classified. The learner can then either broadcast its most confident classifications for use by other learners, or can query neighbors for their classifications of its least confident items. As such, collaborative learning combines elements of both passive (broadcast) and active (query) learning. It also uses ideas from ensemble learning to combine the multiple responses to a given query into a single useful label. This approach has been evaluated against current non-collaborative alternatives, including training a single classifier and deploying it at all nodes with no further learning possible, and permitting learners to learn from their own most confident judgments, absent interaction with their neighbors. On several data sets, it has been consistently found that active collaboration is the best strategy for a distributed learner network. The main advantages include the ability for learning to take place autonomously by collaboration rather than by requiring intervention from an oracle (usually human), and also the ability to learn in a distributed environment, permitting decisions to be made in situ and to yield faster response time.

Wagstaff, Kiri L.↗

Systems and methods for an extensible business application framework

Method and systems for editing data from a query result include requesting a query result using a unique collection identifier for a collection of individual files and a unique identifier for a configuration file that specifies a data structure for the query result. A query result is generated that contains a plurality of fields as specified by the configuration file, by combining each of the individual files associated with a unique identifier for a collection of individual files. The query result data is displayed with a plurality of labels as specified in the configuration file. Edits can be performed by querying a collection of individual files using the configuration file, editing a portion of the query result, and transmitting only the edited information for storage back into a data repository.

Bell, David G.↗

Ameloblastin binding to biomimetic models of cell membranes – A continuum of intrinsic disorder

A 37-residue amino acid sequence corresponding to the segment encoded by exon-5 of murine ameloblastin (Ambn), AB2 (Y67-Q103), has been implicated with membrane association, ameloblastin self-assembly, and amelogenin-binding. Here, our aim was to characterize, at the residue level, the structural behavior of AB2 bound to chemical mimics of biological membranes using NMR spectroscopy. To better define the structure of AB2 using NMR-based methods, recombinant 13 C- and 15 N-labelled AB2 (*AB2) was prepared and data collected free in solution and with deuterated dodecylphosphocholine (dPC) micelles, deuterated bicelles, and both small and large unilamellar vesicles. Amide chemical shift and intensity perturbations observed in 1 H- 15 N HSQC spectra of *AB2 in the presence of bicelles and dPC micelles suggest that a region of *AB2, S6-E36 (murine Ambn S68 – E98), associates with the membrane biomimetics. A CSI-3 analysis of the NMR chemical shift assignments for *AB2 free in solution and bound to dPC micelles indicated the peptide remains disordered except for the adoption of a short, 12-residue α-helix, F10-G21 (murine Ambn F72-G83). In dPC micelles, the NOE NMR data was void of patterns characteristic of long-lived helical structure indicating this helix was transient in nature. A continuum of intrinsic disorder in the membrane-bound state may be responsible for ameloblastin’s ability to dynamically interact with multiple partners at the same site during amelogenesis.

59 BASIC BIOLOGICAL SCIENCES↗

Regime Characterization of Offshore Wind Resource Using Unsupervised Learning

Predictability of wind resource conditions is critical for offshore wind design and operations. While many studies of extreme wind conditions focus on specific events such as low-level jets or ramps, these rely on threshold definitions that limit generality. Here we present a data-driven framework that combines principal component analysis (PCA), self-organizing maps (SOM), and k-means clustering to classify wind resource conditions as typical and anomalous from climatological data. Anomalies are defined not by fixed thresholds but by flagging samples located far from SOM node centers inside the baseline SOM structure. This reframes extremes as rare ebents and hence, likely difficult to anticipate by numerical weather prediction models. We applied this approach to 23 years (2000–2022) of hourly profiles from the NOW-23 hindcast model at the Humboldt Wind Energy Area. Classification is conducted on a feature space consisting of 10 m wind speed and direction, bulk shear and veer across 30–270 m, and a low-level jet index. Dimensionality reduction is achieved through PC. A 2 × 3 OM lattice trained on the PCA vectors identified six baseline regimes spanning weak to strong flow states. High quantization-error profiles are identified and re-clustered into four anomalous regimes. The baseline regimes exhibited clear seasonal and diurnal cycles. Meanwhile, the anomalous regimes represented <10 % of all hours but showed distinct combinations of speed, shear, and veer, when compared to the baseline regimes. Anomalous regimes are typically short-lived (~few hours), yet their transitions can lead to hub-height wind changes of −18 to +9 m s -1 . For a representative 15 MW turbine, these shifts imply rapid swings in capacity factor from near-full output to negligible generation. Validation with lidar buoy data showed 51% agreement in SOM labels across ~6,000 overlapping hours, with most mismatches confined to adjacent speed classes. HRRR comparisons further revealed that anomalous regimes were disproportionately associated with forecast biases exceeding 5 m s -1 . Together, these results reframe extremes in offshore wind from absolute maxima or minima to weather states that are difficult to anticipate from models.

17 WIND ENERGY↗

Metabolomics analysis of P. tremula × P. alba ‘717-1B4’ tissue culture hybrids grown in liquid media supplemented with EGGC and d4-EGGC

A metabolite linked to ethylene metabolism in Populus was recently structurally characterized as 2-hydroxyethyl β-D-glucopyranoside, an ethylene glycol glucose conjugate (EGGC). The dataset presented here is associated with the metabolomics analysis of various tissues (i.e., roots, stems/leaves) of plants grown in liquid media supplemented with EGGC and stable-isotope labelled EGGC (i.e., d4-EGGC). Data were collected using a Thermo Scientific gas chromatograph (GC) coupled to a Q Exactive Orbitrap mass spectrometer (MS). Samples were collected at three different timepoints and silylated prior to GCMS analysis.

09 BIOMASS FUELS↗

Evaluation of high specific-heat ceramic for regenerator use at temperatures between 2-30 K

Specific heat, thermal conductivity (both in the range 2-30 K), and microhardness data were measured on the ceramics labelled LS-8, LS-8A, and LS-8A doped with CsI, SnCl2, and AgCl. A work hardened sample of LS-8A was also studied in an effort to determine the feasibility of using these types of LS-8 materials to replace Pb spheres in the regenerator of the JPL cryocooler. The LS-8A materials are all more than an order of magnitude harder than Pb, and the dopants do not significantly improve the hardness. However, the SnCl2 dopant has a remarkable effect in improving the specific heat and thermal conductivity of LS-8A. The SnCl2 doping level which maximized the regenerator enthalpy change in going from an unloaded to a loaded condition was found to be 0.2 percent SnCl2 in LS-8A. It was also found that the enthalpy change for a regenerator employing the LS-8A material is more than three times larger than for the Pb spheres case. The use of rods, rather than spheres, of optimally doped LS-8A in regenerators is discussed.

Lawless, W. N.↗

Aircraft wing weight build-up methodology with modification for materials and construction techniques

An aircraft wing weight estimating method based on a component buildup technique is described. A simplified analytically derived beam model, modified by a regression analysis, is used to estimate the wing box weight, utilizing a data base of 50 actual airplane wing weights. Factors representing materials and methods of construction were derived and incorporated into the basic wing box equations. Weight penalties to the wing box for fuel, engines, landing gear, stores and fold or pivot are also included. Methods for estimating the weight of additional items (secondary structure, control surfaces) have the option of using details available at the design stage (i.e., wing box area, flap area) or default values based on actual aircraft from the data base.

York, P.↗

Selection of the Australian indicator region

Each Australian state was examined for the availability of LANDSAT data, area, yield, and production characteristics, statistics, crop calendars, and other ancillary data. Agrophysical conditions that could influence labeling and classification accuracies were identified in connection with the highest producing states as determined from available Australian crop statistics. Based primarily on these production statistics, Western Australia and New South Wales were selected as the wheat indicator region for Australia. The general characteristics of wheat in the indicator region, with potential problems anticipated for proportion estimation are considered. The varieties of wheat, the diseases and pests common to New South Wales, and the wheat growing regions of both states are examined.

Reed, C. R.↗

Bar-Code-Scribing Tool

Proposed hand-held tool applies indelible bar code to small parts. Possible to identify parts for management of inventory without tags or labels. Microprocessor supplies bar-code data to impact-printer-like device. Device drives replaceable scribe, which cuts bar code on surface of part. Used to mark serially controlled parts for military and aerospace equipment. Also adapts for discrete marking of bulk items used in food and pharmaceutical processing.

Badinger, Michael A.↗

Fluorescence Studies of Protein Crystal Nucleation

Fluorescence can be used to study protein crystal nucleation through methods such as anisotropy, quenching, and resonance energy transfer (FRET), to follow pH and ionic strength changes, and follow events occurring at the growth interface. We have postulated, based upon a range of experimental evidence that the growth unit of tetragonal hen egg white lysozyme is an octamer. Several fluorescent derivatives of chicken egg white lysozyme have been prepared. The fluorescent probes lucifer yellow (LY), cascade blue, and 5-((2-aminoethyl)aminonapthalene-1-sulfonic acid (EDANS), have been covalently attached to ASP 101. All crystallize in the characteristic tetragonal form, indicating that the bound probes are likely laying within the active site cleft. Crystals of the LY and EDANS derivatives have been found to diffract to at least 1.7 A. A second group of derivatives is to the N-terminal amine group, and these do not crystallize as this site is part of the contact region between the adjacent 43 helix chains. However derivatives at these sites would not interfere with formation of the 43 helices in solution. Preliminary FRET studies have been carried out using N-terminal bound pyrene acetic acid (Ex 340 nm, Em 376 nm) lysozyme as a donor and LY (Ex -425 nm, Em 525 nm) labeled lysozyme as an acceptor. FRET data have been obtained at pH 4.6, 0.1 M NaAc buffer, at 5 and 7% NaCl, 4 C. The corresponding Csat values are 0.471 and 0.362 mg/ml (approximately 3.3 and approximately 2.5 x 10(exp -5) M respectively). The data at both salt concentrations show a consistent trend of decreasing fluorescence intensity of the donor species (PAA) with increasing total protein concentration. This decrease is more pronounced at 7% NaCl, consistent with the expected increased intermolecular interactions at higher salt concentrations reflected in the lower solubility. The calculated average distance between any two protein molecules at 5 x 10(exp -6) M is approximately 70nm, well beyond the range where any FRET can be expected. Results from these and ongoing studies will be presented.

Pusey, Marc L.↗

Model-Driven Development For PDS4 Software And Services

Software and services that access Planetary Data System (PDS) PDS4 data products need to parse product labels to retrieve, interpret, and process the referenced digital objects. Under PDS4 a driving principle is that the product label provide all of the information necessary for these functions to be performed accurately. However, significantly more information is available in the PDS4 Information Model (IM)[1], the controlling document used to define, create, and syntactically and semantically verify the product labels. This additional information in the IM is made available for use, by both software and services, to configure, promote resiliency, and improve interoperability.

Padams, Jordan↗

Correlation between aircraft MSS and LIDAR remotely sensed data on a forested wetland in South Carolina

Wetlands in a portion of the Savannah River swamp forest, the Steel Creek Delta, were mapped using April 26, 1985 high-resolution aircraft multispectral scanner (MSS) data. Due to the complex spectral characteristics of the wetland vegetation, it was necessary to implement several techniques in the classification of the MSS imagery of the Steel Creek Delta. In particular, when performing unsupervised classification, an iterative cluster busting technique was used which simplified the cluster labeling process. In addition to the MSS data, light detecting and ranging (LIDAR) data were acquired by National Aeronautics and Space Administration (NASA) personnel along two flightlines over the Steel Creek Delta. These data were registered with the wetland classification map and correlated. Statistical analyses demonstrated that the laser derived canopy height information was significantly correlated with the Steel Creek Delta wetland classes encountered along the profiling transect of the LIDAR data.

Jensen, John R.↗

Correlation between aircraft MSS and LIDAR remotely sensed data on a forested wetland in South Carolina

Wetlands in a portion of the Savannah River swamp forest, the Steel Creek Delta, were mapped using April 26, 1985 high-resolution aircraft multispectral scanner (MSS) data. Due to the complex spectral characteristics of the wetland vegetation, it was necessary to implement several techniques in the classification of the MSS imagery of the Steel Creek Delta. In particular, when performing unsupervised classification, an iterative cluster busting technique was used which simplified the cluster labeling process. In addition to the MSS data, light detecting and ranging (LIDAR) data were acquired by National Aeronautics and Space Administration (NASA) personnel along two flightlines over the Steel Creek Delta. These data were registered with the wetland classification map and correlated. Statistical analyses demonstrated that the laser derived canopy height information was significantly correlated with the Steel Creek Delta wetland classes encountered along the profiling transect of the LIDAR data.

Jensen, John R.↗