Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Total Dissolved Nitrogen and Ammonia Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for total dissolved nitrogen (TDN) and ammonia concentrations for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. TDN was analyzed using a Shimadzu Total Nitrogen Module (TNM-1) combined with the TOC-VCSH analyzer (Shimadzu Corporation, Japan). TNM-1 is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. Ammonia was determined using a Lachat's QuikChem 8500 Series 2 Flow Injection Analysis System (LACHAT Instruments, QuckChem 8500 series 2, Automated Ion Analyzer, Loveland, Colorado). When ammonia in water samples is heated (60 degrees C) with salicylate and hypochlorite in an alkaline phosphate buffer, an emerald green color is produced which is proportional to the ammonia concentration. The color is intensified by the addition of nitroprusside. Ethylenediaminetetraacetic acid (EDTA) is added to the buffer to prevent the interference of metal ions (Ca, Mg, and Fe etc.). Ammonia-N is then determined by LACHAT flow injection and a colorimetric assay at an absorbance wavelength 660 nm. (Reference: LACHAT Instruments: QuickChem Method 90-107-06-3-A, Determination of Ammonia by Flow Injection Analysis (High Throughput, Salicylate Method/DCIC) (Multi Matrix method). Written by Lynn Egan (Application group), February 08, 2011.) All files are labeled by location and variable, and data reported are the mean values upon replicate measurements. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process as detailed in the methods. This data package contains (1) a zip file (tdn_ammonia_data_2015-2025.zip) containing a total of 299 files: 298 data files of ammonia and TDN data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the determination of Method Detection Limits (MDLs) for TDN data, which has been updated in 2026-08; and (5) PDF and docx files for the detemination of Method Detection Limits (MDLs) for Ammonia and the Interferences by LACHAT Flow Injection Analysis. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 105 locations containing TDN and Ammonia-N data. Update 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses and Determination of Method Detection Limit for Ammonia and the Interferences by LACHAT Flow Injection Analysis documents, which can be accessed as PDFs or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated total dissolved nitrogen and ammonia data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Units were listed incorrectly, but have been fixed to reflect correct units (ug/L). File level metadata (flmd) and data dictionary (dd) files were updated to reflect the updated versions of these files. Available data was added up until 2022-06-01. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-27. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for TDN were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for TDN data were added to this dataset.

54 ENVIRONMENTAL SCIENCES↗

Collaborative Supervised Learning for Sensor Networks

Collaboration methods for distributed machine-learning algorithms involve the specification of communication protocols for the learners, which can query other learners and/or broadcast their findings preemptively. Each learner incorporates information from its neighbors into its own training set, and they are thereby able to bootstrap each other to higher performance. Each learner resides at a different node in the sensor network and makes observations (collects data) independently of the other learners. After being seeded with an initial labeled training set, each learner proceeds to learn in an iterative fashion. New data is collected and classified. The learner can then either broadcast its most confident classifications for use by other learners, or can query neighbors for their classifications of its least confident items. As such, collaborative learning combines elements of both passive (broadcast) and active (query) learning. It also uses ideas from ensemble learning to combine the multiple responses to a given query into a single useful label. This approach has been evaluated against current non-collaborative alternatives, including training a single classifier and deploying it at all nodes with no further learning possible, and permitting learners to learn from their own most confident judgments, absent interaction with their neighbors. On several data sets, it has been consistently found that active collaboration is the best strategy for a distributed learner network. The main advantages include the ability for learning to take place autonomously by collaboration rather than by requiring intervention from an oracle (usually human), and also the ability to learn in a distributed environment, permitting decisions to be made in situ and to yield faster response time.

Wagstaff, Kiri L.↗

Systems and methods for an extensible business application framework

Method and systems for editing data from a query result include requesting a query result using a unique collection identifier for a collection of individual files and a unique identifier for a configuration file that specifies a data structure for the query result. A query result is generated that contains a plurality of fields as specified by the configuration file, by combining each of the individual files associated with a unique identifier for a collection of individual files. The query result data is displayed with a plurality of labels as specified in the configuration file. Edits can be performed by querying a collection of individual files using the configuration file, editing a portion of the query result, and transmitting only the edited information for storage back into a data repository.

Bell, David G.↗

Ameloblastin binding to biomimetic models of cell membranes – A continuum of intrinsic disorder

A 37-residue amino acid sequence corresponding to the segment encoded by exon-5 of murine ameloblastin (Ambn), AB2 (Y67-Q103), has been implicated with membrane association, ameloblastin self-assembly, and amelogenin-binding. Here, our aim was to characterize, at the residue level, the structural behavior of AB2 bound to chemical mimics of biological membranes using NMR spectroscopy. To better define the structure of AB2 using NMR-based methods, recombinant 13 C- and 15 N-labelled AB2 (*AB2) was prepared and data collected free in solution and with deuterated dodecylphosphocholine (dPC) micelles, deuterated bicelles, and both small and large unilamellar vesicles. Amide chemical shift and intensity perturbations observed in 1 H- 15 N HSQC spectra of *AB2 in the presence of bicelles and dPC micelles suggest that a region of *AB2, S6-E36 (murine Ambn S68 – E98), associates with the membrane biomimetics. A CSI-3 analysis of the NMR chemical shift assignments for *AB2 free in solution and bound to dPC micelles indicated the peptide remains disordered except for the adoption of a short, 12-residue α-helix, F10-G21 (murine Ambn F72-G83). In dPC micelles, the NOE NMR data was void of patterns characteristic of long-lived helical structure indicating this helix was transient in nature. A continuum of intrinsic disorder in the membrane-bound state may be responsible for ameloblastin’s ability to dynamically interact with multiple partners at the same site during amelogenesis.

59 BASIC BIOLOGICAL SCIENCES↗

Regime Characterization of Offshore Wind Resource Using Unsupervised Learning

Predictability of wind resource conditions is critical for offshore wind design and operations. While many studies of extreme wind conditions focus on specific events such as low-level jets or ramps, these rely on threshold definitions that limit generality. Here we present a data-driven framework that combines principal component analysis (PCA), self-organizing maps (SOM), and k-means clustering to classify wind resource conditions as typical and anomalous from climatological data. Anomalies are defined not by fixed thresholds but by flagging samples located far from SOM node centers inside the baseline SOM structure. This reframes extremes as rare ebents and hence, likely difficult to anticipate by numerical weather prediction models. We applied this approach to 23 years (2000–2022) of hourly profiles from the NOW-23 hindcast model at the Humboldt Wind Energy Area. Classification is conducted on a feature space consisting of 10 m wind speed and direction, bulk shear and veer across 30–270 m, and a low-level jet index. Dimensionality reduction is achieved through PC. A 2 × 3 OM lattice trained on the PCA vectors identified six baseline regimes spanning weak to strong flow states. High quantization-error profiles are identified and re-clustered into four anomalous regimes. The baseline regimes exhibited clear seasonal and diurnal cycles. Meanwhile, the anomalous regimes represented <10 % of all hours but showed distinct combinations of speed, shear, and veer, when compared to the baseline regimes. Anomalous regimes are typically short-lived (~few hours), yet their transitions can lead to hub-height wind changes of −18 to +9 m s -1 . For a representative 15 MW turbine, these shifts imply rapid swings in capacity factor from near-full output to negligible generation. Validation with lidar buoy data showed 51% agreement in SOM labels across ~6,000 overlapping hours, with most mismatches confined to adjacent speed classes. HRRR comparisons further revealed that anomalous regimes were disproportionately associated with forecast biases exceeding 5 m s -1 . Together, these results reframe extremes in offshore wind from absolute maxima or minima to weather states that are difficult to anticipate from models.

17 WIND ENERGY↗

Metabolomics analysis of P. tremula × P. alba ‘717-1B4’ tissue culture hybrids grown in liquid media supplemented with EGGC and d4-EGGC

A metabolite linked to ethylene metabolism in Populus was recently structurally characterized as 2-hydroxyethyl β-D-glucopyranoside, an ethylene glycol glucose conjugate (EGGC). The dataset presented here is associated with the metabolomics analysis of various tissues (i.e., roots, stems/leaves) of plants grown in liquid media supplemented with EGGC and stable-isotope labelled EGGC (i.e., d4-EGGC). Data were collected using a Thermo Scientific gas chromatograph (GC) coupled to a Q Exactive Orbitrap mass spectrometer (MS). Samples were collected at three different timepoints and silylated prior to GCMS analysis.

09 BIOMASS FUELS↗

Evaluation of high specific-heat ceramic for regenerator use at temperatures between 2-30 K

Specific heat, thermal conductivity (both in the range 2-30 K), and microhardness data were measured on the ceramics labelled LS-8, LS-8A, and LS-8A doped with CsI, SnCl2, and AgCl. A work hardened sample of LS-8A was also studied in an effort to determine the feasibility of using these types of LS-8 materials to replace Pb spheres in the regenerator of the JPL cryocooler. The LS-8A materials are all more than an order of magnitude harder than Pb, and the dopants do not significantly improve the hardness. However, the SnCl2 dopant has a remarkable effect in improving the specific heat and thermal conductivity of LS-8A. The SnCl2 doping level which maximized the regenerator enthalpy change in going from an unloaded to a loaded condition was found to be 0.2 percent SnCl2 in LS-8A. It was also found that the enthalpy change for a regenerator employing the LS-8A material is more than three times larger than for the Pb spheres case. The use of rods, rather than spheres, of optimally doped LS-8A in regenerators is discussed.

Lawless, W. N.↗

Aircraft wing weight build-up methodology with modification for materials and construction techniques

An aircraft wing weight estimating method based on a component buildup technique is described. A simplified analytically derived beam model, modified by a regression analysis, is used to estimate the wing box weight, utilizing a data base of 50 actual airplane wing weights. Factors representing materials and methods of construction were derived and incorporated into the basic wing box equations. Weight penalties to the wing box for fuel, engines, landing gear, stores and fold or pivot are also included. Methods for estimating the weight of additional items (secondary structure, control surfaces) have the option of using details available at the design stage (i.e., wing box area, flap area) or default values based on actual aircraft from the data base.

York, P.↗

Selection of the Australian indicator region

Each Australian state was examined for the availability of LANDSAT data, area, yield, and production characteristics, statistics, crop calendars, and other ancillary data. Agrophysical conditions that could influence labeling and classification accuracies were identified in connection with the highest producing states as determined from available Australian crop statistics. Based primarily on these production statistics, Western Australia and New South Wales were selected as the wheat indicator region for Australia. The general characteristics of wheat in the indicator region, with potential problems anticipated for proportion estimation are considered. The varieties of wheat, the diseases and pests common to New South Wales, and the wheat growing regions of both states are examined.

Reed, C. R.↗

Bar-Code-Scribing Tool

Proposed hand-held tool applies indelible bar code to small parts. Possible to identify parts for management of inventory without tags or labels. Microprocessor supplies bar-code data to impact-printer-like device. Device drives replaceable scribe, which cuts bar code on surface of part. Used to mark serially controlled parts for military and aerospace equipment. Also adapts for discrete marking of bulk items used in food and pharmaceutical processing.

Badinger, Michael A.↗

Fluorescence Studies of Protein Crystal Nucleation

Fluorescence can be used to study protein crystal nucleation through methods such as anisotropy, quenching, and resonance energy transfer (FRET), to follow pH and ionic strength changes, and follow events occurring at the growth interface. We have postulated, based upon a range of experimental evidence that the growth unit of tetragonal hen egg white lysozyme is an octamer. Several fluorescent derivatives of chicken egg white lysozyme have been prepared. The fluorescent probes lucifer yellow (LY), cascade blue, and 5-((2-aminoethyl)aminonapthalene-1-sulfonic acid (EDANS), have been covalently attached to ASP 101. All crystallize in the characteristic tetragonal form, indicating that the bound probes are likely laying within the active site cleft. Crystals of the LY and EDANS derivatives have been found to diffract to at least 1.7 A. A second group of derivatives is to the N-terminal amine group, and these do not crystallize as this site is part of the contact region between the adjacent 43 helix chains. However derivatives at these sites would not interfere with formation of the 43 helices in solution. Preliminary FRET studies have been carried out using N-terminal bound pyrene acetic acid (Ex 340 nm, Em 376 nm) lysozyme as a donor and LY (Ex -425 nm, Em 525 nm) labeled lysozyme as an acceptor. FRET data have been obtained at pH 4.6, 0.1 M NaAc buffer, at 5 and 7% NaCl, 4 C. The corresponding Csat values are 0.471 and 0.362 mg/ml (approximately 3.3 and approximately 2.5 x 10(exp -5) M respectively). The data at both salt concentrations show a consistent trend of decreasing fluorescence intensity of the donor species (PAA) with increasing total protein concentration. This decrease is more pronounced at 7% NaCl, consistent with the expected increased intermolecular interactions at higher salt concentrations reflected in the lower solubility. The calculated average distance between any two protein molecules at 5 x 10(exp -6) M is approximately 70nm, well beyond the range where any FRET can be expected. Results from these and ongoing studies will be presented.

Pusey, Marc L.↗

Model-Driven Development For PDS4 Software And Services

Software and services that access Planetary Data System (PDS) PDS4 data products need to parse product labels to retrieve, interpret, and process the referenced digital objects. Under PDS4 a driving principle is that the product label provide all of the information necessary for these functions to be performed accurately. However, significantly more information is available in the PDS4 Information Model (IM)[1], the controlling document used to define, create, and syntactically and semantically verify the product labels. This additional information in the IM is made available for use, by both software and services, to configure, promote resiliency, and improve interoperability.

Padams, Jordan↗

Correlation between aircraft MSS and LIDAR remotely sensed data on a forested wetland in South Carolina

Wetlands in a portion of the Savannah River swamp forest, the Steel Creek Delta, were mapped using April 26, 1985 high-resolution aircraft multispectral scanner (MSS) data. Due to the complex spectral characteristics of the wetland vegetation, it was necessary to implement several techniques in the classification of the MSS imagery of the Steel Creek Delta. In particular, when performing unsupervised classification, an iterative cluster busting technique was used which simplified the cluster labeling process. In addition to the MSS data, light detecting and ranging (LIDAR) data were acquired by National Aeronautics and Space Administration (NASA) personnel along two flightlines over the Steel Creek Delta. These data were registered with the wetland classification map and correlated. Statistical analyses demonstrated that the laser derived canopy height information was significantly correlated with the Steel Creek Delta wetland classes encountered along the profiling transect of the LIDAR data.

Jensen, John R.↗

Correlation between aircraft MSS and LIDAR remotely sensed data on a forested wetland in South Carolina

Wetlands in a portion of the Savannah River swamp forest, the Steel Creek Delta, were mapped using April 26, 1985 high-resolution aircraft multispectral scanner (MSS) data. Due to the complex spectral characteristics of the wetland vegetation, it was necessary to implement several techniques in the classification of the MSS imagery of the Steel Creek Delta. In particular, when performing unsupervised classification, an iterative cluster busting technique was used which simplified the cluster labeling process. In addition to the MSS data, light detecting and ranging (LIDAR) data were acquired by National Aeronautics and Space Administration (NASA) personnel along two flightlines over the Steel Creek Delta. These data were registered with the wetland classification map and correlated. Statistical analyses demonstrated that the laser derived canopy height information was significantly correlated with the Steel Creek Delta wetland classes encountered along the profiling transect of the LIDAR data.

Jensen, John R.↗

Processing AIRS Scientific Data Through Level 3

The Atmospheric Infra-Red Sounder (AIRS) Science Processing System (SPS) is a collection of computer programs, known as product generation executives (PGEs). The AIRS SPS PGEs are used for processing measurements received from the AIRS suite of infrared and microwave instruments orbiting the Earth onboard NASA's Aqua spacecraft. Early stages of the AIRS SPS development were described in a prior NASA Tech Briefs article: Initial Processing of Infrared Spectral Data (NPO-35243), Vol. 28, No. 11 (November 2004), page 39. In summary: Starting from Level 0 (representing raw AIRS data), the AIRS SPS PGEs and the data products they produce are identified by alphanumeric labels (1A, 1B, 2, and 3) representing successive stages or levels of processing. The previous NASA Tech Briefs article described processing through Level 2, the output of which comprises geo-located atmospheric data products such as temperature and humidity profiles among others. The AIRS Level 3 PGE samples selected information from the Level 2 standard products to produce a single global gridded product. One Level 3 product is generated for each day s collection of Level 2 data. In addition, daily Level 3 products are aggregated into two multiday products: an eight-day (half the orbital repeat cycle) product and monthly (calendar month) product.

Granger, Stephanie↗

Blueprints for Training Information Bottlenecks for Collider Analyses

Dimensionality reduction is a crucial aspect of data analysis in high energy physics, even if accompanied by information loss. Several methods, including histogram- and kernel-based analyses, are only computationally feasible for low-dimensional data. Furthermore, simulation models used in HEP can often only be validated for low-dimensional data. We provide several blueprints for using machine learning to create low-dimensional data representations (continuous event variables and discrete classification labels) for use in signal discovery and parameter estimation tasks. We also describe how to design the learned representation to facilitate a) searches with unknown model parameters and b) validation of simulation models in data control regions.

43 PARTICLE ACCELERATORS↗

AgRISTARS: Foreign commodity production forecasting. The 1980 US corn and soybeans exploratory experiment

The U.S. corn and soybeans exploratory experiment is described which consisted of evaluations of two technology components of a production forecasting system: classification procedures (crop labeling and proportion estimation at the level of a sampling unit) and sampling and aggregation procedures. The results from the labeling evaluations indicate that the corn and soybeans labeling procedure works very well in the U.S. corn belt with full season (after tasseling) LANDSAT data. The procedure should be readily adaptable to corn and soybeans labeling required for subsequent exploratory experiments or pilot tests. The machine classification procedures evaluated in this experiment were not effective in improving the proportion estimates. The corn proportions produced by the machine procedures had a large bias when the bias correction was not performed. This bias was caused by the manner in which the machine procedures handled spectrally impure pixels. The simulation test indicated that the weighted aggregation procedure performed quite well. Although further work can be done to improve both the simulation tests and the aggregation procedure, the results of this test show that the procedure should serve as a useful baseline procedure in future exploratory experiments and pilot tests.

Malin, J. T.↗