Search NASA⌕ Search

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42

DataHub--Knowledge-Based Science Data Management for Exploratory Data Analysis

It is our belief that new modes of research and new tools will be required to handle the massive amount of diverse data that is to be stored, organized, accessed, distributed, visualized, and analyzed. The fundamental innovation required is the integration of three automation technologies, videlicet knowledge-based expert systems, science visualization and science data management. This integration is based on a concept caled the DataHub, which we describe here.

DataHub↗

The Open Data Repository - an Open Science Platform for Long-Tail Research Data

Introduction: ‘Long-tail’ research is performed by individual PIs and small research teams, producing what are often highly diverse, but relatively small da-tasets that span a variety of traditional scientific disci-plines. ‘Long-tail’ research is a fundamental part of realizing NASA’s goals in planetary sciences and a key input feeding into mission life cycles. The ‘long-tail’ has traditionally lacked the resources (available to larg-er groups and missions) to overcome barriers inhibit-ing the adoption of open science practices, including lack of acknowledgment, time, money, guidance, ex-pertise, and trust in available platforms1. Here we show how the Open Data Repository’s (ODR) data publish-ing platform could help lower some of these barriers and accelerate the adoption of open-science practices in planetary science.

‘Long-tail’ research↗

Data‐Driven Insights into Rare Earth Mineralization: Machine Learning Applications Using Functional Material Synthesis Data

Understanding rare‐earth element (REE) mineralization mechanisms is essential for developing efficient separation strategies. Although the geochemical pathways that generate REE deposits are qualitatively known, quantitative links between specific conditions and mineralization outcomes remain limited. Herein, the repurpose laboratory REE hydrothermal synthesis data—originally collected for functional‐materials fabrication—as a surrogate for studying mineralization with data‐driven methods. The compiled 1,200+ hydrothermal reaction records and trained three machine‐learning models—K‐nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGB)—to predict product elements and phases from precursors, additives, reaction conditions, and engineered features. Validation shows XGB achieves the highest accuracy. Feature importance indicates thermodynamic properties of cations and anions dominate model decisions. Correlations reveal positive relationships among precursor concentration, reaction time, pH, and temperature, consistent with classical crystallization behavior. XGB‐based regressors are built to predict crystallization temperature and pH from precursor/product attributes. Performance is strongest when similar training examples exist, while accuracy declines for underrepresented reactions, notably REE carbonates and heavy‐REE systems. Overall, the study shows that functional‐materials datasets can illuminate REE mineralization and provide priors for exploration and processing. Expanding datasets with less‐studied chemistries and conditions will improve generality and support deposit discovery and more efficient REE recovery.

feature importance analysis↗

Utilization of Data Augmentation Techniques in Automated Inspection Systems for Defect Detection in Metals With Limited Data

Accurate identification of defects on metal surfaces is of great interest to many industry sectors, such as the automotive and aerospace industries. In contrast to conventional manual inspection techniques, recent automated inspection systems employ deep learning models trained to detect defects rapidly and precisely. The development of these models often requires a substantial image dataset to acquire adequate knowledge of defect features and enhance their predictive accuracy. When data is limited, augmentation techniques are often used to improve the precision and accuracy of defect detection systems. This study examined the prediction performance of two object detection models, namely Faster Region‐based Convolutional Neural Network (Faster R‐CNN) and You Only Look Once version 8 (YOLOv8), to identify dent defects in limited images of cast iron cylinder head surfaces. The original image set contains 46 images with 563 dents. To overcome limited data availability, common image augmentation techniques along with a copy‐paste method were applied. Results show that standard augmentation improved YOLOv8 accuracy by 8.00% and average precision (AP) by 3.00%. On the other hand, the copy‐paste technique achieved a 20.00% increase in accuracy and a 1% increase in AP with just 200 synthetic dents. Furthermore, these results provide support for using the copy‐paste augmentation strategy to enhance defect detection performance, with a limited dataset, contributing to more accurate defect identification in remanufacturing processes.

36 MATERIALS SCIENCE↗

Impact of new experimental data on the C2HDM: the strong interdependence between LHC Higgs data and the electron EDM

The complex two-Higgs doublet model (C2HDM) is one of the simplest extensions of the Standard Model with a source of CP-violation in the scalar sector. It has a $\mathbb{Z}$ 2 symmetry, softly broken by a complex coefficient. There are four ways to implement this symmetry in the fermion sector, leading to models known as Type-I, Type-II, Lepton Specific and Flipped. In the latter three models, there is a priori the surprising possibility that the 125 GeV Higgs boson couples mostly as a scalar to top quarks, while it couples mostly as a pseudoscalar to bottom quarks. This “maximal” scenario was still possible with the data available in 2017. Since then, there have been more data on the 125 GeV Higgs boson, direct searches for CP-violation in angular correlations of $τ$-leptons produced in Higgs boson decays, new results on the electron electric dipole moment, new constraints from LHC searches for additional Higgs bosons and new results on $b$ → $sγ$ transitions. Highlighting the crucial importance of the physics results of LHC’s Run 2, we combine all these experiments and show that the “maximal” scenario is now excluded in all models. Still, one can have a pseudoscalar component in $hτ\overline{τ}$ couplings in the Lepton-Specific case as large as 87% of the scalar component for all mass orderings of the neutral scalar bosons.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

SIMS Data Correction Procedure for Quasi‐Simultaneous Arrival ( QSA ) Under‐counting and Ramifications of Misapplication

Isotope geochemistry requires isotope ratios measured using secondary ion mass spectrometry (SIMS) to be made with optimal precision and accuracy. Under some analytical conditions when using electron multiplier detectors, secondary ions may be under‐counted because of quasi‐simultaneous arrival (QSA) at the first dynode. The relative magnitude of the associated QSA correction to raw measured isotopic ratios can be up to seventy permil or more. Therefore, not applying the correction, or misapplication of it could lead to significant inaccuracies in published isotope ratio data. Examples and ramifications of the latter are described in addition to a straightforward procedure for QSA under‐counting correction.

Geochemistry & Geophysics↗

Data for The Stem Cell-Type Transcriptome of Bioenergy Sorghum Reveals the Spatial Regulation of Secondary Cell Wall Networks

Bioenergy sorghum is a low-input, drought-resilient, deep-rooting annual crop that has high biomass yield potential enabling the sustainable production of biofuels, biopower, and bioproducts. Bioenergy sorghum’s 4-5 m stems account for ~80% of the harvested biomass. Stems accumulate high levels of sucrose that could be used to synthesize bioethanol and useful biopolymers if information about stem cell-type gene expression and regulation was available to enable engineering. To obtain this information, Laser Capture Microdissection (LCM) was used to isolate and collect transcriptome profiles from five major cell types that are present in stems of the sweet sorghum Wray. Transcriptome analysis identified genes with cell-type specific and cell-preferred expression patterns that reflect the distinct metabolic, transport, and regulatory functions of each cell type. Analysis of cell-type specific gene regulatory networks (GRNs) revealed that unique TF families contribute to distinct regulatory landscapes, where regulation is organized through various modes and identifiable network motifs. Cell-specific transcriptome data was combined with a stem developmental transcriptome dataset to identify the GRN that differentially activates the secondary cell wall (SCW) formation in stem xylem sclerenchyma and epidermal cells. The cell-type transcriptomic dataset provides a valuable source of information about the function of sorghum stem cell types and GRNs that will enable the engineering of bioenergy sorghum stems.

Software↗

Data for Propagation Method and Planting Density Influence Canopy Developmental Transition and Biomass Productivity in Miscanthus × giganteus

Understanding how establishment practices influence the mechanisms underlying Miscanthus × giganteus (miscanthus) productivity and canopy development is critical for optimizing management. Data was collected during the juvenile (2011–2013) and mature (2024) phases of a long-term field experiment established in Urbana, Illinois, to evaluate the effects of propagation method (plug propagation [PP] and rhizome propagation [RP]), planting density (1.0, 0.75, and 0.25 plants m⁻²), and nitrogen application (0 and 67 kg N ha⁻¹) on end-of-season biomass yield, tiller mass, tiller density, and tiller height. Linear regression models identified the dominant predictors of yield across stand ages and management regimes. Planting density, nitrogen (N) application, and propagation method significantly influenced early yield and canopy development. During the juvenile phase, biomass yield was driven by tiller density due to canopy expansion; in the mature phase, yield became driven by tiller mass. The PP plots produced higher tiller density than the RP plots, resulting in faster canopy closure and higher juvenile-phase yields. Rhizome-propagated (RP) plots produced lower tiller density, but individual tillers were 3.3–6.4 g tiller−1 heavier than PP tillers. After the canopy reached equilibrium, the PP and RP yields were similar because greater RP tiller mass compensated for its lower tiller density. Higher planting density resulted in greater yield and tiller density during the second year (2012), but this effect was absent from the third year (2013) onward. In the juvenile phase, N fertilization enhanced yield by 1.6–3.4 Mg ha−1. Initiating fertilization in 2013 on unfertilized plots produced biomass similar to that in fertilized plots, suggesting yield recovery in the mature phase. These findings revealed that establishment strategies, including propagation method and planting density, influence juvenile miscanthus canopy development and productivity, transitioning from tiller-density- to mass-dominated yields, but not mature phase productivity.

Miscanthus↗

Data for Creating Yellow Seed Camelina sativa with Enhanced Oil Accumulation by CRISPR-Mediated Disruption of Transparent Testa 8

Camelina ( Camelina sativa L.), a hexaploid member of the Brassicaceae family, is an emerging oilseed crop being developed to meet the increasing demand for plant oils as biofuel feedstocks. In other Brassicas, high oil content can be associated with a yellow seed phenotype, which is unknown for camelina. We sought to create yellow seed camelina using CRISPR/Cas9 technology to disrupt its Transparent Testa 8 (TT8) transcription factor genes and to evaluate the resulting seed phenotype. We identified three TT8 genes, one in each of the three camelina subgenomes, and obtained independent CsTT8 lines containing frameshift edits. Disruption of TT8 caused seed coat colour to change from brown to yellow reflecting their reduced flavonoid accumulation of up to 44%, and the loss of a well-organized seed coat mucilage layer. Transcriptomic analysis of CsTT8-edited seeds revealed significantly increased expression of the lipid-related transcription factors LEC1, LEC2, FUS3, and WRI1 and their downstream fatty acid synthesis-related targets. These changes caused metabolic remodelling with increased fatty acid synthesis rates and corresponding increases in total fatty acid (TFA) accumulation from 32.4% to as high as 38.0% of seed weight, and TAG yield by more than 21% without significant changes in starch or protein levels compared to parental line. These data highlight the effectiveness of CRISPR in creating novel enhanced-oil germplasm in camelina. The resulting lines may directly contribute to future net-zero carbon energy production or be combined with other traits to produce desired lipid-derived bioproducts at high yields.

Biofuels↗

Data for NB6 HBRR Science Design ORNL/TM-2025/3807

Data for the report (ORNL/TM-2025/3807) that describes the calculations and the Monte Carlo Ray Tracing simulations performed using the McStas package to determine the coatings and geometry for the NB-6 guide. It provides the information to inform the mechanical design, validation tests and verification that it meets the science requirements.

47 OTHER INSTRUMENTATION↗

Dual-Fuel Ammonia Equivalence Ratio Sweep Data

Ammonia (NH3) has garnered significant interest as an alternative fuel for meeting international emissions reduction mandates in sectors with high weight and distance requirements, such as shipping. Technical barriers and unanswered questions remain on the combustion strategies that can maximize ammonia utilization and minimize emissions. Prior research studies at the US Department of Energy’s Oak Ridge National Laboratory have shown strong performance with NH3 under dual-fuel mode using conventional diesel combustion (CDC) manifold air pressure settings. Diesel airflow was initially used to simplify retrofitting (no turbocharger modification), which resulted in air-fuel equivalence ratios (λ) greater than 1.5. To characterize potential improvements in dual-fuel NH3 combustion performance at richer in-cylinder conditions, a global λ sweep compared the use of early (E-pilot) and late (L-pilot) single diesel injections. The experiments were conducted at 1200 RPM and 12.8 ± 0.2 bar (75 % load), and λ was varied by decreasing the commanded air flow to the engine at greater than 90 % ammonia energy substitution level. A diesel injection timing sweep was conducted for both the injection strategies at fixed λ, and the timing with the lowest engine-out N2O emissions was identified. The results indicated an optimal balance between CO2,eq and thermal efficiency benefits both E-pilot and L-pilot injection strategy cases compared with CDC at a λ of 1.4. The indicated nitrogen-based emissions exhibited a strong correlation to the ratio of CA5–50 and ignition delay for L-pilot, but no apparent trend emerged for the E-pilot injection strategy at the tested boundary conditions. This dataset includes the raw experimental data and documentation of the experiment conditions and methods.

30 DIRECT ENERGY CONVERSION↗

EAGLE-I Power Outage Data 2025

The provided EAGLE-I historic dataset includes power outage information at the county level for 2025 at 15-minute intervals collected by the EAGLE-I program at ORNL. The data has been collected from utility's public outage maps using an ETL process. The dataset details FIPS code, county name, state name, total number of customers without power, and a date/timestamp. For detailed metadata, refer to the linked metadata DOI.

EAGLE-I↗

Open-Source Data for MAC-POSTS: Mobility Data Analytics Center - Prediction, Optimization, and Simulation Toolkit for Transportation Systems

MAC-POSTS (Mobility Data Analytics Center - Prediction, Optimization, and Simulation toolkit for Transportation Systems) is a toolkit for dynamic transportation network modeling. Developed by the Mobility Data Analytics Center (MAC) at Carnegie Mellon University, this package implements many classic dynamic transportation network models, as well as new models proposed by MAC members. It has served as one building block for many other models and research projects. As such, this package used to be treated as an internal research project of the MAC lab, and admittedly, the code base is messy, and the interface is hard to use. However, we are working hard to make it a generally usable and useful toolkit for dynamic transportation network modeling. We would really appreciate any feedback, comments, suggestions, or criticisms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗