Search NASA⌕ Search

SEARCH · Search NASA

Results for “data statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41

Using the Bootstrap Method for a Statistical Significance Test of Differences between Summary Histograms

A new method is proposed to compare statistical differences between summary histograms, which are the histograms summed over a large ensemble of individual histograms. It consists of choosing a distance statistic for measuring the difference between summary histograms and using a bootstrap procedure to calculate the statistical significance level. Bootstrapping is an approach to statistical inference that makes few assumptions about the underlying probability distribution that describes the data. Three distance statistics are compared in this study. They are the Euclidean distance, the Jeffries-Matusita distance and the Kuiper distance. The data used in testing the bootstrap method are satellite measurements of cloud systems called cloud objects. Each cloud object is defined as a contiguous region/patch composed of individual footprints or fields of view. A histogram of measured values over footprints is generated for each parameter of each cloud object and then summary histograms are accumulated over all individual histograms in a given cloud-object size category. The results of statistical hypothesis tests using all three distances as test statistics are generally similar, indicating the validity of the proposed method. The Euclidean distance is determined to be most suitable after comparing the statistical tests of several parameters with distinct probability distributions among three cloud-object size categories. Impacts on the statistical significance levels resulting from differences in the total lengths of satellite footprint data between two size categories are also discussed.

Xu, Kuan-Man↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗

Orbital and cloud cover sampling analyses for space lidar missions

The sampling capabilities of an orbital lidar mission are evaluated. Spatial and temporal sampling data from a lidar spacecraft orbit simulation are combined with global, statistical cloud cover data to yield a quantification of lidar measurement opportunities for both partly cloudy and mostly overcast viewing conditions. The optimum launch time (month and local hour) is determined to maximize lidar measurement opportunities for specified cloud cover conditions. Results indicate that the time of year selected for the lidar mission is very important in maximizing acceptable data return, whereas the effect of launch time of day on mission optimization is generally not as strong as the seasonal effect.

Lawrence, G. F.↗

Improving Weather and Climate Prediction with the AIRS on Aqua

The Atmospheric Infrared Sounder (AIRS) on the EOS Aqua Spacecraft was launched on May 4, 2002. Early in the mission, the AIRS instrument demonstrated its value to the weather forecasting community with better than 6 hours of improvement on the 5 day forecast. Now with over six years of consistent and stable data from AIRS, scientists are able to examine processes governing weather and climate and look at seasonal and interannual trends from the AIRS data with high statistical confidence. Naturally, long-term climate trends require a longer data set, but indications are that the Aqua spacecraft and the AIRS instrument should last beyond 2016. This paper briefly describes the AIRS products, reviews past science and weather accomplishments from AIRS data product users and highlights recent findings in these areas.

Temperature↗

A Statistical Analysis of Impact Ice Adhesion Strength Data Acquired with a Modified Lap Joint Test

Numerous methodologies have been utilized to measure the adhesion strength of impact ice, and the data reported in the literature varies significantly from method to method. In order to initiate an investigation to determine the cause of this disparity, a lap-joint shear test methodology that was recently developed was utilized in the Icing Research Tunnel at the NASA Glenn Research Center. Data was obtained while varying the temperature, test section velocity, liquid water content, and cloud droplet median volumetric diameter, among other parameters, over five campaigns. A new data set acquired using this new test method is presented. The results are analyzed with standard statistical methodologies and demonstrate a strong correlation between the apparent adhesion strength of ice and temperature, annealing time, and heating time. Observations during the test and analysis of the results suggest the presence of large residual stresses in the samples, which is in agreement with prior work. Due to the nature of the test methodology, and all known ice adhesion test methodologies, the measured, or apparent, adhesion strength is geometry dependent, a fact emphasized by the results presented here. The data was organized both cumulatively and independently by result code. Trends in the data are discussed in the context of causal physical mechanisms. The data shows the apparent adhesion strength was not linear with temperature, liquid water content, annealing time, and other variables. Short-term heating of the samples was shown to have negligible effect, and the apparent adhesion strength had a flat trend with velocity.

Ice Adhesion↗

Parallel line analysis: multifunctional software for the biomedical sciences

An easy to use, interactive FORTRAN program for analyzing the results of parallel line assays is described. The program is menu driven and consists of five major components: data entry, data editing, manual analysis, manual plotting, and automatic analysis and plotting. Data can be entered from the terminal or from previously created data files. The data editing portion of the program is used to inspect and modify data and to statistically identify outliers. The manual analysis component is used to test the assumptions necessary for parallel line assays using analysis of covariance techniques and to determine potency ratios with confidence limits. The manual plotting component provides a graphic display of the data on the terminal screen or on a standard line printer. The automatic portion runs through multiple analyses without operator input. Data may be saved in a special file to expedite input at a future time.

NASA Discipline Cell Biology↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN↗

Statistical Analysis of a Large Sample Size Pyroshock Test Data Set Including Post Flight Data Assessment

The Earth Observing System (EOS) Terra spacecraft was launched on an Atlas IIAS launch vehicle on its mission to observe planet Earth in late 1999. Prior to launch, the new design of the spacecraft's pyroshock separation system was characterized by a series of 13 separation ground tests. The analysis methods used to evaluate this unusually large amount of shock data will be discussed in this paper, with particular emphasis on population distributions and finding statistically significant families of data, leading to an overall shock separation interface level. The wealth of ground test data also allowed a derivation of a Mission Assurance level for the flight. All of the flight shock measurements were below the EOS Terra Mission Assurance level thus contributing to the overall success of the EOS Terra mission. The effectiveness of the statistical methodology for characterizing the shock interface level and for developing a flight Mission Assurance level from a large sample size of shock data is demonstrated in this paper.

Hughes, William O.↗

Statistical Studies on Thin Cirrus from MODIS Data

The 1.38 micron channel on the MODerate resolution Imaging Spectroradiomater (MODIS) is an ideal channel to identify and quantify thin cirrus on a global basis. This channel is used to produce the cirrus reflectance product in MOD06 and also used extensively by the MODIS aerosol algorithms to mask clouds for the MOD04 product. The aerosol product uses a lower threshold of the 1.38 micron channel reflectance of 0.01. A cirrus channel reflectance of 0.01 corresponds to approximately an aerosol optical thickness of 0.10. Therefore, the ambiguity due to the minor cirrus contamination may introduce artificial optical thickness in the aerosol products. The questions arise: How prevalent are the thinnest cirrus clouds over the globe? Do they persist over specific regions and seasons? Can we distinguish between the noise of the channel and the actual cloudiness by extrapolating the cloudiness signal to very dark scenes, statistically. We analyze the Terra data, over land and ocean to answer these questions.

Li, Rong-Rong↗

Testing hadronic-model predictions of depth of maximum of air-shower profiles and ground-particle signals using hybrid data of the Pierre Auger Observatory

We test the predictions of hadronic interaction models regarding the depth of maximum of air-shower profiles, X max , and ground-particle signals in water-Cherenkov detectors at 1000 m from the shower core, S ( 1000 ) , using the data from the fluorescence and surface detectors of the Pierre Auger Observatory. The test consists of fitting the measured two-dimensional ( S ( 1000 ) , X max ) distributions using templates for simulated air showers produced with hadronic interaction models pos-, et--04, 2.3d and leaving the scales of predicted X max and the signals from hadronic component at ground as free-fit parameters. The method relies on the assumption that the mass composition remains the same at all zenith angles, while the longitudinal shower development and attenuation of ground signal depend on the mass composition in a correlated way. The analysis was applied to 2239 events detected by both the fluorescence and surface detectors of the Pierre Auger Observatory with energies between 10 18.5 eV to 10 19.0 eV and zenith angles below 60°. We found, that within the assumptions of the method, the best description of the data is achieved if the predictions of the hadronic interaction models are shifted to deeper X max values and larger hadronic signals at all zenith angles. Given the magnitude of the shifts and the data sample size, the statistical significance of the improvement of data description using the modifications considered in the paper is larger than 5 σ even for any linear combination of experimental systematic uncertainties. Published by the American Physical Society 2024

79 ASTRONOMY AND ASTROPHYSICS↗

Algorithms for Non-Negative Matrix Factorization on Noisy Data With Negative Values

Non-negative matrix factorization (NMF) is a dimensionality reduction technique that has shown promise for analyzing noisy data, especially astronomical data. For these datasets, the observed data may contain negative values due to noise even when the true underlying physical signal is strictly positive. Prior NMF work has not treated negative data in a statistically consistent manner, which becomes problematic for low signal-to-noise data with many negative values. In this paper we present two algorithms, Shift-NMF and Nearly-NMF, that can handle both the noisiness of the input data and also any introduced negativity. Both of these algorithms use the negative data space without clipping or masking and recover non-negative signals without any introduced positive offset that occurs when clipping or masking negative data. We demonstrate this numerically on both simple and more realistic examples, and prove that both algorithms have monotonically decreasing update rules.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Statistical Analysis of Imaging Laser Scan Data of an Exhaust Tunnel at the SRS

• The SRS H-Canyon Building is a critical facility under the responsibility of DOE-EM. • It includes an Air Exhaust Tunnel (HCAEX) that allows for ventilation of the process airflow. • Inspections are performed remotely because of hazards, e.g. radioactivity, debris, high airflow, and nitric acid vapors.

Wells, William Willie [Savannah River National Lab↗

ATS-F millimeter wave propagation experiment data processing.

Review of the data processing program for the Applications Technology Satellite (ATS-F) scheduled for launch in 1973. The program consists of: (1) short term processing, which utilizes analog data recording and produces amplitude scintillation and time-frequency correlation studies; and (2) long term processing, which utilizes digitally recorded data to provide hourly and daily statistical output functions. The data acquisition and processing program planned for the ATS-F experiment provides a highly automated environment which emphasizes the unique requirements for the characterization of the propagation medium in the frequency bands required for the next generation of space communication systems.

Ippolito, L. J.↗

Apollo experience report: Flight instrumentation calibration

Three types of instrumentation-calibration data were used in the Apollo Program to provide the correct engineering data for tests and mission support. The command and service module instrumentation-component procurement specifications required individual-component calibration, and calibration data for these individual components (conventional-calibration data) were always used for mission data support. A mean standard type of calibration data derived from a statistical sampling of conventional-calibration data was used for test and checkout during the latter part of the Apollo Program. The lunar module instrumentation procurement specification permitted the use of standard-calibration data. These data were applicable to similarly instrumented measurements. The definition, merit, and application of each type of data are discussed.

Demoss, J. F.↗

Statistical processing of Pioneer front film data, part 1

A program was constructed to read the data on impacts and positional information on Pioneer and to classify these events according to a number of different criteria. The program is flexible enough to permit the introduction of further criteria and additional classifications, should this appear desirable. Not all cards correspond to particle impacts on the Pioneer sensors, many are inserted only to supply Pioneer position information.

Wolf, H.↗

Electron precipitation patterns and substorm morphology.

Statistical analysis of data from the auroral particles experiment aboard OGO 4, performed in a statistical framework interpretable in terms of magnetospheric substorm morphology, both spatial and temporal. Patterns of low-energy electron precipitation observed by polar satellites are examined as functions of substorm phase. The implications of the precipitation boundaries identifiable at the low-latitude edge of polar cusp electron precipitation and at the poleward edge of precipitation in the premidnight sector are discussed.

Hoffman, R. A.↗

A failure model for sealed nickel-cadmium batteries

A model has been developed to describe failure in electrochemical batteries. The model is based on the concept of the existence and subsequent growth of flaws which ultimately lead to battery failure. This model provides, in a natural way, for the statistical variability of lifetime data. The model as applied to the Crane data indicates that when the effects of temperature and depth of discharge are taken into account, the observed variability in lifetime data is due almost entirely to statistical variability inherent in the battery itself.

Fedors, R. F.↗

Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs

This paper presents a methodology for debugging the performance of message-passing programs on both tightly coupled and loosely coupled distributed-memory machines. The AIMS (Automated Instrumentation and Monitoring System) toolkit, a suite of software tools for measurement and analysis of performance, is introduced and its application illustrated using several benchmark programs drawn from the field of computational fluid dynamics. AIMS includes (i) Xinstrument, a powerful source-code instrumentor, which supports both Fortran77 and C as well as a number of different message-passing libraries including Intel's NX Thinking Machines' CMMD, and PVM; (ii) Monitor, a library of timestamping and trace -collection routines that run on supercomputers (such as Intel's iPSC/860, Delta, and Paragon and Thinking Machines' CM5) as well as on networks of workstations (including Convex Cluster and SparcStations connected by a LAN); (iii) Visualization Kernel, a trace-animation facility that supports source-code clickback, simultaneous visualization of computation and communication patterns, as well as analysis of data movements; (iv) Statistics Kernel, an advanced profiling facility, that associates a variety of performance data with various syntactic components of a parallel program; (v) Index Kernel, a diagnostic tool that helps pinpoint performance bottlenecks through the use of abstract indices; (vi) Modeling Kernel, a facility for automated modeling of message-passing programs that supports both simulation -based and analytical approaches to performance prediction and scalability analysis; (vii) Intrusion Compensator, a utility for recovering true performance from observed performance by removing the overheads of monitoring and their effects on the communication pattern of the program; and (viii) Compatibility Tools, that convert AIMS-generated traces into formats used by other performance-visualization tools, such as ParaGraph, Pablo, and certain AVS/Explorer modules.

Yan, Jerry C.↗