Search NASA⌕ Search

SEARCH · Search NASA

Results for “Column subset selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Algorithm For Solution Of Subset-Regression Problems

Reliable and flexible algorithm for solution of subset-regression problem performs QR decomposition with new column-pivoting strategy, enables selection of subset directly from originally defined regression parameters. This feature, in combination with number of extensions, makes algorithm very flexible for use in analysis of subset-regression problems in which parameters have physical meanings. Also extended to enable joint processing of columns contaminated by noise with those free of noise, without using scaling techniques.

Verhaegen, Michel↗

On the reliable and flexible solution of practical subset regression problems

A new algorithm for solving subset regression problems is described. The algorithm performs a QR decomposition with a new column-pivoting strategy, which permits subset selection directly from the originally defined regression parameters. This, in combination with a number of extensions of the new technique, makes the method a very flexible tool for analyzing subset regression problems in which the parameters have a physical meaning.

Verhaegen, M. H.↗

On the reliable and flexible solution of practical subset regression problems

A new algorithm for solving subset regression problems is described. The algorithm performs a QR decomposition with a new column-pivoting strategy, which permits subset selection directly from the originally defined regression parameters. This, in combination with a number of extensions of the new technique, makes the method a very flexible tool for analyzing subset regression problems in which the parameters have a physical meaning.

Verhaegen, M.↗

Vector-Ordering Filter Procedure for Data Reduction

The vector-ordering filter (VOF) technique involves a procedure for sampling a large population of data vectors to select a subset of data vectors that fully characterize the state space of the large population. The VOF technique enables a large reduction of the volume of data that must be handled in the automated monitoring system and method discussed in the two immediately preceding articles. In so doing, the VOF technique enables the development of data-driven mathematical models of a monitored asset from sets of data that would otherwise exceed the memory capacities of conventional engineering computers. Data-driven mathematical models have been shown to offer high fidelity for purposes of control and monitoring of assets. In practice, a collection of asset-operating observations is acquired with the intention that the collection contain observations characteristic of the full dynamic range of operation of the asset. Often, such a collection contains an extremely large number of observations, many of which are redundant. The VOF technique fills the need for a means to extract, from the original collection of observational data, a reduced data matrix that excludes redundant data while maintaining the full statistical character and dynamic range of the original data. The reduced data matrix can then be used as the input data for development of a mathematical model of the monitored asset, or as training data for a neural-network substitute for an explicit mathematical model of the asset. Alternatively, the reduced data matrix can, itself, be used directly as a mathematical model of the monitored asset, as is commonly done in multivariate state-estimation techniques. The original data are collected from the asset over a range of operating states and are put in matrix form. Each column vector in the original data matrix represents the signal values acquired at a particular operational state of the asset. Thus, the number of columns of the original data matrix equals the number of observed states and the number of rows in this matrix equals the number of signals acquired at each observation. In the VOF technique, one extracts the reduced data matrix from the original data matrix through the selection of a representative subset of the column (state) vectors.

Bickford, Randall L.↗

Selecting Data from a Star Catalog

MCDUMP is a computer program that selects data from the SKYMAP SKY2000 Master Star Catalog a database about 150 MB in size, stored on a computer hard drive. The database describes about 300,000 stars, each by means of a 500-byte entry. MCDUMP reads all 300,000 entries, then generates an output file that comprises a subset of entries selected according to one or more criteria entered by the user. Examples of criteria that could be entered include: location in a selected portion of the sky; constancy or a specified degree of variability of brightness; absence of nearby, bright companion stars; a particular surface temperature; and brightness sufficient to enable detection by a specified astronomical instrument. The output of MCDUMP can be in the form of either a single 520-column file or multiple files that contain fewer columns to facilitate printing. MCDUMP has been configured and tested for use under the HP-UX 10.20 operating system (a Hewlett-Packard version of the UNIX operating system). It should also be possible to adapt MCDUMP to other versions of UNIX.

Tracewell, David A.↗

Remote Sensing Data from CLARET: A Prototype Cart Data Set

A data set containing radiation, meteorological, and cloud sensor observations is documented. It was prepared for use by the Department of Energy's Atmospheric Radiation Measurement (ARM) program and other interested scientists. These data are a precursor of the types of data that ARM Cloud And Radiation Testbed (CART) sites will provide. The data are from the Cloud Lidar And Radar Exploratory Test (CLARET) conducted by the Wave Propagation Laboratory during autumn 1989 in the Denver-Boulder area of Colorado primarily for the purpose of developing new cloud-sensing techniques on cirrus. After becoming aware of this experiment, ARM scientists requested archival of subsets or the data to assist in the developing ARM program. Five CLARET cases were selected: two with cirrus, one with stratus, one with mixed-phase clouds, and one with clear skies. The cases range from 2 to 9.5 h in length. A pyranometer, pyrgeometer, pyrheliometer, and an infrared radiometer constituted the ensemble of instruments that provided surface radiation data. A lidar, radar, and ceilometer observed the cloud geometrical structure, and visual reports and all-sky camera observations were assimilated to provide cloud cover data. Radiosondes, wind profiler, RASS (profiling virtual temperature), microwave radiometers (observing column integrated liquid water and water vapor), and standard surface measurements provided meteorological data. Satellite data from the stratus case and one cirrus case were analyzed for statistics on cloud cover and top height. The main body of the selected data are available on diskette from the Wave Propagation Laboratory or Los Alamos National Laboratory. In addition to documenting the data set, this report describes CLARET and gives a bibliography of publications associated with the project. Some preliminary results of CLARET' research are also summarized. Simultaneous CO 2 lidar and radar backscatter measurements were shown to provide estimates of the effective radius of ice particles. Simultaneous radar and infrared radiometer data appear useful for estimating column-integrated numbers and average sizes of ice cloud particles. Ice water content obtained with this method compared favorably with values from another empirical technique using radar data alone. Depolarization of the CO 2 lidar signal from ice clouds was surprisingly small, suggesting that calculation of backscatter from nonspherical particles for this lidar is a tractable problem. Examples are also cited of CO 2 lidar measurements of the effective radius of water cloud drop size distributions and of inference of the size of pristine ice crystals that assume a particular orientation in the air. These parameters are all important to radiative transfer through clouds.

Clouds (Meteorology)↗

An Examination of the Recent Stability of Ozonesonde Global Network Data

The recent Assessment of Standard Operating Procedures for OzoneSondes (ASOPOS 2.0; WMO/GAW Report #268) addressed questions of homogeneity and long-term stability in global electrochemical concentration cell (ECC) ozone sounding network time series. Among its recommendations was adoption of a standard for evaluating data quality in ozonesonde time series. Total column ozone (TCO) derived from the sondes compared to TCO from Aura’s Ozone Monitoring Instrument (OMI) is a primary quality indicator. Comparisons of sonde ozone with Aura’s Microwave Limb Sounder (MLS) are used to assess the stability of stratospheric ozone. This paper provides a comprehensive examination of global ozonesonde network data stability and accuracy since 2004 in light of the sudden post-2013 TCO “dropoff” of ~3-4% that was reported previously at select stations (Stauffer et al., 2020). Comparisons with Aura OMI TCO averaged across the network of 60 stations are stable within about ±2% over the past 18 years. Sonde TCO has similar stability compared to three other TCO satellite instruments, and the stratospheric ozone measurements average to within ±5% of MLS from 50 to 10 hPa. Thus, sonde data are reliable for trends, but with a caveat applied for a subset of dropoff stations in the tropics and subtropics. The dropoff is associated with only one of two major ECC instrument types. A detailed examination of ECC serial numbers pinpoints the timing of the dropoff. However, we find that overall, ozonesonde data are stable and accurate compared to independent measurements over the past two decades.

Ryan M. Stauffer↗

Engine Icing Data - An Analytics Approach

Engine icing researchers at the NASA Glenn Research Center use the Escort data acquisition system in the Propulsion Systems Laboratory (PSL) to generate and collect a tremendous amount of data every day. Currently these researchers spend countless hours processing and formatting their data, selecting important variables, and plotting relationships between variables, all by hand, generally analyzing data in a spreadsheet-style program (such as Microsoft Excel). Though spreadsheet-style analysis is familiar and intuitive to many, processing data in spreadsheets is often unreproducible and small mistakes are easily overlooked. Spreadsheet-style analysis is also time inefficient. The same formatting, processing, and plotting procedure has to be repeated for every dataset, which leads to researchers performing the same tedious data munging process over and over instead of making discoveries within their data. This paper documents a data analysis tool written in Python hosted in a Jupyter notebook that vastly simplifies the analysis process. From the file path of any folder containing time series datasets, this tool batch loads every dataset in the folder, processes the datasets in parallel, and ingests them into a widget where users can search for and interactively plot subsets of columns in a number of ways with a click of a button, easily and intuitively comparing their data and discovering interesting dynamics. Furthermore, comparing variables across data sets and integrating video data (while extremely difficult with spreadsheet-style programs) is quite simplified in this tool. This tool has also gathered interest outside the engine icing branch, and will be used by researchers across NASA Glenn Research Center. This project exemplifies the enormous benefit of automating data processing, analysis, and visualization, and will help researchers move from raw data to insight in a much smaller time frame.

Engine Icing↗

New Features of the NEQAIR Radiation Code

The longest-lived code for predicting shock layer radiation, NEQAIR, is now in its 5th decade of service. Substantial changes to the code have been made over the previous decade, the most recent report of which was at the 5th Workshop on Radiation in High Temperature Gases in 2014, for the version referred to as NEQAIR14. This paper will review some of the improvements made to the NEQAIR code since then, which is now at v15.2. Some of these features are discussed briefly below. NEQAIR15 and subsequent versions have enabled parallel evaluation of multiple lines of sight. This is accomplished by utilizing the HDF5 file format and placing multiple lines into a single file, LOS.h5, which is used for both input and output. This approach enables straightforward parallel execution both over the number of lines of sight and the number of points per line. For large problems, runtime reduces linearly with the number of nodes deployed since each line is processed independently by a subset of MPI ranks. Three applications of the multi-line solver are discussed. The first has to do with performing loosely coupled radiation-flowfield solutions. In this case the computed absorption and emission coefficients are used to evaluate the total energy absorbed or emitted at each point, allowing evaluation of the volumetric source term in the flowfield. The second computation is for obtaining heat flux from nonuniform flows, which require integration over spherical co-ordinates. These are of particular interest for evaluating radiation on the vehicle backshell. This 3D option improves the angular integration scheme and allows adaptive line selection that together reduce the number of lines required by about an order of magnitude. The final application is for remote observation, which is essentially the 3D integration problem over a small solid angle. For all three of these computations, data can be stored in the HDF5 file which allows a NEQAIR run to be restarted when it times out, or to add atmospheric absorption or instrument scan functions. An additional level of parallelism is enabled in NEQAIR15.2 using GPU routines. The GPU parallelism has realized up to 8x speed-up when running on a single core but diminishes as CPU parallelism is increased. For running multi-line simulations, it may be easier to reserve a large number of CPU nodes than to obtain the number of GPU nodes required for similar performance. A GUI, known as NEQTPY, allows for reading and creating input files, running NEQAIR, and displaying results. A significant feature of NEQTPY is the ability to perform spectral fits to data. The fits can operate on a single line spectrum (radiance vs. wavelength) or a 3D input file with multiple columns of data. Other new features include improved constants, additional species, more detailed non-Boltzmann modelling, advanced user controls, the ability to read and calculate spectra from HITRAN datafiles, photodissociation and photoionization cross-sections. A “fast” automatic grid option may reduce the size and time of spectral calculations while still maintaining good accuracy for total heat flux.

Brett A Cruden↗