Search NASA⌕ Search

SEARCH · Search NASA

Results for “Automatic Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

Surface melting over the Greenland ice sheet from enhanced resolution passive microwave brightness temperatures (1979–2019)

Surface melting is a major component of the Greenland ice sheet (GrIS) surface mass balance, affecting sea level rise through direct runoff and the modulation on ice dynamics and hydrological processes, supraglacially, englacially and subglacially. Passive microwave (PMW) brightness temperature observations are of paramount importance in studying the spatial and temporal evolution of surface melting in view of their long temporal coverage (1979–to date) and high temporal resolution (daily). However, a major limitation of PMW datasets has been the relatively coarse spatial resolution, being historically of the order of tens of kilometres. Here, we use a newly released passive microwave dataset (37 GHz, horizontal polarization) made available through the NASA MeASUREs program to study the spatiotemporal evolution of surface melting over the GrIS at an enhanced spatial resolution of 3.125 Km. We assess the outputs of different detection algorithms through data collected by Automatic Weather Stations (AWS) and the outputs of the MAR regional climate model. We found that surface melting is well captured using a dynamic algorithm based on the outputs of MEMLS model, capable to detect sporadic and persistent melting. Our results indicate that, during the reference period 1979–2019 (1988–2019), surface melting over the GrIS increased in terms of both duration, up to ~4.5 (2.9) days per decade, and extension, up to 6.9 % (3.6 %) of the GrIS surface extent per decade, according to the MEMLS algorithm. Furthermore, the melting season has started up to ~4 (2.5) days earlier and ended ~7 (3.9) days later per decade. We also explored the information content of the enhanced resolution dataset with respect to the one at 25 km and MAR outputs through a semi-variogram approach. We found that the enhanced product is more sensitive to local scale processes, hence confirming the potential interest of this new enhanced product for studying surface melting over Greenland at a higher spatial resolution than the historical products and monitor its impact on sea level rise. This offers the opportunity to improve our understanding of the processes driving melting, to validate modelled melt extent at high resolution and potentially to assimilate this data in climate models.

Surface melting↗

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB↗

Performance evaluation of a simulated data-flow computer with low-resolution actors

Basic problems related to the exploitation of parallelism in a program include sequencing of the instructions and communication of the data. It is pointed out that the data-flow approach offers an elegant solution to the sequencing problem, since all data dependencies are automatically handled and only instructions with ready input sets are activated. It is shown that a change in the level of subcomputations (actors) affects communications costs. The concept of variable resolution is discussed, and the testbed environment is examined. Attention is given to the architecture of the processing elements, the communication network, and the simulators. A description of the analytical model is also provided. Simulation and results are discussed, taking into account test programs and allocation, the variation of the number of processing elements, the variation of the resolution in directed acyclic graphs, performance in processing loops, and array handling.

Gaudiot, J. L.↗

Data-Base Software For Tracking Technological Developments

Technology Tracking System (TechTracS) computer program developed for use in storing and retrieving information on technology and related patent information developed under auspices of NASA Headquarters and NASA's field centers. Contents of data base include multiple scanned still images and quick-time movies as well as text. TechTracS includes word-processing, report-editing, chart-and-graph-editing, and search-editing subprograms. Extensive keyword searching capabilities enable rapid location of technologies, innovators, and companies. System performs routine functions automatically and serves multiple users.

Aliberti, James A.↗

Automated Discovery of Flight Track Anomalies

As new technologies are developed to handle the complexities of the Next Generation Air Transportation System (NextGen), it is increasingly important to address both current and future safety concerns along with the operational, environmental, and efficiency issues within the National Airspace System (NAS). In recent years, the Federal Aviation Administration’s (FAA) safety offices have been researching ways to utilize the many safety databases maintained by the FAA, such as those involving flight recorders, radar tracks, weather, and many other high- volume sensors, in order to monitor this unique and complex system. Although a number of current technologies do monitor the frequency of known safety risks in the NAS, very few methods currently exist that are capable of analyzing large data repositories with the purpose of discovering new and previously unmonitored safety risks. While monitoring the frequency of known events in the NAS enables mitigation of already identified problems, a more proactive approach of finding unidentified issues still needs to be addressed. This is especially important in the proactive identification of new, emergent safety issues that may result from the planned introduction of advanced NextGen air traffic management technologies and procedures. Development of an automated tool that continuously evaluates the NAS to discover both events exhibiting flight characteristics indicative of safety-related concerns as well as operational anomalies will heighten the awareness of such situations in the aviation community and serve to increase the overall safety of the NAS. This paper discusses the extension of previous anomaly detection work to identify operationally significant flights within the highly complex airspace encompassing the New York area of operations, focusing on the major airports of Newark International (EWR), LaGuardia International (LGA), and John F. Kennedy International (JFK). In addition, flight traffic in the vicinity of Denver International (DEN) airport/airspace is also investigated to evaluate the impact on operations due to variances in seasonal weather and airport elevation. From our previous research, subject matter experts determined that some of the identified anomalies were significant, but could not reach conclusive findings without additional supportive data. To advance this research further, causal examination using domain experts is continued along with the integration of air traffic control (ATC) voice data to shed much needed insight into resolving which flight characteristic(s) may be impacting an aircraft's unusual profile. Once a flight characteristic is identified, it could be included in a list of potential safety precursors. This paper also describes a process that has been developed and implemented to automatically identify and produce daily reports on flights of interest from the previous day.

Matthews, Bryan↗

The northernmost hyperspectral FLoX sensor dataset for monitoring of high-Arctic tundra vegetation phenology and Sun-Induced Fluorescence (SIF)

A hyperspectral field sensor (FloX) was installed in Adventdalen (Svalbard, Norway) in 2019 as part of the Svalbard Integrated Arctic Earth Observing System (SIOS) for monitoring vegetation phenology and Sun-Induced Chlorophyll Fluorescence (SIF) of high-Arctic tundra. This northernmost hyperspectral sensor is located within the footprint of a tower for long-term eddy covariance flux measurements and is an integral part of an automatic environmental monitoring system on Svalbard (AsMovEn), which is also a part of SIOS. One of the measurements that this hyperspectral instrument can capture is SIF, which serves as a proxy of gross primary production (GPP) and carbon flux rates. This paper presents an overview of the data collection and processing, and the 4-year (2019–2021) datasets in processed format are available at: https://thredds.met.no/thredds/catalog/arcticdata/infranor/NINA-FLOX/raw/catalog.html associated with https://doi.org/10.21343/ZDM7-JD72 under a CC-BY-4.0 license. Results obtained from the first three years in operation showed interannual variation in SIF and other spectral vegetation indices including MERIS Terrestrial Chlorophyll Index (MTCI), EVI and NDVI. Synergistic uses of the measurements from this northernmost hyperspectral FLoX sensor, in conjunction with other monitoring systems, will advance our understanding of how tundra vegetation responds to changing climate and the resulting implications on carbon and energy balance.

hyperspectral field sensor↗

SymbolFit: Automatic Parametric Modeling with Symbolic Regression

We introduce SymbolFit (API: https://github.com/hftsoi/symbolfit), a framework that automates parametric modeling by using symbolic regression to perform a machine-search for functions that fit the data while simultaneously providing uncertainty estimates in a single run. Traditionally, constructing a parametric model to accurately describe binned data has been a manual and iterative process, requiring an adequate functional form to be determined before the fit can be performed. The main challenge arises when the appropriate functional forms cannot be derived from first principles, especially when there is no underlying true closed-form function for the distribution. In this work, we develop a framework that automates and streamlines the process by utilizing symbolic regression, a machine learning technique that explores a vast space of candidate functions without requiring a predefined functional form because the functional form itself is treated as a trainable parameter, making the process far more efficient and effortless than traditional regression methods. We demonstrate the framework in high-energy physics experiments at the CERN Large Hadron Collider (LHC) using five real proton-proton collision datasets from new physics searches, including background modeling in resonance searches for high-mass dijet, trijet, paired-dijet, diphoton, and dimuon events. We show that our framework can flexibly and efficiently generate a wide range of candidate functions that fit a nontrivial distribution well using a simple fit configuration that varies only by random seed, and that the same fit configuration, which defines a vast function space, can also be applied to distributions of different shapes, whereas achieving a comparable result with traditional methods would have required extensive manual effort.

Tsoi, Ho Fung [Univ. of Pennsylvania, Philadelphia↗

Fault management for data systems

Issues related to automating the process of fault management (fault diagnosis and response) for data management systems are considered. Substantial benefits are to be gained by successful automation of this process, particularly for large, complex systems. The use of graph-based models to develop a computer assisted fault management system is advocated. The general problem is described and the motivation behind choosing graph-based models over other approaches for developing fault diagnosis computer programs is outlined. Some existing work in the area of graph-based fault diagnosis is reviewed, and a new fault management method which was developed from existing methods is offered. Our method is applied to an automatic telescope system intended as a prototype for future lunar telescope programs. Finally, an application of our method to general data management systems is described.

Boyd, Mark A.↗

Digital Analytics, Causal Knowledge Acquisition and Reasoning for Technical Language Processing

Complex engineering systems such as nuclear power plants (NPPs) generate and collect large amounts of equipment reliability (ER) data elements that contain information on the status of components, assets, and systems. Some of this information is textual in form and can be found in documents such as incident reports (IRs) and work orders (WOs). Analyses of textual data in current NPPs-using natural language processing (NLP) methods-have been expanded over the last decade, and it is only recently that the true potential of such analyses has emerged. So far, applications of NLP methods have mostly been limited to classification and prediction, the goal being to identify the nature of the textual element (e.g., safety or non-safety related). Here, we target a more complex problem: automatically extracting knowledge from a textual element in order to assist system engineers in conducting system health assessments. Knowledge extraction is a very broad concept, and its definition may vary depending on the application context. Our methods are a blend of both rule-based and machine learning (ML) algorithms. For our purposes, knowledge extraction means identifying the systems or assets mentioned in a given textual element, as well as the type of event described (e.g., component failure or maintenance activity). In addition, we want to capture details such as measured quantities and the temporal/cause-effect relations between events. In this tool, we also demonstrate how textual data elements are preprocessed in order to handle typos, acronyms, and abbreviations. One main feature of these methods is that they are not based solely on data, but are in fact model-based. In other words, they also rely on MBSE models that are designed to capture-from a functional point of view-the architecture of the systems/assets under consideration. The main purpose of such models is to digitally emulate system engineers' knowledge of system and asset architecture and to identify dependencies among systems, assets, and components. Provided these models, analyses of textual and numeric ER data can be performed by first identifying the OPM model elements to which the ER data elements are referring. The relationships between ER data elements are then identified by checking for any temporal or logical dependencies.

Mandelli, Diego [Idaho National Laboratory (INL), ↗

Reconstruction framework advancements to support streaming for the ePIC detector at the EIC

The ePIC collaboration adopted the JANA2 framework to manage its reconstruction algorithms. This framework has since evolved substantially in response to ePIC’s needs. There have been three main design drivers: integrating cleanly with the Podio-based data models and other layers of the key4hep stack, enabling external configuration of existing components, and supporting timeframe splitting for streaming readout. The result is a unified component model featuring a new declarative interface for specifying inputs, outputs, parameters, services, and resources. This interface enables the user to instantiate, configure, and wire components via an external file. One critical new addition to the component model is a hierarchical decomposition of data boundaries into levels such as Run, Timeframe, PhysicsEvent, and Subevent. Two new component abstractions, Folder and Unfolder, are introduced in order to traverse this hierarchy, e.g. by splitting or merging. The pre-existing components can now operate at different event levels, and JANA2 will automatically construct the corresponding parallel processing topology. This means that a user may write an algorithm once, and configure it at runtime to operate on timeframes or on physics events. Overall, these changes mean that the user requires less knowledge about the framework internals, obtains greater flexibility with configuration, and gains the ability to reuse the existing abstractions in new streaming contexts.

Brei, Nathan [Thomas Jefferson National Accelerato↗

ARMing the Edge: Demonstration of Edge Computing Field Campaign Report

Edge computing enables “next-to-instrument” control and intelligent data volume reduction and the potential for autonomous, adaptive measurement strategies such as for automated control of scan strategies for Doppler lidar (DL). Instruments with narrow bandwidth connections (e.g., ship and remote sites) can do scene determination and save phenomenon-appropriate data. For example, Doppler spectrum can be saved when clouds are detected by the instrument or automatic moment detection can take place in camera images and only preserve spectrum when non-monomodal spectra are detected. Automated control at the edge involves changing the sampling (temporal or scanning strategy) of an instrument to suit the phenomena both present and being studied (Jackson et al. 2020). Both data processing and instrument control introduces the possibility of a software-defined instrument.

54 ENVIRONMENTAL SCIENCES↗

The Viking seismometer

The Viking seismometer, operating on the Martian surface in the region of Utopia Planitia, includes three inertial velocity transducers, amplifiers, filters, automatic event detectors, data compactors, and temporary data storage capability within a unit weighing 2.2 kg. The seismometer functions at a peak magnification of 218,000 with a ground resolution of 2 nanosec and 3 Hz. Since the telemetry capacities associated with the system are limited, the seismometer instrument package employs three data processing modes: a high sampling rate mode; a heavily filtered background monitoring mode; and a frequently-used data compacting mode.

Miller, W. F.↗

Passive microwave remote sensing for sea ice research

Techniques for gathering data by remote sensors on satellites utilized for sea ice research are summarized. Measurement of brightness temperatures by a passive microwave imager converted to maps of total sea ice concentration and to the areal fractions covered by first year and multiyear ice are described. Several ancillary observations, especially by means of automatic data buoys and submarines equipped with upward looking sonars, are needed to improve the validation and interpretation of satellite data. The design and performance characteristics of the Navy's Special Sensor Microwave Imager, expected to be in orbit in late 1985, are described. It is recommended that data from that instrument be processed to a form suitable for research applications and archived in a readily accessible form. The sea ice data products required for research purposes are described and recommendations for their archival and distribution to the scientific community are presented.

Source record↗

The Cooperative Huntsville Meteorological Experiment (COHMEX)

The Satellite Precipitation and Cloud Experiment, the Microburst and Severe Thunderstorm, and the FAA Lincoln Laboratories Operational Weather Studies of the COHMEX are described. The precipitation and cloud experiment focuses on the prestorm period in order to observe the physical processes leading to the formation of small convective systems. Aircraft, remote sensing and rewinsonde data are utilized to determine various storm/environment characteristics. Doppler velocity and reflectivity of microburst clouds are studied to evaluate the three-dimensional structure of microbursts from thunderstorms. The weather studies are designed to develop and test automatic algorithms for wind shear detection using Doppler weather radars. The application of satellite systems to data collection for these experiments is discussed.

Dodge, J.↗

Ames vision group research overview

A major goal of the reseach group is to develop mathematical and computational models of early human vision. These models are valuable in the prediction of human performance, in the design of visual coding schemes and displays, and in robotic vision. To date researchers have models of retinal sampling, spatial processing in visual cortex, contrast sensitivity, and motion processing. Based on their models of early human vision, researchers developed several schemes for efficient coding and compression of monochrome and color images. These are pyramid schemes that decompose the image into features that vary in location, size, orientation, and phase. To determine the perceptual fidelity of these codes, researchers developed novel human testing methods that have received considerable attention in the research community. Researchers constructed models of human visual motion processing based on physiological and psychophysical data, and have tested these models through simulation and human experiments. They also explored the application of these biological algorithms to applications in automated guidance of rotorcraft and autonomous landing of spacecraft. Researchers developed networks for inhomogeneous image sampling, for pyramid coding of images, for automatic geometrical correction of disordered samples, and for removal of motion artifacts from unstable cameras.

Watson, Andrew B.↗

Rate 3/4 coded 16-QAM for uplink applications

First phase development of an advanced modulation technology which synergistically combines coding and modulation to achieve 2 bits per second per Hertz bandwidth efficiency in satellite demodulators is nearing completion. A proof-of-concept model is being developed to demonstrate technology feasibility, establish practical bandwidth efficiency limitations, and provide a data base for the design and development of engineering model satellite demodulators. The basic considerations leading to the choice of 4 x 4 quadrature amplitude modulation (16-QAM) and its associated coding format are discussed, along with the basic implementation of the carrier and clock recovery, automatic gain control, and decoding process. Preliminary performance results are presented. Spectra for the modulated signal shows the effects of the square root Nyquist filters in the modulation. Bit error rate (BER) results for the encoder/decoder subsystem show near ideal results, although power consumption is high and baseband BER performance of the Nyquist filter set is poor. Recommendations regarding the present system to improve BER performance and acquisition speed are given.

Mrozek, Eric M.↗

The Overgrid Interface for Computational Simulations on Overset Grids

Computational simulations using overset grids typically involve multiple steps and a variety of software modules. A graphical interface called OVERGRID has been specially designed for such purposes. Data required and created by the different steps include geometry, grids, domain connectivity information and flow solver input parameters. The interface provides a unified environment for the visualization, processing, generation and diagnosis of such data. General modules are available for the manipulation of structured grids and unstructured surface triangulations. Modules more specific for the overset approach include surface curve generators, hyperbolic and algebraic surface grid generators, a hyperbolic volume grid generator, Cartesian box grid generators, and domain connectivity: pre-processing tools. An interface provides automatic selection and viewing of flow solver boundary conditions, and various other flow solver inputs. For problems involving multiple components in relative motion, a module is available to build the component/grid relationships and to prescribe and animate the dynamics of the different components.

Chan, William M.↗