Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Hyperspectral segmentation of plants in fabricated ecosystems

Hyperspectral imaging provides a powerful tool for analyzing above-ground plant characteristics in fabricated ecosystems, offering rich spectral information across diverse wavelengths. This study presents an efficient workflow for hyperspectral data segmentation and subsequent data analytics, minimizing the need for user annotation through the use of ensembles of sparse mixed scale convolution neural networks. The segmentation process leverages the diversity of ensembles to achieve high accuracy with minimal labeled data, reducing labor-intensive annotation efforts. To further enhance robustness, we incorporate image alignment techniques to address spatial variability in the dataset. Downstream analysis focuses on using the segmented data for processing spectral data, enabling monitoring of plant health. This approach provides a scalable solution for spectral segmentation, and facilitates actionable insights into plant conditions in complex, controlled environments. Our results demonstrate the utility of combining advanced machine learning techniques with hyperspectral analytics for high-throughput plant monitoring.

Zwart, Petrus H.↗

A Performance Model of In-Situ Techniques

The computational capacity of High-Performance Computing (HPC) systems increases continuously with the rapid development of central processing units (CPUs) and graphic processing units (GPUs), while the in-/output (IO) subsystem develops relatively slowly and storage capacity is also limited. Data-intensive applications, which are designed to leverage the high computational capacity of HPC resources, typically generate a considerable amount of data for post-processing visualizations and data analytics. The limited IO speed and storage space could lead to constraints in the actual performance of these applications and, therefore, scientific discovery. In-situ techniques, where data is visualized/analysed while still in memory rather than through disk, can contribute to alleviating these problems as they can reduce or even fully avoid data writing/reading through the IO subsystem to/from storage. However, the overall efficiency of insitu techniques crucially depends on the characteristics of both the in-situ tasks and the applications, and the resource distribution among them. Therefore, choosing the right in-situ approach (synchronous, asynchronous, or hybrid) and resource allocation is essential to minimize overhead and maximize the benefits of concurrent execution. In this paper, we present a performance model of in-situ techniques to find the most beneficial in-situ approach and the preferred resource configuration. We verify the high accuracy of our approach with over 6800 measurements and provide use cases with different applications.

Ju, Yi [Max Planck Computing and Data Facility, Ga↗

JSC OCFO Cloud Integrated Budget Analytical Toolbox

The JSC OCFO analyst's job function relies heavily on analytics to provide every directorate customer with quality service in managing their respective budgets. This technology, configured on the Salesforce cloud platform, enhances the JSC OCFO customer user interface, provides state of the art data analytics and statistics, reduces IT costs, increases productivity, eliminates duplicate entry and errors, and utilizes machine learning and artificial intelligence to process day to day transactional tasks automatically without analyst manipulation. The unique configuration of the COTS technology allows for highly advanced analysis and reporting of government budget data. This provides the ability for complex budgets to be fully understood in detail by non-budget organizations and eliminates time consuming analysis on the OCFO analyst's behalf.

Robertson, Arden↗

Modal characteristics of a stiffened composite cylinder with open and closed end conditions

The structural and acoustic modal frequencies and the loss factors of a composite fuselage model under free-free end and support-end conditions are evaluated. The fuselage model is a composite filament wound, cylindrical structure composed of carbon fibers embedded in an epoxy resin and has a ply orientation of +,-45, +,-32, 0, -,+32, and -,+45. A computer aided test system was utilized to obtain resonance frequencies, modal damping, and mode shape coefficients. Sound pressure and phase survey data are utilized to estimate acoustic modal freuencies and reverberation time was measured to calculate the acoustic loss factors. The derived structural and acoustic modal parameters are compared to analytical data from a computer prediction model. The data reveal that the acoustic modal frequencies and mode shape correlate well with predicted data, and there is fair agreement between structural modal frequencies and mode shape for the two data sets.

Grosveld, F. W.↗

Practical Aspects of the Frequency Domain Approach for Aircraft System Identification

Practical aspects of the frequency-domain approach for aircraft system identification are explained and demonstrated. Topics related to experiment design, flight data analysis, and dynamic modeling are included. For demonstration purposes, simulated time series data and simulated flight data from an F-16 nonlinear simulation with realistic noise are used. This approach enables detailed evaluations of the techniques and results, because the true characteristics of the data and aircraft dynamics are known for the simulated data. Analytical techniques and practical considerations are examined for the finite Fourier transform, nonparametric frequency response estimation, parametric modeling in the frequency domain, experiment design for frequency-domain modeling, data analysis and modeling in the frequency domain, and real-time calculations. Flight data from a subscale jet transport aircraft are used to demonstrate some of the techniques and technical issues.

Morelli, Eugene A.↗

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics↗

Results of geodetic processing and analysis of Skylab altimetry data

A geodetic analysis of Skylab S-193 altimeter preliminary data from mission SL/2 and EREP pass 9 is considered. The overall objective of the investigation was a demonstration of the feasibility of a use of altimeter data for the determination of the geoid in ocean areas. The geoid is the equipotential surface that would coincide with an 'undisturbed' mean sea level of the earth's gravity field. Analytical data handling formulations are discussed.

Fubara, D. M. J.↗

The BioMole Facility: Advancement of In Situ Microbiome Analysis for the International Space Station

Characterization of the International Space Station (ISS) microbiome has been enabled by sample return and Earth-based analysis. As human exploration pushes beyond low-Earth orbit, microbial-related crew health, planetary protection, and space research requires in situ capabilities. Steps toward reducing Earth-dependence for complex sample analysis began in 2016 with the amplification of DNA within the miniPCR thermal cycler and DNA sequencing with the MinION sequencer onboard the ISS; for both, samples were prepared on Earth. In 2017, these platforms synergistically enabled the in-situ identification of unknown bacteria collected and cultured from ISS surfaces, thereby shifting the paradigm that microbial cultures had to be returned to Earth. The following year, a culture-independent, swab-to-sequencer method further advanced spaceflight microbiology, demonstrating that culturing could be excluded and provided enhanced insight into the bacterial profile of ISS surfaces. Based on the success of these payloads in confirming the ability to meet crew health identification requirements and the benefits accompanying a culture-independent method, the BioMole Facility was established by the medical operations Crew Health Care Systems team. BioMole is the set of hardware, consumables, and procedures required to support sample preparation and nanopore sequencing onboard the ISS. BioMole goals include expanding sample sources, comparing data to previous methods, demonstrating onboard data analytics, and validating new hardware. To date, comparative surface analysis, molecular- and culture-based, has been completed. Additionally, the demonstration of a sample-to-answer process was achieved when BioMole data was processed onboard using the IBM Open Data and AI Edge software platform installed on the ISS-residing Spaceborne Computer-2. The taxonomic profiles generated from the edge analysis were as expected and paralleled that of the downlinked processed data. Future BioMole efforts involve microbial profiling of the ISS water system, ISS validation of the MinION Mk1C, and an expansion to a research facility available to investigators.

Sarah L. Castro-Wallace↗

Accelerated Simulation of Air Pollution Using NVIDIA RAPIDS

Atmospheric chemistry models are a central tool to study and forecast the impact of air pollution on the environment, vegetation, and human health. However, the numerical simulation of chemical kinetics is computationally expensive due to the stiffness of the system of ordinary differential equations that describes atmospheric chemistry. Here we present an alternative approach to the computation of atmospheric chemistry based on machine learning. Our training data set is produced using the NASA Goddard Earth Observing System (GEOS) model with GEOS-Chem chemistry, run on the NASA Center for Climate Simulation (NCCS) Discover supercomputing cluster on 384 Intel Xeon Haswell cores. This model spends more than 50% of total run time on solving atmospheric chemistry. The data set contains as input features the air pollution concentrations before solving the differential equations, together with some key physical parameters such as temperature and sun intensity. As target variables we define the air pollution concentrations after solving the differential equations. Using Dask-cuDF and Dask-XGBoost on the NVIDIA RAPIDS platform on 8 Tesla V100 GPUs, we generate from this training set gradient boosted decision tree models that can reproduce the simulation of chemical kinetics. We do this on the NCCS Advanced Data Analytics Platform (ADAPT) science cloud environment. Our application takes full advantage of recent advances in Dask-XGBoost, such as multi-node and multi-GPU scaling for distributed training with large data sets. The increase in training data size enabled by this is critical to capture the full range of chemical environments encountered across the globe and all annual seasons.The boosted tree models offer good predictability and show many of the features of the full chemistry reference simulation. Further improvements can be achieved through mass balance considerations and by accounting for error correlations. We incorporate the boosted tree models into the GEOS reference model using XGBoost's C API. This enables a seamless integration of the GPU trained models into GEOS-Chem, which is written in Fortran and optimized for use in a massively parallel CPU environment. We show the benefits of this approach and discuss the potential speedup of this machine learning accelerated atmospheric chemistry model.

Keller, Christoph A.↗

Introduction to NASA Goddard Workshop on Artificial Intelligence

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few.This workshop will be investigating how AI technologies can be adapted or developed to address the following challenges: Discover events of interest and correlations in large amounts of science data; improve the outcomes of science modeling and data assimilation using improved data processing, integration, and analysis. Design advisors for mission planning and operations, including anomaly detection and spacecraft health monitoring. Develop tools for engineering support, including advanced manufacturing, orbit determination, new component design and system engineering. Customize intelligent user interfaces, including visual analytics and natural language processing.

Le Moigne, Jacqueline↗

Development of National New Construction Weighting Factors for the Commercial Building Prototype Analyses (2008-2022)

The U.S. Department of Energy (DOE) tasked Pacific Northwest National Laboratory (PNNL) with updating commercial building construction weights for the purpose of estimating national and state-by-state energy savings impacts of changes made to various commercial energy codes and standards. A similar activity was last completed by PNNL in 2020 using disaggregate construction volume data acquired from the Dodge Data & Analytics database (formerly McGraw Hill) for the years 2003-2018 (Lei et al, 2020). As time passes, changes in economic and social demand reshape construction volume trends. For the current update, PNNL reviewed the same data source with the latest construction data for the years 2008-2022. For commercial building analyses, PNNL typically uses a suite of 16 prototype buildings simulated in the 19 ASHRAE climate zones with 16 of them present in the United States. The 2008-2022 commercial building weighting factors were derived using the same approach employed to develop the 2003-2018 set (Lei et al, 2020). Applying the construction volume data from the database to the prototypes and climate zones resulted in the following new construction area-based weighting factors. Table ES.1 shows the weighting factors including all building categories found in the database, and Table ES.2 shows the weighting factors normalized to include only buildings represented by the 16 prototypes. Section 3.0 also includes national- and state-level weighting factors by area and building count.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The Development of the Ducted Fan Noise Propagation and Radiation Code CDUCT-LaRC

The development of the ducted fan noise propagation and radiation code CDUCT-LaRC at NASA Langley Research Center is described. This code calculates the propagation and radiation of given acoustic modes ahead of the fan face or aft of the exhaust guide vanes in the inlet or exhaust ducts, respectively. This paper gives a description of the modules comprising CDUCT-LaRC. The grid generation module provides automatic creation of numerical grids for complex (non-axisymmetric) geometries that include single or multiple pylons. Files for performing automatic inviscid mean flow calculations are also generated within this module. The duct propagation is based on the parabolic approximation theory of R. P. Dougherty. This theory allows the handling of complex internal geometries and the ability to study the effect of non-uniform (i.e. circumferentially and axially segmented) liners. Finally, the duct radiation module is based on the Ffowcs Williams-Hawkings (FW-H) equation with a penetrable data surface. Refraction of sound through the shear layer between the external flow and bypass duct flow is included. Results for benchmark annular ducts, as well as other geometries with pylons, are presented and compared with available analytical data.

Nark, Douglas M.↗

A digital twin platform for building performance monitoring and optimization: Performance simulation and case studies

Advancements in sensor technology, data analytics, affordable compute, and communication infrastructure have paved the way for Digital Twin technology in optimizing building operations and controls. This study presents the development of an open and interoperable web-based Digital Twin platform for integrating diverse data streams and facilitating effective user interactions. The platform utilizes modern technologies for the web framework and time-series data management, ensuring scalability and responsiveness. The backend supports seamless integration of diverse data sources and emulators, incorporating data from building sensors and meters, external weather Application Programming Interfaces, and advanced EnergyPlus simulation models of the building and its energy systems including the Distributed Energy Resources that are formulated in Functional Mockup Units. A simulation case study was conducted with FlexLab, a test facility on Lawrence Berkeley National Laboratory campus. The case study includes normal operations, Distributed Energy Resource integration, and power outage scenarios, to illustrate the Digital Twin’s ability to provide critical insights into energy performance and thermal resilience. The results demonstrated the platform’s potential as a decision-support tool for optimizing building energy performance and enhancing resilience against extreme weather events. Future work will focus on deploying the Digital Twin platform to a real building for field validation, extending its capabilities to cover more scenarios such as bidirectional Electric Vehicle interactions, and enhancing user engagement.

EnergyPlus↗

The Vertebrate Breed Ontology: Toward Effective Breed Data Standardization

Abstract Background Limited universally-adopted data standards in veterinary medicine hinder data interoperability and therefore integration and comparison; this ultimately impedes the application of existing information-based tools to support advancement in diagnostics, treatments, and precision medicine. Hypothesis/Objectives A single, coherent, logic-based standard for documenting breed names in health, production, and research-related records will improve data use capabilities in veterinary and comparative medicine. Animals No live animals were used. Methods The Vertebrate Breed Ontology (VBO) was created from breed names and related information compiled from the Food and Agriculture Organization of the United Nations, breed registries, communities, and experts, using manual and computational approaches. Each breed is represented by a VBO term that includes breed information and provenance as metadata. VBO terms are classified using description logic to allow computational applications and Artificial Intelligence–readiness. Results VBO is an open, community-driven ontology representing over 19 500 livestock and companion animal breed concepts covering 49 species. Breeds are classified based on community and expert conventions (e.g., cattle breed) and supported by relations to the breed's genus and species indicated by National Center for Biotechnology Information (NCBI) Taxonomy terms. Relationships between VBO terms (e.g., relating breeds to their foundation stock) provide additional context to support advanced data analytics. VBO term metadata includes synonyms, breed identifiers/codes, and attributed cross-references to other databases. Conclusion and Clinical Importance The adoption of VBO as a standard for breed names in databases and veterinary electronic health records enhances veterinary data interoperability and computability, supporting precision medicine.

Veterinary Sciences↗

AI Driven Optimization of Public Transit

This project explores the application of AI-driven methods to optimize public transit operations for the Chattanooga Area Regional Transportation Authority (CARTA). By leveraging data analytics, machine learning, and predictive modeling, the initiative seeks to enhance system efficiency, improve rider experience, and support sustainability goals. This research, supported by the National Science Foundation and the U.S. Department of Energy, integrates real-time transit data with advanced computational tools to inform decision-making, optimize routes, and balance operational demands. The work exemplifies a forward-looking model for mid-sized cities aiming to modernize mobility systems through intelligent technology integration.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Throughput of the Composite Infrared Spectrometer (CIRS) Mid Infrared (MIR) Channel for the Cassini Mission to Saturn

The Composite Infrared Spectrometer (CIRS) of the Cassini mission to Saturn has two interferometers covering the far infrared and mid infrared wavelength region. The instrument was aligned at ambient temperature, but operates at 170 Kelvin and has challenging interferometric alignment tolerances. Cryogenic alignment tests of the instrument indicate that it should suffer minimal degradation due to the cooldown from ambient to operational temperature. System level tests performed by the calibration team indicated a lower than expected signal level on the MIR channel, while providing ambiguous optical throughput data. Therefore it became imperative to develop a metric that could be used to determine the instrument performance at both the instrument and system levels, at ambient and cryogenic temperature. Modulation efficiency and throughput measurements were performed and new analytical models developed to evaluate the status of the instrument. Empirical and analytical data were eventually reconciled and deviations from the design values explained.

Hagopian, John G.↗

Earth System Digital Twins (ESDT) Technology for NASA Earth Science

For NASA's Advanced Information Systems Technology (AIST) Program, an Earth System Digital Twin (ESDT) is defined as an interactive and integrated multidomain, multiscale, digital replica of the state and temporal evolution of Earth systems. It dynamically integrates: relevant Earth system models and simulations; other relevant models (e.g., related to the world's infrastructure); continuous and timely (including near real time and direct readout) observations (e.g., space, air, ground, over/underwater, Internet of Things (IoT), socioeconomic); long-time records; as well as analytics and artificial intelligence tools. Effective ESDTs enable users to run hypothetical scenarios to improve the understanding, prediction of and mitigation/response to Earth system processes, natural phenomena and human activities as well as their many interactions. An ESDT is a type of integrated information system that, for example, enables continuous assessment of impact from naturally occurring and/or human activities on physical and natural environments. AIST ESDT strategic goals are to: 1. Develop information system frameworks to provide continuous and accurate representations of systems as they change over time; 2. Mirror various Earth Science systems and utilize the combination of Data Analytics, Artificial Intelligence, Digital Thread, and state-of-the-art models to help predict the Earth’s response to various phenomena; 3. Provide the tools to conduct "what if" investigations that can result in actionable predictions. The AIST ESDT thrust is developing capabilities toward the development of future digital twins of the Earth or of subcomponents of the Earth. This will enable the development of an overarching framework that will integrate New Observing Strategies (NOS) to enable new observation measurements, i.e., multi-source, coordinated, dynamic and responsive to needs and requests defined by Analytic Collaborative Frameworks (ACF) that enable agile science investigations fusing and analyzing very large amounts of diverse data. NOS and ACF capabilities along with open access to various science, infrastructure and human data, interconnected modeling, data assimilation, simulations, surrogate modeling, high-performance computing and advanced visualization, will define a powerful framework that could be utilized for local, regional or global and/or thematic digital twins. This presentation will describe a general overview of the AIST ESDT vision including prior work done in the areas of NOS and ACF as well as current and upcoming ESDT projects.

Jacqueline Le Moigne↗

Analyzing a 35-Year Hourly Data Record: Why So Difficult?

At the Goddard Distributed Active Archive Center, we have recently added a 35-Year record of output data from the North American Land Assimilation System (NLDAS) to the Giovanni web-based analysis and visualization tool. Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) offers a variety of data summarization and visualization to users that operate at the data center, obviating the need for users to download and read the data themselves for exploratory data analysis. However, the NLDAS data has proven surprisingly resistant to application of the summarization algorithms. Algorithms that were perfectly happy analyzing 15 years of daily satellite data encountered limitations both at the algorithm and system level for 35 years of hourly data. Failures arose, sometimes unexpectedly, from command line overflows, memory overflows, internal buffer overflows, and time-outs, among others. These serve as an early warning sign for the problems likely to be encountered by the general user community as they try to scale up to Big Data analytics. Indeed, it is likely that more users will seek to perform remote web-based analysis precisely to avoid the issues, or the need to reprogram around them. We will discuss approaches to mitigating the limitations and the implications for data systems serving the user communities that try to scale up their current techniques to analyze Big Data.

computational performance↗