Search NASA⌕ Search

SEARCH · Search NASA

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Data Sharing in Radiation Biology: Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally "Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth independence and autonomy of mission operations. Here we present an overview of AI/ML architecture to support deep space mission goals, developed with leaders in the field. First, we focus on the fundamental biological research that supports our understanding of physiological responses to spaceflight, and we describe current efforts to support AI/ML research including data standardization and data engineering through maximally open and FAIR (findable, accessible, interoperable, reusable) databases and the generation of AI-ready datasets for reuse and analysis. We also discuss remote data management frameworks for research data as well as environmental and health data that are generated during deep space missions. We highlight several research projects that leverage data standardization and management for fundamental biological discovery to uncover the complex effects of space travel on living systems. Next, we provide an overview of cutting-edge AI/ML approaches that can be integrated to support remote monitoring and analysis during deep space missions, including generative models and large language models to learn the underlying biomedical patterns and predict outcomes or answer questions during off world medical scenarios. We also describe current AI/ML methods to support this research and monitoring through automated cloud-based labs which enable limited human intervention and closed-loop experimentation in remote settings. These labs could support mission autonomy by analyzing environmental data streams, and would be facilitated through in situ analytics capabilities to avoid sending large raw data files through low bandwidth communications. Finally, in the context of deep space missions with limited communications or access to medical advice from Earth, we describe a solution for integrated, real-time mission biomonitoring across hierarchical levels from continuous environmental monitoring, to wearables and point-of-care devices, to molecular and physiological monitoring. We introduce a precision space health system that will ensure that the future of space health is predictive, preventative, participatory and personalized.

artificial intelligence↗

Image Labeler: A Web Interface to Catalog Earth Science Events

Advances in machine learning (ML) have made it possible to automatically detect Earth science phenomena from satellite imagery. While useful, ML algorithms typically require an extensive dataset containing labeled images for training. Systematic labeling and management of such datasets is quite cumbersome. With this in mind, we present the Image Labeler. Image Labeler is a fast and scalable cloud-based tool that facilitates the rapid development of Earth science event databases, in order to aid automated ML-based image classification.

Case Study↗

Climatology of Global Precipitation Measurement Mission Precipitation Regimes and Implications for Global Estimates of Vertical Winds

The Global Precipitation Measurement (GPM) mission Validation Network (VN) framework leverages over 118 ground-based polarimetric Doppler radars to validate a large subset of precipitation measurements and retrievals from the GPM Dual-frequency Precipitation Radar (DPR). Recently, GPM DPR reflectivity profiles within the VN have been classified according to their convective regime using unsupervised machine learning techniques. The archetypal regimes are stratiform, convective, mixed stratiform-convective (e.g., transition regions), and “other” (e.g., peripheral regions of light precipitation). Subcategories within these four primary regimes vary according to the characteristic depth of included reflectivity profiles, resulting in 12 main GPM DPR precipitation profile categories. Polarimetry of ground-based Doppler radars in the VN offers additional insights into the types of precipitation, while pairs of radars positioned near each other enable retrieval of vertical winds via dual-Doppler analysis. Geometrically matched to the DPR reflectivity profiles in the GPM VN, these ground-based data and retrievals contribute more detailed characterization of the distinct kinematic and microphysical structures associated with each of the 12 DPR precipitation regimes. DPR reflectivity profiles linked with wind in the VN are restricted to GPM overpasses of proximal radar pairs that allow dual-Doppler analysis. Although a limited subset of DPR profiles in the VN are matched with vertical motion, agreement between the reflectivity structures paired with wind data and those of the greater DPR dataset in the VN suggest that estimates of vertical motion may be inferred in regions without ground-based measurements. We present a climatology of the 12 convective regimes identified within the DPR VN dataset as well as early efforts to estimate the kinematic and microphysical structures of precipitation profiles within the greater GPM DPR dataset by applying machine learning techniques. Precipitation data paired with global estimates of vertical winds from these efforts offer early insight to and support upcoming missions to retrieve convective mass flux, including the Investigation of Convective Updrafts (INCUS) in the Tropics and the global Atmosphere Observing System (AOS).

Precipitation↗

Toward a New Outlook on Primate Learning and Behavior: Complex Learning and Emergent Processes in Comparative Perspective

Primate research of the 20th century has established the validity of Darwin's postulation of psychological as well as biological continuity between humans and other primates, notably the great apes. Its data make clear that Descartes' view of animals as unfeeling 'beast-machines' is invalid and should be discarded. Traditional behavioristic frameworks, that emphasize the concepts of stimulus, response, and reinforcement and an 'empty-organism' psychology, are in need of major revisions. Revised frameworks should incorporate the fact that, in contrast to the lifeless databases of the 'hard' sciences, the database of psychology entails properties novel to life and its attendant phenomena. The contributions of research this century, achieved by field and laboratory researchers from around the world, have been substantial, indeed revolutionary. It is time to celebrate the progress of our field, to anticipate its significance, and to emphasize conservation of primates in their natural habitats.

Rumbaugh, Duane M.↗

Populating a Graph Database to Run a Usage-Based Discovery Tool

Most dataset discovery tools for Earth Observation data rely on descriptions and other metadata of the datasets, using keyword searches or attribute filtering to determine relevance. However, these descriptions often do not include the potential uses of the data. Thus, a user working on floods will rarely see few if any rainfall datasets show up in such a search. The Usage Based Discovery tool, on the other hand, offers usage instances to the user, either research articles or applications, along with the datasets that those usage instances used. This allows a user, particularly one new to the world of Earth Observation data, to investigate which datasets are used in similar cases. The information that powers Usage-Based Discovery is a graph database of relationships of usage to dataset and usage to topic, allowing the user to narrow their search for similar cases. In order to scale out to a graph database rich enough to provide a satisfactory user experience, we combine manual and automated processes to populate the graph. The initial content of the graph has been seeded primarily via human-aided data curation methods, using sites like Google Scholar. To scale up this effort, we’ve employed crowdsourcing. It is easy for anyone to contribute to our graph using their Open Researcher and Contributor Identifier for authorization. We’re now experimenting with Machine Learning and Natural Language Processing to help automate population of the graph, starting with the classification of research articles by topic. Finding adequate training data in the absence of a comprehensive and open research article API continues to be a significant challenge.

Vincent Inverso↗

Aerospace Engineering Systems

Continuous improvement of aerospace product development processes is a driving requirement across much of the aerospace community. As up to 90% of the cost of an aerospace product is committed during the first 10% of the development cycle, there is a strong emphasis on capturing, creating, and communicating better information (both requirements and performance) early in the product development process. The community has responded by pursuing the development of computer-based systems designed to enhance the decision-making capabilities of product development individuals and teams. Recently, the historical foci on sharing the geometrical representation and on configuration management are being augmented: Physics-based analysis tools for filling the design space database; Distributed computational resources to reduce response time and cost; Web-based technologies to relieve machine-dependence; and Artificial intelligence technologies to accelerate processes and reduce process variability. Activities such as the Advanced Design Technologies Testbed (ADTT) project at NASA Ames Research Center study the strengths and weaknesses of the technologies supporting each of these trends, as well as the overall impact of the combination of these trends on a product development event. Lessons learned and recommendations for future activities will be reported.

VanDalsem, William R.↗

Decision Manifold Approximation for Physics-Based Simulations

With the recent surge of success in big-data driven deep learning problems, many of these frameworks focus on the notion of architecture design and utilizing massive databases. However, in some scenarios massive sets of data may be difficult, and in some cases infeasible, to acquire. In this paper we discuss a trajectory-based framework that quickly learns the underlying decision manifold of binary simulation classifications while judiciously selecting exploratory target states to minimize the number of required simulations. Furthermore, we draw particular attention to the simulation prediction application idealized to the case where failures in simulations can be predicted and avoided, providing machine intelligence to novice analysts. We demonstrate this framework in various forms of simulations and discuss its efficacy.

Wong, Jay Ming↗

KARL: A Knowledge-Assisted Retrieval Language

Data classification and storage are tasks typically performed by application specialists. In contrast, information users are primarily non-computer specialists who use information in their decision-making and other activities. Interaction efficiency between such users and the computer is often reduced by machine requirements and resulting user reluctance to use the system. This thesis examines the problems associated with information retrieval for non-computer specialist users, and proposes a method for communicating in restricted English that uses knowledge of the entities involved, relationships between entities, and basic English language syntax and semantics to translate the user requests into formal queries. The proposed method includes an intelligent dictionary, syntax and semantic verifiers, and a formal query generator. In addition, the proposed system has a learning capability that can improve portability and performance. With the increasing demand for efficient human-machine communication, the significance of this thesis becomes apparent. As human resources become more valuable, software systems that will assist in improving the human-machine interface will be needed and research addressing new solutions will be of utmost importance. This thesis presents an initial design and implementation as a foundation for further research and development into the emerging field of natural language database query systems.

Dominick, Wayne D.↗

Aerospace Engineering Systems and the Advanced Design Technologies Testbed Experience

Continuous improvement of aerospace product development processes is a driving requirement across much of the aerospace community. As up to 90% of the cost of an aerospace product is committed during the first 10% of the development cycle, there is a strong emphasis on capturing, creating, and communicating better information (both requirements and performance) early in the product development process. The community has responded by pursuing the development of computer-based systems designed to enhance the decision-making capabilities of product development individuals and teams. Recently, the historical foci on sharing the geometrical representation and on configuration management are being augmented: 1) Physics-based analysis tools for filling the design space database; 2) Distributed computational resources to reduce response time and cost; 3) Web-based technologies to relieve machine-dependence; and 4) Artificial intelligence technologies to accelerate processes and reduce process variability. The Advanced Design Technologies Testbed (ADTT) activity at NASA Ames Research Center was initiated to study the strengths and weaknesses of the technologies supporting each of these trends, as well as the overall impact of the combination of these trends on a product development event. Lessons learned and recommendations for future activities are reported.

VanDalsem, William R.↗

The Large Footprint of Small-scale Artisanal Gold Mining in Ghana

Gold mining has played a significant role in Ghana's economy for centuries. Regulation of this industry has varied over time and while industrial mining is prevalent in the country, the expansion of artisanal mining, or Galamsey has escalated in recent years. Many of these artisanal mines are not only harmful to human health due to the use of Mercury (Hg) in the amalgamation process, but also leave a significant footprint on terrestrial ecosystems, degrading and destroying forested ecosystems in the region. In this study, the Landsat image archive available through Google Earth Engine was used to quantify the total footprint of vegetation loss due to artisanal goldmines in Ghana from 2005 to 2019 and understand how conversion of forested regions to mining has changed over a decadal period from 2007 to 2017. A combination of machine learning and change detection algorithms were used to calculate different land cover conversions and the timing of conversion annually. Within the study area of southwestern Ghana, our results indicate that approximately 47,000 ha (⨦2218 ha) of vegetation were converted to mining at an average rate of ~2600 ha yr−1. The results indicate that a high percentage(~50%) of this mining occurred between 2014 and 2017. Around 700 ha of this mining occurred within protected areas as mapped by the World Database of Protected Areas. In addition to deforestation, increased artisanal mining activity in recent years has the potential to affect human health, access to drinking water resources and food security. This work expands upon limited research into the spatial footprint of Galamseyin Ghana, complements mapping efforts by local geographers, and will support efforts by the government of Ghana to monitor deforestation caused by artisanal mining.

Abigail Barenblitt↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

A Survey Protocol to Assess Meaningfulness and Usefulness of Automated Topic Finding in the NASA Aviation Safety Reporting System

Context: The NASA Aviation Safety Reporting System (ASRS) is a voluntary confidential aviation safety reporting system. The ASRS receives reports from pilots, air traffic controllers, flight attendants and other involved in aviation operations. The reports are de-identified and coded by ASRS expert safety analysts and a short descriptive synopsis is written to describe the safety issue. The de-identified reports are then disseminated to the aviation community in a number of ways including entry into an online database, Safety Alert Bulletins and For Your Information Notices, and the CALLBACK newsletter. Key to these publications are the timely processing (de-identification, coding and summarization) of new reports, which is currently done by ASRS expert safety analysts. Thus, we believe topic modelling could decrease effort in ASRS, if topics are comprehensible. Aim: We propose a methodology to evaluate whether automated topic finding using topic modelling provides meaningful and useful topics. Method: We extend the total error survey methodology to evaluate user topic comprehension of machine learning outputs. To accomplish this we performed a literature review to identify existing methods and define a construct for topic comprehension, utilizing existing ASRS synopsis writing practices to more precisely define meaningfulness and usefulness. Results: A survey protocol was created that addresses the limitations of other survey protocols found in the literature review, which we found lacking in rationale and clear protocol definition. Conclusion: The surveying of user understanding in machine learning outputs presents challenges due to the explosion of parameters to control for and the lack of systematic approach presented in the literature. More reproducible work and survey protocols are needed in the literature and our work is one step towards that direction.

topic finding↗

The Spaceport Command and Control System Security Assessor Project

This Summer, I worked as a National Aeronautics and Space Administration (NASA) Internships and Fellowships (NIF) intern under my mentor, Jill Giles within the Software Engineering Branch. Within this project, I worked alongside the Cyber Security branch to identify a list of Commercial Off the Shelf (COTS) software to analyze, research, and gain insight about potential vulnerabilities within the software that could become a threat of attack. After identifying the list of COTS software, my team and I used Microsoft Excel to create a worksheet to easily organize and design a questionnaire about the software. Security reports weregiven to us to identify the software used on the machines in the firing rooms. With these reports, we created a script that would populate the database with the software information to identify potential security weaknesses of COTS software.The goal of the project was to produce a final report, summarizing the most vulnerable launch control system servers and configurations and document vulnerabilities, residual risk, likelihood, and consequence. This project is important for the Cyber Security and Information Technology branches because it will identify security weaknesses and help to mitigate risk. From the Spaceport Command and Control System Security Assessor Project, I learned how to properly identify weaknesses and vulnerabilities within software and how to mitigate the risks within the software. This project also taught me how to create databases using scripts and input files.

Destani Satora Van Arsdalen↗

Requirement Discovery Using Embedded Knowledge Graph with ChatGPT

The field of Advanced Air Mobility (AAM) is witnessing a transformation with innovations such as electric aircraft and increasingly automated airspace operations. Within AAM, the Urban Air Mobility (UAM) concept focuses on providing air-taxi services in densely populated urban areas. This research introduces the utilization of Large Language Models (LLMs), such as OpenAI's GPT-4, to enhance the UAM Requirement discovery process. This study explores two distinct approaches to leverage LLMs in the context of UAM Requirement discovery. The first approach evaluates the LLM's ability to provide responses without relying on additional outside systems, such as a relational or graph database. Instead, a vector store provides relevant information to the LLM based on the user’s question, a process known as Retrieval Augmented Generation (RAG). The second approach integrates the LLM with a graph database. The LLM acts as an intermediary between the user and the graph database, translating user questions into cypher queries for the database and database responses into human-readable answers for the user. Our team implemented and tested both solutions to analyze requirements within a UAM dataset. This paper will talk about our approaches, implementations, and findings related to both approaches.

systems engineering↗

Requirement Discovery Using Embedded Knowledge Graph With ChatGPT

The field of Advanced Air Mobility (AAM) is witnessing a transformation with innovations such as electric aircraft and increasingly automated airspace operations. Within AAM, the Urban Air Mobility (UAM) con-cept focuses on providing air-taxi services in densely populated urban areas. This research introduces the utilization of Large Language Models (LLMs), such as OpenAI's GPT-4, to enhance the UAM Requirement discovery process. This study explores two distinct approaches to leverage LLMs in the context of UAM Requirement discovery. The first approach evaluates the LLM's ability to provide responses without relying on additional outside systems, such as a relational or graph database. Instead, a vector store provides relevant information to the LLM based on the user’s question, a process known as Retrieval Augmented Generation (RAG). The second approach integrates the LLM with a graph database. The LLM acts as an intermediary between the user and the graph database, translating user questions into cypher queries for the database and database responses into human-readable answers for the user. Our team implemented and tested both solutions to analyze require-ments within a UAM dataset. This paper will talk about our approaches, implementations, and findings related to both approaches.

systems engineering↗

Requirement Discovery Using Embedded Knowledge Graph With ChatGPT - Poster

The field of Advanced Air Mobility (AAM) is witnessing a transformation with innovations such as electric aircraft and increasingly automated airspace operations. Within AAM, the Urban Air Mobility (UAM) con-cept focuses on providing air-taxi services in densely populated urban areas. This research introduces the utilization of Large Language Models (LLMs), such as OpenAI's GPT-4, to enhance the UAM Requirement discovery process. This study explores two distinct approaches to leverage LLMs in the context of UAM Requirement discovery. The first approach evaluates the LLM's ability to provide responses without relying on additional outside systems, such as a relational or graph database. Instead, a vector store provides relevant information to the LLM based on the user’s question, a process known as Retrieval Augmented Generation (RAG). The second approach integrates the LLM with a graph database. The LLM acts as an intermediary between the user and the graph database, translating user questions into cypher queries for the database and database responses into human-readable answers for the user. Our team implemented and tested both solutions to analyze require-ments within a UAM dataset. This paper will talk about our approaches, implementations, and findings related to both approaches.

systems engineering↗

Sub-Continental-Scale Carbon Stocks of Individual Trees in African Drylands

The distribution of dryland trees and their density, cover, size, mass and carbon content are not well known at sub-continental to continental scales. This information is important for ecological protection, carbon accounting, climate mitigation and restoration efforts of dryland ecosystems. We assessed more than 9.9 billion trees derived from more than 300,000 satellite images, covering semi-arid sub-Saharan Africa north of the Equator. We attributed wood, foliage and root carbon to every tree in the 0–1,000 mm year −1 rainfall zone by coupling field data, machine learning, satellite data and high-performance computing. Average carbon stocks of individual trees ranged from 0.54 Mg C ha −1 and 63 kg C tree −1 in the arid zone to 3.7 Mg C ha −1 and 98 kg tree −1 in the sub-humid zone. Overall, we estimated the total carbon for our study area to be 0.84 (±19.8%) Pg C. Comparisons with 14 previous TRENDY numerical simulation studies23 for our area found that the density and carbon stocks of scattered trees have been underestimated by three models and overestimated by 11 models, respectively. This benchmarking can help understand the carbon cycle and address concerns about land degradation. We make available a linked database of wood mass, foliage mass, root mass and carbon stock of each tree for scientists, policymakers, dryland-restoration practitioners and farmers, who can use it to estimate farmland tree carbon stocks from tablets or laptops.

Compton Tucker↗