Search NASASearch

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins

The CEOS Data Cube Portal: A User-Friendly, Open Source Software Solution for the Distribution, Exploration, Analysis, and Visualization of Analysis Ready Data

There is an urgent need to increase the capacity of developing countries to take part in the study and monitoring of their environments through remote sensing and space-based Earth observation technologies. The Open Data Cube (ODC) provides a mechanism for efficient storage and a powerful framework for processing and analyzing satellite data. While this is ideal for scientific research, the expansive feature space can also be daunting for end-users and decision-makers who simply require a solution which provides easy exploration, analysis, and visualization of Analysis Ready Data (ARD). Utilizing innovative web-design and a modular architecture, the Committee on Earth Observation Satellites (CEOS) has created a web-based user interface (UI) which harnesses the power of the ODC yet provides a simple and familiar user experience: the CEOS Data Cube (CDC). This paper presents an overview of the CDC architecture and the salient features of the UI. In order to provide adaptability, flexibility, scalability, and robustness, we leverage widely-adopted and well-supported technologies such as the Django web framework and the AWS Cloud platform. The fully-customizable source code of the UI is available at our public repository. Interested parties can download the source and build their own UIs. The UI empowers users by providing features that assist with streamlining data preparation, data processing, data visualization, and sub-setting ARD products in order to achieve a wide variety of Earth imaging objectives through an easy to use web interface.

User Interface

In Silico Human Mobility Data Science: Leveraging Massive Simulated Mobility Data (Vision Paper)

Human mobility data science using trajectories or check-ins of individuals has many applications. Recently, we have seen a plethora of research efforts that tackle these applications. However, research progress in this field is limited by a lack of large and representative datasets. The largest and most commonly used dataset of individual human trajectories captures fewer than 200 individuals, while datasets of individual human check-ins capture fewer than 100 check-ins per city per day. Thus, it is not clear if findings from the human mobility data science community would generalize to large populations. Since obtaining massive, representative, and individual-level human mobility data is hard to come by due to privacy considerations, the vision of this work is to embrace the use of data generated by large-scale socially realistic microsimulations. Informed by both real data and leveraging social and behavioral theories, massive spatially explicit microsimulations may allow us to simulate entire megacities at the person level. The simulated worlds, which do not capture any identifiable personal information, allow us to perform “in silico” experiments using the simulated world as a sandbox in which we have perfect information and perfect control without jeopardizing the privacy of any actual individual. In silico experiments have become commonplace in other scientific domains such as chemistry and biology, permitting experiments that foster the understanding of concepts without any harm to individuals. This work describes challenges and opportunities for leveraging massive and realistic simulated alternate worlds for in silico human mobility data science.

97 MATHEMATICS AND COMPUTING

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL),

Raw Lidar and Camera Data Synchronized with Precipitation and Present Weather Data

As part of the sensor characterization task of the SMART 2.0 project, this dataset includes raw data from three spinning lidars ([Ouster OS2-128](https://ouster.com/products/scanning-lidar/os2-sensor/), [Velodyne Puck (VLP-16)](https://velodynelidar.com/products/puck/), and [Velodyne Ultra Puck (VLP-32)](https://velodynelidar.com/products/ultra-puck/)), one camera ([Mako G-319](https://www.alliedvision.com/en/camera-selector/detail/mako/g-319/)), and one present weather sensor ([Vaisala FD-70](https://www.vaisala.com/en/products/weather-environmental-sensors/forward-scatter-fd70)). All data were synchronized, with the log start time indicated in the file name (HHMMSS). The data can be filtered by date, log time (HHMMSS), sensor, frame ID, and weather classification. These data were gathered statically at the Argonne Testbed for Multiscale Observational Science (ATMOS). Two target stop signs were placed in view of the sensors to contribute a target for comparing sensor data under different conditions. The weather data for each day are stored in netCDF “.nc” files. The lidar data contain the X, Y, Z, intensity, reflectivity, and ring from Ouster OS2-128 rev6, Velodyne VLP-16, and Velodyne VLP-32 lidars. ![raw lidar image](LiDAR_pointcloud_ATMOS.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the Geological Society of America Connects 2024 Annual Meeting in Anaheim, California, 22-25 September 2024.

Creason, Christopher

The data array, a tool to interface the user to a large data base

Aspects of the processing of spacecraft data is considered. Use of the data array in a large address space as an intermediate form in data processing for a large scientific data base is advocated. Techniques for efficient indexing in data arrays are reviewed and the data array method for mapping an arbitrary structure onto linear address space is shown. A compromise between the two forms is given. The impact of the data array on the user interface are considered along with implementation.

Foster, G. H.

Data availability and the role of the earth resources observation systems data center

With the launch of LANDSAT-1 in July 1972, and the follow-on launch of LANDSAT-2 in January of this year, routine availability of satellite imagery and electronic data of the earth's resources has become a reality. Federal data centers provide LANDSAT data to resource managers and the general public. These data centers have to date provided almost 500,000 frames of LANDSAT data at a cost of more than $2,000,000. Data from the LANDSAT satellite program, along with data and information from the Skylab manned program, are available over any location to anyone for the cost of reproduction.

Watkins, A. H.

The role of World Data Centers and the Lunar and Planetary Institute in the international exchange of lunar and planetary data

The success of many lunar and planetary investigations has resulted in the accumulation of a mass of data in a myriad of formats and medias. The application of these data to comparative planetology, origin of the solar system, and potential industrial applications of space has made it necessary for scientists from many disciplines to have access to the data. The collections of these data are so diverse that it is often difficult to select what is needed based on catalog information alone. The Lunar and Planetary Institute (LPI) proposes to bring the user and the data together. As a research support organization operated by the Universities Space Research Association, the LPI through its active Visiting Scientist Program, a balanced program of study workshops and topical conferences, and organized and supervised data collections has assisted scientists, educators, and students to review, study, and obtain the data necessary to the pursuit of their research.

Waranius, F. B.

Compatibility study of the Magsat data and aeromagnetic data in the eastern Piedmont of US

Data from a 2 day period recorded by Magsat were used to produce world magnetic maps of the scalar total field and three vector component total fields. Subtracting the reference field of Magsat 6/80, a scalar anomalous field and three vector component anomalous fields were also mapped. After removing 718 bad points from the original data, every fifth point was picked for contouring. While the main geomagnetic field of the Earth is surprisingly well mapped considering the short data period, the anomaly maps suffer from data sparseness. The entire Magsat file collected at altitudes of 500-700 m in nonmountainous terrain and at 900-1,000 m in mountainous terrain was averaged to reduce the total data to 6,500 measurements, yielding a 0.1 deg sampling interval along the flight path. A U.S. aeromagnetic anomaly surface map was produced and the field was upward continued to a 300 km altitude. Differences in anomaly structure between the POGO data and the map produced were attributed to insufficient removal of the reference field. Reprocessing of the data using the GSFC reference field (9/80-2) should remove the low harmonic field and improve the anomalous field structure.

Won, I. J.

Data engineering systems: Computerized modeling and data bank capabilities for engineering analysis

The Data Engineering System (DES) is a computer-based system that organizes technical data and provides automated mechanisms for storage, retrieval, and engineering analysis. The DES combines the benefits of a structured data base system with automated links to large-scale analysis codes. While the DES provides the user with many of the capabilities of a computer-aided design (CAD) system, the systems are actually quite different in several respects. A typical CAD system emphasizes interactive graphics capabilities and organizes data in a manner that optimizes these graphics. On the other hand, the DES is a computer-aided engineering system intended for the engineer who must operationally understand an existing or planned design or who desires to carry out additional technical analysis based on a particular design. The DES emphasizes data retrieval in a form that not only provides the engineer access to search and display the data but also links the data automatically with the computer analysis codes.

Kopp, H.

RIM as the data base management system for a material properties data base

Relational Information Management (RIM) was selected as the data base management system for a prototype engineering materials data base. The data base provides a central repository for engineering material properties data, which facilitates their control. Numerous RIM capabilities are exploited to satisfy prototype data base requirements. Numerical, text, tabular, and graphical data and references are being stored for five material types. Data retrieval will be accomplished both interactively and through a FORTRAN interface. The experience gained in creating and exercising the prototype will be used in specifying requirements for a production system.

Karr, P. H.

Standard format data units - Tools for automatic exchange of space mission data

A set of standard formatting rules for the data sets, and a standard computer-readable language with which to describe the data, are two tools which are used to create the Standard Format Data Unit (SFDU). The NASA/JPL proposal for creation and utilization of SFDUs is presented, and its relationship to recommendations from the Consultative Committee for Space Data Systems (CCSDS) is discussed. Several current and planned implementations of the SFDU concept among major space flight projects are identified. The purpose of creating the concept of an SFDU is to allow members of the science community to share national and global resource data independently of project or program. The feedback from SFDU implementation efforts is considered an essential part of the CCSDS activity. Even though the CCSDS specifically deals with space data systems, the SFDU concept can be applied to practically every data system on an open network. The SFDU is in the early phase of CCSDS standard definition work, and must go through several other phases before being formally recommended as an international standard.

Willett, J. B.

Data selection techniques in the interpretation of MAGSAT data over Australia

The MAGSAT data require critical selection in order to produce a self-consistent data set suitable for map construction and subsequent interpretation. Interactive data selection techniques are described which involve the use of a special-purpose profile-oriented data base and a colour graphics display. The careful application of these data selection techniques permits validation every data value and ensures that the best possible self-consistent data set is being used to construct the maps of the magnetic field measured at satellite altitudes over Australia.

Johnson, B. D.

Mathematical analysis study for radar data processing and enhancement. Part 1: Radar data analysis

A study is performed under NASA contract to evaluate data from an AN/FPS-16 radar installed for support of flight programs at Dryden Flight Research Facility of NASA Ames Research Center. The purpose of this study is to provide information necessary for improving post-flight data reduction and knowledge of accuracy of derived radar quantities. Tracking data from six flights are analyzed. Noise and bias errors in raw tracking data are determined for each of the flights. A discussion of an altiude bias error during all of the tracking missions is included. This bias error is defined by utilizing pressure altitude measurements made during survey flights. Four separate filtering methods, representative of the most widely used optimal estimation techniques for enhancement of radar tracking data, are analyzed for suitability in processing both real-time and post-mission data. Additional information regarding the radar and its measurements, including typical noise and bias errors in the range and angle measurements, is also presented. This is in two parts. This is part 1, an analysis of radar data.

James, R.