Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

TESS Spots a Hot Jupiter with an Inner Transiting Neptune

Hot Jupiters are rarely accompanied by other planets within a factor of a few in orbital distance. Previously, only two such systems have been found. Here, we report the discovery of a third system using data from the Transiting Exoplanet Survey Satellite (TESS). The host star, TOI-1130, is an eleventh magnitude K-dwarf in Gaia G-band. It has two transiting planets: a Neptune-sized planet (3.65±0.10 Rꚛ) with a 4.1 days period, and a hot Jupiter (-1.50(+0.22,-0.27) R(J)) with an 8.4 days period. Precise radial-velocity observations show that the mass of the hot Jupiter is -0.974(+0.044, 0.043) M(J). For the inner Neptune, the data provide only an upper limit on the mass of 0.17M(J)(3σ). Nevertheless, we are confident that the inner planet is real, based on follow-up ground-based photometry and adaptive-optics imaging that rule out other plausible sources of the TESS transit signal. The unusual planetary architecture of and the brightness of the host star make TOI-1130 a good test case for planet formation theories, and an attractive target for future spectroscopic observations.

Exoplanet astronomy↗

Transform Your Search Capability Using Insight Engines

Enable rapid discovery of NASA’s open science data, software and documentation. Support NASA’s open science goals and infrastructure. Promote interdisciplinary science. Prototype emerging technologies and search techniques including Large Language Models.

Kaylin Bugbee↗

HPC Campaign Management: Remote data access with user-defined error bound using ADIOS and ZFP

Remote access to large-scale scientific datasets, like those generated by combustion simulations or other high-performance computing (HPC) applications, presents a significant challenge. Downloading entire datasets is often impractical due to their size and the bandwidth limitations of typical networks. To address this challenge, we propose a novel approach that enables efficient remote access to large datasets distributed across multiple facilities. Our method enables technologies to download only the data values of a select variable, in a select region of interest, to a user-defined accuracy. For this purpose, we extended the ADIOS IO library to provide read functions with user-defined accuracy, a remote data server that understands multidimensional selections of specific variables, steps and accuracy from an ADIOS dataset, and which uses lossy compression on the remote site to reduce the data to be transferred back to the client. In addition, our extension of the ADIOS library collects metadata from multiple datasets in small files called Campaign Archives, which can be shared among project participants on any HPC, cloud or laptop, and which can easily facilitate the discovery of content and pointers to the data location as well as remote access to the data by local tools as if data was local. This feature called Campaign Management, enables a group of scientists to manage related datasets stored in multiple files, across multiple facilities as if it was in a single file/database. We demonstrate the effectiveness of our approach using a 1.5 TB dataset from the S3D combustion simulation on Frontier at the Oak Ridge Leadership Facility. Even a single variable from this dataset, at 64 GB, is too large to be processed on a standard laptop. We show two different reading patterns for 2D plots and 3D visualization, with careful settings that a scientist studying combustion data would do and show that running the same Python scripts on Frontier directly takes comparable time than running them on the local laptop with remote access to the data on Frontier.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X↗

Hypersonic Navier-Stokes Comparisons to Orbiter Flight Data

During the STS-119 flight of Space Shuttle Discovery, two sets of surface temperature measurements were made. Under the HYTHIRM program3 quantitative thermal images of the windward side of the Orbiter with a were taken. In addition, the Boundary Layer Transition Flight Experiment 4 made thermocouple measurements at discrete locations on the Orbiter wind side. Most of these measurements were made downstream of a surface protuberance designed to trip the boundary layer to turbulent flow. In this paper, we use the US3D computational fluid dynamics code to simulate the Orbiter flow field at conditions corresponding to the STS-119 re-entry. We employ a standard two-temperature, five-species finite-rate model for high-temperature air, and the surface catalysis model of Stewart.1 This work is similar to the analysis of Wood et al . 2 except that we use a different approach for modeling turbulent flow. We use the one-equation Spalart-Allmaras turbulence model8 with compressibility corrections 9 and an approach for tripping the boundary layer at discrete locations. In general, the comparison between the simulations and flight data is remarkably good

Candler, Graham V.↗

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING↗

The northeast materials database for magnetic materials

The discovery of magnetic materials with high operating temperature ranges and optimized performance is essential for advanced applications. Current data-driven approaches are limited by the lack of accurate, comprehensive, and feature-rich databases. This study aims to address this challenge by using Large Language Models (LLMs) to create a comprehensive, experiment-based, magnetic materials database named the Northeast Materials Database (NEMAD), which consists of 67,573 magnetic materials entries (www.nemad.org). The database incorporates chemical composition, magnetic phase transition temperatures, structural details, and magnetic properties. Enabled by NEMAD, we trained machine learning models to classify materials and predict transition temperatures. Our classification model achieved an accuracy of 90% in categorizing materials as ferromagnetic (FM), antiferromagnetic (AFM), and non-magnetic (NM). The regression models predict Curie (Néel) temperature with a coefficient of determination (R 2 ) of 0.87 (0.83) and a mean absolute error (MAE) of 56K (38K). These models identified 25 (13) FM (AFM) candidates with a predicted Curie (Néel) temperature above 500K (100K) from the Materials Project. This work shows the feasibility of combining LLMs for automated data extraction and machine learning models to accelerate the discovery of magnetic materials.

Ferromagnetism↗

Planetary image conversion task

The Planetary Image Conversion Task group processed 12,500 magnetic tapes containing raw imaging data from JPL planetary missions and produced an image data base in consistent format on 1200 fully packed 6250-bpi tapes. The output tapes will remain at JPL. A copy of the entire tape set was delivered to US Geological Survey, Flagstaff, Ariz. A secondary task converted computer datalogs, which had been stored in project specific MARK IV File Management System data types and structures, to flat-file, text format that is processable on any modern computer system. The conversion processing took place at JPL's Image Processing Laboratory on an IBM 370-158 with existing software modified slightly to meet the needs of the conversion task. More than 99% of the original digital image data was successfully recovered by the conversion task. However, processing data tapes recorded before 1975 was destructive. This discovery is of critical importance to facilities responsible for maintaining digital archives since normal periodic random sampling techniques would be unlikely to detect this phenomenon, and entire data sets could be wiped out in the act of generating seemingly positive sampling results. Reccomended follow-on activities are also included.

Martin, M. D.↗

Data Grid Management Systems

The "Grid" is an emerging infrastructure for coordinating access across autonomous organizations to distributed, heterogeneous computation and data resources. Data grids are being built around the world as the next generation data handling systems for sharing, publishing, and preserving data residing on storage systems located in multiple administrative domains. A data grid provides logical namespaces for users, digital entities and storage resources to create persistent identifiers for controlling access, enabling discovery, and managing wide area latencies. This paper introduces data grids and describes data grid use cases. The relevance of data grids to digital libraries and persistent archives is demonstrated, and research issues in data grids and grid dataflow management systems are discussed.

Moore, Reagan W.↗

Geologic and mineral and water resources investigations in western Colorado, using Skylab EREP data

The author has identified the following significant results. Discovery of three major north-trending, throughgoing faults in the Front Range, previously mapped only as isolated segments, demonstrates the utility of space photography and may lead to reinterpretation of the Front Range tectonic style. Faulting and alteration appear to be the most useful indicators of mineralization in central Colorado. These phenomena appear on Skylab photography as tonal lineaments and color anomalies. Twenty-three lineaments have been mapped in the San Juan Mountains, the longest of which is 156 km long. Twelve lineaments intersect or are tangent to calderas. Intrusive domes are aligned along lineaments, but calderas appear to occur at the intersections of major lineaments. Lineaments can be recognized on some EREP passes but not on other passes over the same area. The difference is attributed to solar elevation effects. Bedding attitudes can be photogeologically estimated down to surprisingly low dips, on the order of + or - 1-2 deg, and attitudes can be subdivided easily into quantitative groups. The primary application of Skylab photography to geologic mapping in montane areas is clearly limited to regional mapping at scales smaller than 1:24,000.

Lee, K.↗

The processing and transmission of EEG data

Interest in sleep research was stimulated by the discovery of a number of physiological changes that occur during sleep and by the observed effects of sleep on physical and mental performance and status. The use of the relatively new methods of EEG measurement, transmission, and automatic scoring makes sleep analysis and categorization feasible. Sleep research involving the use of the EEG as a fundamental input has the potential of answering many unanswered questions involving physical and mental behavior, drug effects, circadian rhythm, and anesthesia.

Schulze, A. E.↗

Joint Inversion and Forward Modeling of Gravity and Magnetic Data in the Ismenius Region of Mars

The unexpected discovery of remanent crustal magnetism on Mars was one of the most intriguing results from the Mars Global Surveyor mission. The origin of the pattern of magnetization remains elusive. Correlations with gravity and geology have been examined to better understand the nature of the magnetic anomalies. In the area of the Martian dichotomy between 50 and 90 degrees E (here referred to as the Ismenius Area), we find that both the Bouguer and the isostatic gravity anomalies appear to correlate with the magnetic anomalies and a buried fault, and allow for a better constraint on the magnetized crust].

Milbury, C. A.↗

STS-114: Discovery Mission Status Briefing

Paul Hill, STS-114 Lead Shuttle Flight Director, talks about the imagery that was captured during the first twenty four hours in orbit. He expresses that new data was captured during ascent of the Space Shuttle Discovery through imagery and in-orbit handheld photographs of the Orbital Maneuvering system (OMS) pods were taken. Laser surveys were also taken of both the wing leading edges and nose caps using the Orbiter Boom and Sensor System (OBSS) instrument package. He presents raw footage from the Laser Dynamic Range Imager (LDRI) at the end of the OBSS showing RCC panels and tiles. He also answers questions from the news media about the possible damage to the Orbiter and debris during lift-off.

Source record↗

Basin & Range Investigation for Developing Geothermal Energy

Hidden geothermal systems represent a potentially prolific energy resource that could support critical U.S. public and government energy priorities. Basin and Range Investigations for Developing Geothermal Energy (BRIDGE) addressed some the challenges associated with hidden system exploration by prioritizing cost-effective exploration early on through strategic workflow and informed decision-making that mitigates early risk and shifts resources to later exploration stages (e.g., drilling). Sandia National Laboratories partnered with U.S. Navy Geothermal Office, Geologic Geothermal Group, and independent consultants, with additional collaboration with U.S. Geological Survey and private industry. The primary tool of the BRIDGE project was to deploy a regional-scale airborne electromagnetic method to investigate the shallow resistivity structure in areas with high prospectivity. This was followed up at several prospects by a multidisciplinary exploration approach, including additional geologic, geophysical and geochemical studies. A central tenet to the BRIDGE methodology is that zones of low resistivity frequently occur over geothermal systems in the Basin and Range, and when paired with other data constraints, imaging these zones can enable discovery of these systems. In addition to exploring greenfield areas (i.e., Grover Point), the BRIDGE project also flew HTEM resistivity surveys over known geothermal systems including those with established power plants (Don A. Campbell and Salt Wells) and prospects that are known to the literature but remain undeveloped, at least in part, due to a lack of understanding on the location of their producible reservoirs. BRIDGE produced a comprehensive set of data from prospects identified in the Nevada Play Fairway Analysis along with conceptual models for top ranking prospects, wherein all of the observations are used to inform an interpreted model of the system. These models present a range of possible system parameters such as temperature and size, and they are further informed by system analogues in the Basin and Range province and elsewhere. The results of this work leave space for further exploration that may now occur at prospects ‘down the list’ rather than distribution exploration resources evenly across all prospects.

15 GEOTHERMAL ENERGY↗

The NASA Open Science Data Repository: Biomedical Data, Analysis Tools, and Informatic Collaborations

Increased biomedical risks and challenges associated with deep space missions require knowledge discovery, health countermeasures, and biomedical support capabilities. Maximally open-access and reusable data is needed by developers, scientists, and engineers to develop these systems. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database (ie., findable, accessible, interoperable, and reusable), and meets various scientific, technical, and operational needs. It offers users and submitters the ability to upload, download, search, share, analyze, cite, and visualize data across ‘omics, physiological, phenotypic, payload, hardware, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR is an expanded database, based upon the successes of NASA GeneLab. OSDR has >460 studies with datasets covering model organisms to non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets with raw files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) developed from industry norms. OSDR is collecting and curating biomedical human data from a new sub-orbital research flight and is open to more space life science/biomedical submissions from the international and commercial sectors. OSDR also recently began a collaboration with the European Space Agency (ESA) to collect and curate >200 terabytes of human and model organism data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics and ~50 physiological-phenotypic-imaging assay data types. Tools available for OSDR users include: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, and 3) a Multi-study visualization tool which enables users to look across and combine ‘omics datasets. There are ~600 volunteer OSDR Analysis Working Group (AWG) members providing feedback on scientific data/metadata standards and collaborating to mine-reuse OSDR in research. OSDR/GeneLab has enabled ~60 publications reusing data as of October 2023.

space biology↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

The existing and forthcoming data bases from NASA missions contain an abundance of information whose complexity cannot be efficiently tapped with simple statistical techniques. Powerful multivariate statistical methods already exist which can be used to harness much of the richness of these data. Automatic classification techniques have been developed to solve the problem of identifying known types of objects in multi parameter data sets, in addition to leading to the discovery of new physical phenomena and classes of objects. We propose an exploratory study and integration of promising techniques in the development of a general and modular classification/analysis system for very large data bases, which would enhance and optimize data management and the use of human research resources.

Djorgovski, Stanislav↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

The existing and forthcoming data bases from NASA missions contain an abundance of information whose complexity cannot be efficiently tapped with simple statistical techniques. Powerful multivariate statistical methods already exist which can be used to harness much of the richness of these data. Automatic classification techniques have been developed to solve the problem of identifying known types of objects in multiparameter data sets, in addition to leading to the discovery of new physical phenomena and classes of objects. We propose an exploratory study and integration of promising techniques in the development of a general and modular classification/analysis system for very large data bases, which would enhance and optimize data management and the use of human research resource.

Djorgovski, George↗

A Virtual Bioinformatics Knowledge Environment for Early Cancer Detection

Discovery of disease biomarkers for cancer is a leading focus of early detection. The National Cancer Institute created a network of collaborating institutions focused on the discovery and validation of cancer biomarkers called the Early Detection Research Network (EDRN). Informatics plays a key role in enabling a virtual knowledge environment that provides scientists real time access to distributed data sets located at research institutions across the nation. The distributed and heterogeneous nature of the collaboration makes data sharing across institutions very difficult. EDRN has developed a comprehensive informatics effort focused on developing a national infrastructure enabling seamless access, sharing and discovery of science data resources across all EDRN sites. This paper will discuss the EDRN knowledge system architecture, its objectives and its accomplishments.

knowledge systems↗