Search NASASearch

SEARCH · Search NASA

Results for “data discover”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Bioactivity Profiling of Chemical Mixtures for Hazard Characterization

Abstract The assessment and regulation of chemical toxicity to protect human health and the environment are done one chemical at a time and seldom at environmentally relevant concentrations. However, chemicals are found in the environment as mixtures, and their toxicity is largely unknown. Understanding the hazard posed by chemicals within the mixture is critical to enforce protective measures. Here, we demonstrate the application of bioactivity profiling of environmental water samples using the sentinel and ecotoxicology model species Daphnia to reveal the biomolecular response induced by exposure to real-world mixtures. We exposed a Daphnia strain to 30 sampled waters of the Chaobai River and measured the gene expression response profiles. Using a multiblock correlation analysis, we establish correlations between chemical mixtures identified in 30 water samples with gene expression patterns induced by these chemical mixtures. We identified 80 metabolic pathways putatively activated by mixtures of inorganic ions, heavy metals, polycyclic aromatic hydrocarbons, industrial chemicals, and a set of biocides, pesticides, and pharmacologically active substances. Our data-driven approach discovered both known bioactivity signatures with previously described modes of action and new pathways linked to undiscovered potential hazards. This study demonstrates the feasibility of reducing the complexity of real-world mixture toxicity to characterize the biomolecular effects of a defined number of chemical components based on gene expression monitoring of the sentinel species Daphnia.

Engineering

Data for reproducing the figures of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas

This deposit contains the raw data for reproducing research results of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas. The main contribution of this work is to utilize machine learning techniques to reconstruct and enhance the resolution of a diagnostic measurement from other available diagnostics in a system. The proposed techniques is called Diag2Diag.

diag2diag

Ultra-faint Milky Way Satellites Discovered in Carina, Phoenix, and Telescopium with DELVE Data Release 3

We report the discovery of three Milky Way satellite candidates: Carina IV, Phoenix III, and DELVE 7, in the third data release of the DECam Local Volume Exploration survey (DELVE). The candidate systems were identified by cross-matching results from two independent search algorithms. All three are extremely faint systems composed of old, metal-poor stellar populations (τ ≳ 10 Gyr, [Fe/H] ≲−1.4). Carina IV (M V = −2.8; r 1/2 = 40 pc) and Phoenix III (M V = −1.2; r 1/2 = 19 pc) have half-light radii that are consistent with the known population of dwarf galaxies, while DELVE 7 (M V = 1.2; r 1/2 = 2 pc) is very compact and seems more likely to be a star cluster, though its nature remains ambiguous without spectroscopic follow-up. The Gaia proper motions of stars in Carina IV ($M_{\star} = 2250^{+1180}_{-830} M_⊙$) indicate that it is unlikely to be associated with the LMC, while DECam CaHK photometry confirms that its member stars are metal poor. Phoenix III ($M_{\star} = 520^{+660}_{-290} M_⊙$) is the faintest known satellite in the extreme outer stellar halo (D GC > 100 kpc), while DELVE 7 ($M_{\star} = 60^{+120}_{-40} M_⊙$) is the faintest known satellite with D GC > 20 kpc.

Tan, Chin Yi [Univ. of Chicago, IL (United States)

Elimination LArTPC Simulation Uncertainty

Liquid Argon Time Projection Chambers (LArTPC) are essential for detecting muons and neutrinos by capturing electrons released during particle collisions, which drift toward wire planes under an electric field and induce currents measured to reconstruct particle paths. However, LArTPCs face challenges from effects such as electron-ion recombination, electron diffusion, and electron attenuation, complicating data simulation. The Short Baseline Neutrino (SNB) detector aims to measure neutrinos before oscillation occurs. To bridge the gap between simulation and actual data, we propose modifying the amplitude and width of signals on the TPC wires, addressing uncertainties by adjusting signal characteristics to better match observed data. A Gaussian fit to current waveforms produces hits with associated charge and width, and by comparing data and simulated values, discrepancies highlight areas where the model fails. Initial results indicate the current modification algorithm may increase divergence between simulation and data, necessitating further refinement. A discovered bug in the WireModMakeHist_plug.cpp file, which incorrectly computed simulation and data ratios, underscores the need for precise algorithm adjustments. Future work involves correcting code errors, fine-tuning the model, and conducting multiple simulation runs to enhance statistical confidence and reduce uncertainties, ultimately aiming for accurate LArTPC operation and reliable neutrino detection.

Mkrtchyan, Ka'ren

The z ∼ 1.03 Merging Cluster SPT-CL J0356–5337: New Strong Lensing Analysis with HST and MUSE

We present a strong lensing analysis and reconstruct the mass distribution of SPT-CL J0356−5337, a galaxy cluster at redshift z = 1.034. Our model supersedes previous models by making use of new multiband Hubble Space Telescope data and Multi-Unit Spectroscopic Explorer (MUSE) spectroscopy. We identify two additional lensed galaxies to inform a more well-constrained model using 12 sets of multiple images in five separate lensed sources. The three previously known sources were spectroscopically confirmed by G. Mahler et al. at redshifts of z = 2.363, z = 2.364, and z = 3.048. We measured the spectroscopic redshifts of two of the newly discovered arcs using MUSE data, at z = 3.0205 and z = 5.3288. We increase the number of cluster member galaxies by a factor of 3 compared to previous work. We also report the detection of extended Lyα emission from several background galaxies. We measure the total projected mass density of the two major subcluster components, one dominated by the brightest cluster galaxy and the other by a compact group of luminous red galaxies. We find ${M}_{{\rm{B}}{\rm{C}}{\rm{G}}}(\lt 80\,\rm{kpc})=3.9{3}_{-0.14}^{+0.21}\times 1{0}^{13}$ M⊙ and ${M}_{{\rm{LRG}}}(\lt 80\,\rm{kpc})=2.9{2}_{-0.23}^{+0.16}\times 1{0}^{13}$ M⊙, yielding a mass ratio of $1.3{5}_{-0.08}^{+0.16}$ . The strong lensing constraints offer a robust estimate of the projected mass density regardless of modeling assumptions; allowing more substructure in this line of sight does not change the results or conclusions. Our results corroborate the conclusion that SPT-CL J0356−5337 is dominated by two mass components and is likely undergoing a major merger on the plane of the sky.

79 ASTRONOMY AND ASTROPHYSICS

The ECP SICM project: Managing complex memory hierarchies for exascale applications

The Exascale Computing Project (ECP)’s Simplified Interface to Complex Memories (SICM) effort focuses on developing universal interfaces for discovering, managing, and sharing data across complex memory hierarchies. These facilitate the exploitation of emerging memory technologies and support precise control over their various trade-offs such as high-bandwidth versus low-latency, persistent versus ephemeral, high-capacity versus low-capacity, and near-CPU versus near-GPU. SICM comprises three interrelated components: a low-level interface, a high-level interface, and a persistent-heap interface. The low-level SICM interface is intended for system and run-time developers as well as expert application developers who prefer full control of the memory objects used within their application. The high-level SICM interface builds upon the low-level interface, employing application-level profiling and analysis to optimize data management for complex memory hierarchies. The persistent-heap interface provides applications with a persistent memory allocator that can allocate custom C++ data structures in both block-storage and byte-addressable persistent memories.

97 MATHEMATICS AND COMPUTING

Characterising the Higgs boson with ATLAS data from the LHC Run-2

The Higgs boson was discovered by the ATLAS and CMS Collaborations in 2012 using data from Run 1 of the Large Hadron Collider (2010 2012). In Run 2 (2015 2018), about 140 fb -1 of proton–proton collisions at a centre-of-mass energy of 13 TeV were collected by the ATLAS experiment. This review presents the most important Run 2 results obtained by the ATLAS Collaboration regarding the properties of the Higgs boson and its interactions with other particles. The performed studies significantly enhance the understanding of the Higgs boson, while hunting for deviations from the predictions of the Standard Model of particle physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR

The silicon citizen naturalist

Smartphone-wielding citizen scientists and an AI called FLORIST are transforming ecology at the continental scale. Here, in this issue of Cell, when Tibbs-Cortes et al. pair the crowdsourced data with controlled genetics, they discover how switchgrass times its flowering to outwit both frost and heat, depending on latitude.

Hudson, Matthew E. [University of Illinois at Urba

Subsurface Energy Systems Mapping Inquiry Tool (MapIT)

The Subsurface Energy Systems Mapping Inquiry Tool (MapIT) is an online web mapping tool designed to help users discover available public-sourced data to facilitate data exploration for subsurface energy exploration and characterization efforts for resource identification (e.g. critical minerals, hydrocarbons, geothermal) as well as injection of geologic sequestration of carbon dioxide (e.g. enhanced oil recovery, saline storage, etc.). Modules within the tool curate data related to geology, faults, fractures, injection and confining zones, hydrologic information, groundwater, groundwater wells, geomechanical and petrophysical data, and geochemical data. User documentation on how to use the tool is also provided. Data have been collected from authoritative national, state, and local sources and made available in this tool. The data is also available as a data catalog and Esri Geodatabase at: https://edx.netl.doe.gov/dataset/mapit-database Disclaimer: There is no guarantee of completeness or appropriateness for individual user’s requirements. Use of this tool is solely at the discretion of the user. See full Federal Disclaimer for further information (https://netl.doe.gov/home/disclaimer). This project was funded by the United States Department of Energy, National Energy Technology Laboratory, in part, through a site support contract. Neither the United States Government nor any agency thereof, nor any of their employees, nor the support contractor, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. https://www.netl.doe.gov/home/disclaimer

Carbon Sequestration

Carbon Storage Site Mapping Inquiry Tool (MapIT)

The Carbon Storage Site Mapping Inquiry Tool (MapIT) is an online web mapping tool designed to help users discover available public-sourced data to facilitate data exploration in support of Underground Injection Control (UIC) Program Class VI Well Site permitting for the geologic sequestration of carbon dioxide.

Pantaleone, Scott

TSDC: Transportation Secure Data Center: Real-World Data for Planning, Modeling, and Analysis

The Transportation Secure Data Center is a centralized repository for high-resolution transportation data from hundreds of travel and transit surveys and studies. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. It houses surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Meanwhile, the Livewire Data Platform empowers research, industry, and academic partners to easily and securely preserve, maintain, share, discover, and gain access to transportation and mobility data. Livewire accommodates a range of datasets, including behavioral, experimental, model, analytical, and raw data at the vehicle, traveler, and system levels. Datasets support mobility research and planning spanning urban science, connected and automated vehicles, fueling and charging infrastructure, mobility decision science, multimodal transportation, vehicle efficiency, and more.

33 ADVANCED PROPULSION SYSTEMS

DESI Strong Lens Foundry. III. Keck Spectroscopy for Strong Lenses Discovered Using Residual Neural Networks

We present spectroscopic data of strong lenses and their source galaxies using the Keck Near-Infrared Echellette Spectrometer (NIRES) and the Dark Energy Spectroscopic Instrument (DESI), providing redshifts necessary for nearly all strong-lensing applications with these systems, especially the extraction of physical parameters from lensing modeling. These strong lenses were found in the DESI Legacy Imaging Surveys using residual neural networks and followed up by our Hubble Space Telescope program, with all systems displaying unambiguous lensed arcs. With NIRES, we target eight lensed sources at redshifts difficult to measure in the optical range and determine the source redshifts for six, between z s = 1.675 and 3.332. DESI observed one of the remaining source redshifts, as well as an additional source redshift within the six systems. The two systems with nondetections by NIRES were observed for a considerably shorter 600 s at high airmass. Combining NIRES infrared spectroscopy with optical spectroscopy from our DESI Strong Lensing Secondary Target Program, these results provide the complete lens and source redshifts for six systems, a resource for refining automated strong lens searches in future deep- and wide-field imaging surveys and addressing a range of questions in astrophysics and cosmology.

Agarwal, Shrihan [University of Chicago, IL (Unite

Collaborative: in situ visual analytics technologies for extreme scale combustion simulations

This project aims to drastically enhance the usability of in situ analysis and visualization for extreme-scale scientific simulations. Current exascale computing capabilities promise to offer greater predictive ability of simulations and to further push the frontiers of science and technology. However, to validate the simulation output at extreme scale, examine the modeled phenomena, and discover previously unknowns from the output data, the output must be reduced or transformed in situ as it is being generated during the simulation such that the amount of data to examine and store is kept to a minimum. Such in situ approaches allow us to process and analyze the data and any embedded geometry to an extent that would be prohibitively expensive, if not impossible, to perform as a post hoc task. While in situ processing has been demonstrated to be a feasible and promising approach, its full potential has not yet been leveraged. In this project, we have developed comprehensive enhancements to in situ technology based on probability distributions in data. Our research focuses on jointly developing new ways of interacting with massive statistical samples while creatively utilizing new state-of-the-art computational resources to push the boundaries of in situ exploration. Moreover, we have developed new time-dependent techniques to enable previously unattainable capabilities in areas such as intelligent simulation steering and precise feature identification. We have experimentally studied our design and implementation at NERSC and OLCF, and are able to leverage existing in situ infrastructures whenever possible. While the exemplar in this project is combustion, many other fields for which turbulent transport is important, e.g., fusion, climate, astrophysics among others, encounter similar issues as simulations scale up to the exascale. This project shows its potential to generate high impact on DOE missions since the resulting technology promises to improve scientists’ ability to rapidly and correctly interpret and tune extreme-scale simulations, leading to new scientific understanding and advancements.

97 MATHEMATICS AND COMPUTING

Risks Associated with Sharing the MOSSAIC APIs

The MOSSAIC APIs contain two files which, in theory, could be used to discover information about the pathology report data from the SEER registries on which the AI models were trained. In this document, we explain the contents of these files and assess the associated risk. APPENDIX A contains a set of slides to aid in the dissemination of this information.

97 MATHEMATICS AND COMPUTING

Physical discovery in representation learning via conditioning on prior knowledge

Recent advances in electron, scanning probe, optical, and chemical imaging and spectroscopy yield bespoke data sets containing the information of structure and functionality of complex systems. In many cases, the resulting data sets are underpinned by low-dimensional simple representations encoding the factors of variability within the data. The representation learning methods seek to discover these factors of variability, ideally further connecting them with relevant physical mechanisms. However, generally, the task of identifying the latent variables corresponding to actual physical mechanisms is extremely complex. Here, we present an empirical study of an approach based on conditioning the data on the known (continuous) physical parameters and systematically compare it with the previously introduced approach based on the invariant variational autoencoders. The conditional variational autoencoder (cVAE) approach does not rely on the existence of the invariant transforms and hence allows for much greater flexibility and applicability. Interestingly, cVAE allows for limited extrapolation outside of the original domain of the conditional variable. However, this extrapolation is limited compared to the cases when true physical mechanisms are known, and the physical factor of variability can be disentangled in full. We further show that introducing the known conditioning results in the simplification of the latent distribution if the conditioning vector is correlated with the factor of variability in the data, thus allowing us to separate relevant physical factors. We initially demonstrate this approach using 1D and 2D examples on a synthetic data set and then extend it to the analysis of experimental data on ferroelectric domain dynamics visualized via piezoresponse force microscopy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND