Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discover”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

From clutter to clarity: Emergent neural operators via questionnaire metrics

Real-world datasets in chemical engineering and bioengineering processes—such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials—can often be unlabeled or disorganized, rendering the training of existing supervised learning models ineffective at learning the underlying dynamics. To salvage these datasets for decision-making, we first seek to obtain clarity from the cluttered data. Here, we present a framework for developing “structural” generative models, discovering emergent equations, and constructing efficient emulators from scrambled datasets by integrating unsupervised organizational learning techniques (Questionnaires) with advanced deep learning architectures (Deep Hidden Physics Models and Deep Operator Networks). Our approach is demonstrated on two illustrative model systems: (a) a 1D advection–diffusion partial differential equation representing a winding underground pipe and (b) an ensemble of Stuart–Landau oscillators, an agent-based system of coupled ordinary differential equations. In both cases, we successfully reconstruct meaningful spatial, temporal, and parameter embeddings from scrambled data, enabling good predictions of system dynamics. As a result, we highlight the framework’s potential for broader applications, enabling data-driven system identification in fields with inherently disorganized or hidden parameter spaces.

42 ENGINEERING↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Factorization Machine‐Based Active Learning for Functional Materials Design with Optimal Initial Data

The optimization of functional materials is important to enhance their properties, but their complex geometries pose great challenges to optimization. Data-driven algorithms efficiently navigate such complex design spaces by learning relationships between material structures and performance metrics to discover high-performance functional materials. Surrogate-based active learning, continually improving its surrogate model by iteratively including high-quality data points, has emerged as a cost-effective data-driven approach. Furthermore, it can be coupled with quantum computing to enhance optimization processes, especially when paired with a special form of surrogate model (i.e., quadratic unconstrained binary optimization), formulated by factorization machine (FM). However, current practices often overlook the variability in design space sizes when determining the initial data size for optimization. In this work, we investigate the optimal initial data sizes required for efficient convergence across various design space sizes. By employing averaged piecewise linear regression, we identify initiation points where convergence begins, highlighting the crucial role of employing adequate initial data in achieving efficient optimization. These results contribute to the efficient optimization of functional materials by ensuring faster convergence and reducing computational costs in FM-based active learning.

active learning↗

The Dark Energy Survey: Cosmology Results with ∼1500 New High-redshift Type Ia Supernovae Using the Full 5 yr Data Set

Abstract We present cosmological constraints from the sample of Type Ia supernovae (SNe Ia) discovered and measured during the full 5 yr of the Dark Energy Survey (DES) SN program. In contrast to most previous cosmological samples, in which SNe are classified based on their spectra, we classify the DES SNe using a machine learning algorithm applied to their light curves in four photometric bands. Spectroscopic redshifts are acquired from a dedicated follow-up survey of the host galaxies. After accounting for the likelihood of each SN being an SN Ia, we find 1635 DES SNe in the redshift range 0.10 < z < 1.13 that pass quality selection criteria sufficient to constrain cosmological parameters. This quintuples the number of high-quality z > 0.5 SNe compared to the previous leading compilation of Pantheon+ and results in the tightest cosmological constraints achieved by any SN data set to date. To derive cosmological constraints, we combine the DES SN data with a high-quality external low-redshift sample consisting of 194 SNe Ia spanning 0.025 < z < 0.10. Using SN data alone and including systematic uncertainties, we find Ω M = 0.352 ± 0.017 in flat ΛCDM. SN data alone now require acceleration ( q 0 < 0 in ΛCDM) with over 5 σ confidence. We find ( Ω M , w ) = ( 0.264 − 0.096 + 0.074 , − 0.80 − 0.16 + 0.14 ) in flat w CDM. For flat w 0 w a CDM, we find ( Ω M , w 0 , w a ) = ( 0.495 − 0.043 + 0.033 , − 0.36 − 0.30 + 0.36 , − 8.8 − 4.5 + 3.7 ) , consistent with a constant equation of state to within ∼2 σ . Including Planck cosmic microwave background, Sloan Digital Sky Survey baryon acoustic oscillation, and DES 3 × 2pt data gives (Ω M , w ) = (0.321 ± 0.007, −0.941 ± 0.026). In all cases, dark energy is consistent with a cosmological constant to within ∼2 σ . Systematic errors on cosmological parameters are subdominant compared to statistical errors; these results thus pave the way for future photometrically classified SN analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Unsupervised discovery of extreme weather events using universal representations of emergent organization

Spontaneous self-organization is ubiquitous in systems far from thermodynamic equilibrium. While organized structures that emerge dominate transport properties, universal representations that identify and describe these key objects remain elusive. Here, we introduce a theoretically grounded framework for describing emergent organization that, via data-driven algorithms, is constructive in practice. Its building blocks are spacetime lightcones that embody how information propagates across a system through local interactions. We show that predictive equivalence classes of lightcones—local causal states—capture organized behaviors in complex spatiotemporal systems. Employing an unsupervised physics-informed machine learning algorithm and a high-performance computing implementation, we demonstrate automatically discovering organized structures in two real-world domain science problems. We show that local causal states identify vortices and track their power-law decay behavior in two-dimensional fluid turbulence. We then show how to detect and track familiar extreme weather events—hurricanes and atmospheric rivers—and discover other novel structures associated with precipitation extremes in high-resolution climate data at the grid-cell level.

Rupe, Adam [Pacific Northwest National Laboratory ↗

Venturing into Unexplored Phase Space: Synthesis, Structure, and Properties of MgCo 3 B 2 Featuring a Rumpled Kagomé Network

MgCo 3 B 2 , a novel ternary boride in a previously unexplored phase space, was synthesized using the hydride route. In situ powder X-ray diffraction and DFT calculations aided in the discovery of this compound, whose structure was then determined by single-crystal X-ray diffraction. Like the closely related CeCo 3 B 2 , MgCo 3 B 2 crystallizes in centrosymmetric space group P6/mmm (a = 4.883(2) Å, c = 2.926(2) Å at 210 K, Z = 1). Unlike CeCo 3 B 2 , however, it adopts a disordered structure that features a rumpled Kagomé network of Co atoms, and Mg atoms fill the channels of a Co–B framework. Although the structural disorder leads to motifs that are similar to those observed in MgNi 3 B 2 and other related ternary borides, no evidence of an ordered superstructure was found by single-crystal X-ray diffraction or high-resolution powder X-ray diffraction. In the case of CeCo 3 B 2 , boron atoms occupy the center of regular Co 6 trigonal prisms; in MgCo 3 B 2 , boron atoms are shifted from the center of the prism to form B–B dimers with roughly the same length as those found in MgNi 3 B 2 . Magnetic susceptibility data exhibit an unusual temperature dependence that cannot be convincingly modeled by the modified Curie–Weiss equation, consistent with DFT calculations predicting a nonmagnetic ground state. Intrinsic susceptibility at 300 K is 1.42 × 10 –3 emu/mol Oe, which is comparable to that of paramagnetic YCo 3 B 2 and CeCo 3 B 2 with a similar structure and composition. Here, this study showcases the efficacy of combining several methodologies to discover new solids in unexplored phase spaces. This approach includes in situ PXRD data to monitor reactions of precursors upon heating, a diffusion-enhanced synthesis method, and DFT assessment of compound stability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

In-Depth Proteome Profiling of the Hippocampus of LDLR Knockout Mice Reveals Alternation in Synaptic Signaling Pathway

The low-density lipoprotein receptor (LDLR) is a major apolipoprotein receptor that regulates cholesterol homeostasis. LDLR deficiency is associated with cognitive impairment by the induction of synaptopathy in the hippocampus. Despite the close relationship between LDLR and neurodegenerative disorders, proteomics research for protein profiling in the LDLR knockout (KO) model remains insufficient. Therefore, understanding LDLR KO-mediated differential protein expression within the hippocampus is crucial for elucidating a role of LDLR in neurodegenerative disorders. In this study, we conducted first-time proteomic profiling of hippocampus tissue from LDLR KO mice using tandem mass tag (TMT)-based MS analysis. LDLR deficiency induces changes in proteins associated with the transport of diverse molecules, and activity of kinase and catalyst within the hippocampus. Additionally, significant alterations in the expression of components in the major synaptic pathways were found. Furthermore, these synaptic effects were verified using a data-independent acquisition (DIA)-based proteomic method. In conclusion, our data will serve as a valuable resource for further studies to discover the molecular function of LDLR in neurodegenerative disorders.

60 APPLIED LIFE SCIENCES↗

SODAs: sparse optimization for the discovery of differential and algebraic equations

Differential-algebraic equations (DAEs) integrate ordinary differential equations (ODEs) with algebraic constraints, providing a fundamental framework for developing models of dynamical systems characterized by time-scale separation, conservation laws and physical constraints. While sparse optimization has revolutionized model development by allowing data-driven discovery of parsimonious models from a library of possible equations, existing approaches for dynamical systems assume DAEs can be reduced to ODEs by eliminating variables before model discovery. This assumption limits the applicability of such methods for DAE systems with unknown constraints and time scales. We introduce sparse optimization for differential-algebraic systems (SODAs), a data-driven method for the identification of DAEs in their explicit form. By discovering the algebraic and dynamic components sequentially without prior identification of the algebraic variables, this approach leads to a sequence of convex optimization problems. It has the advantage of discovering interpretable models that preserve the structure of the underlying physical system. To this end, SODAs improves since SODAs is singular numerical stability when handling high correlations between library terms, caused by near-perfect algebraic relationships, by iteratively refining the conditioning of the candidate library. We demonstrate the performance of our method on biological, mechanical and electrical systems, showcasing its robustness to noise in both simulated time series and real-time experimental data.

DAE↗

Multimodal super-resolution: discovering hidden physics and its application to fusion plasmas

Understanding complex physical systems often requires integrating data from multiple diagnostics, each with limited resolution or coverage. We present a machine learning framework that reconstructs synthetic high-temporal-resolution data for a target diagnostic using information from other diagnostics, without direct target measurements during the inference. This multimodal super-resolution technique improves diagnostic robustness and enables monitoring even in case of measurement failures or degradation. Applied to fusion plasmas, our method targets edge-localized modes (ELMs), which can damage plasma-facing materials. By reconstructing super-resolution Thomson Scattering data from complementary diagnostics, we uncover fine-scale plasma dynamics and validate the role of resonant magnetic perturbations (RMPs) in ELM suppression through magnetic island formation. The approach provides new observation supporting the plasma profile flattening due to these islands. Our results demonstrate the framework’s ability to generate high-fidelity synthetic diagnostics, offering a powerful tool for ELM control development in future reactors like ITER. The approach is broadly transferable to other domains facing sparse, incomplete, or degraded diagnostic data, opening new avenues for discovery.

Jalalvand, Azarakhsh [Princeton Univ., NJ (United ↗

Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents

Membership inference attacks (MIAs), which enable adversaries to determine whether specific data points were part of a model's training dataset, have emerged as an important framework to understand, assess, and quantify the potential information leakage associated with machine learning systems. Designing effective MIAs is a challenging task that usually requires extensive manual exploration of model behaviors to identify potential vulnerabilities. In this paper, we introduce AutoMIA -- a novel framework that leverages large language model (LLM) agents to automate the design and implementation of new MIA signal computations. By utilizing LLM agents, we can systematically explore a vast space of potential attack strategies, enabling the discovery of novel strategies. Our experiments demonstrate AutoMIA can successfully discover new MIAs that are specifically tailored to user-configured target model and dataset, resulting in improvements of up to 0.18 in absolute AUC over existing MIAs. This work provides the first demonstration that LLM agents can serve as an effective and scalable paradigm for designing and implementing MIAs with SOTA performance, opening up new avenues for future exploration.

Tran, Toan Viet [Emory University]↗

MIJAMP

Here, we develop MIJAMP (MIJAMP Is Just A MethylBED Parser), a software package that was developed to discover methylated motifs from the output of ONT’s Modkit or other data in the methylBED format. MIJAMP employs a human-driven refinement strategy that empirically validates all motifs against genome-wide methylation data, thus eliminating false, under-, or overexplained motifs. MIJAMP can also report methylation data on a specific, user-defined motif.

Alexander, William↗

Optimal Control of an Oscillating Surge Wave Energy Converter

During this project, we experimentally investigated the hydrodynamics and performance of a laboratory-scale oscillating surge wave energy converter (OSWEC).We looked at how flap buoyancy and driveline losses (primarily in the form of stiction) affected the dynamics and performance of the device. In addition, we assessed the influence of flap profile (rounded vs. square edges) on OSWEC hydrodynamics. Through this, we were able to develop a deeper understanding of OSWEC performance and provide guidance on strategies to counteract artifacts that may be present in laboratory models, but are absent in field-scale devices. To do this, we tested a laboratory-scale OSWEC in the Sea Wave Environmental Lab (SWEL) wave tank at the National Renewable Energy Laboratory (NREL). We ran several types of experiments to investigate the hydrodynamics and performance of the device. Overall, we achieved the overall goal of experimentally investigating the hydrodynamics and performance of this device. We discovered important and unexpected trends in performance, and collected time-resolved data to help us further investigate the underlying hydrodynamics responsible for these trends. In addition, we are currently using the time-resolved data from these experiments to build data-driven models of the dynamics, which can in turn be used to inform data-driven model predictive control of this device and address this objective in the future.

16 TIDAL AND WAVE POWER↗

Learning new physics from data: A symmetrized approach

Thousands of person years have been invested in searches for new physics (NP), the majority of them motivated by theoretical considerations. Yet, no evidence of beyond the Standard Model physics has been found. This suggests that model-agnostic searches might be an important key to explore NP, and help discover unexpected phenomena which can inspire future theoretical developments. A possible strategy for such searches is identifying asymmetries between data samples that are expected to be symmetric within the Standard Model. We propose exploiting neural networks (NNs) to quickly fit and statistically test the differences between two samples. Our method is based on an earlier work, originally designed for inferring the deviations of an observed dataset from that of a much larger reference dataset. We present a symmetric formalism, generalizing the original one, avoiding fine-tuning of the NN parameters and any constraints on the relative sizes of the samples. Our formalism could be used to detect small symmetry violations, extending the discovery potential of current and future particle physics experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Cocytos Stream: A Disrupted Globular Cluster from our Last Major Merger?

The census of stellar streams and dwarf galaxies in the Milky Way provides direct constraints on galaxy formation models and the nature of dark matter. The DESI Milky Way survey (with a footprint of 14,000$~deg{^2}$ and a depth of $r<19$ mag) delivers the largest sample of distant metal-poor stars compared to previous optical fiber-fed spectroscopic surveys. This makes DESI an ideal survey to search for previously undetected streams and dwarf galaxies. We present a detailed characterization of the Cocytos stream, which was re-discovered using a clustering analysis with a catalog of giants in the DESI year 3 data, supplemented with Magellan/MagE spectroscopy. Our analysis reveals a relatively metal-rich ([Fe/H]$=-1.3$) and thick stream (width$=1.5^\circ$) at a heliocentric distance of $\approx 25$ kpc, with an internal velocity dispersion of 6.5-9 km s$^{-1}$. The stream's metallicity, radial orbit, and proximity to the Virgo stellar overdensities suggest that it is most likely a disrupted globular cluster that came in with the Gaia-Enceladus merger. We also confirm its association with the Pyxis globular cluster. Our result showcases the ability of wide-field spectroscopic surveys to kinematically discover faint disrupted dwarfs and clusters, enabling constraints on the dark matter distribution in the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

Tools for probing new physics with newly discovered gamma-ray targets

Here, we present a computational tool, TweedleDEE, for empirically modeling diffuse gamma-ray background emission in a 1° region of the sky, using publicly available gamma-ray data off-axis from the region of interest. This background model allows a user to perform a purely data-driven search for anomalous localized sources of gamma-ray emission, including new physics. A major application of this tool would be in searching for dark matter annihilation in newly discovered astrophysical targets. For this purpose, we derive a scaling relation for determining velocity-dependent 𝐽-factors using only the stellar parameters, which can be broadly applied to obtain dark matter constraints from new targets. As an application of these tools, we use TweedleDEE and MADHATv2 to perform the first search for dark matter annihilation in the newly discovered Leo VI dwarf spheroidal galaxy, and present model constraints for a variety of choices of the annihilation channel and velocity dependence of the cross section.

dark matter↗

Vera C. Rubin Observatory Prompt Products

Data products produced by prompt and daily processing of images obtained in the Legacy Survey of Space and Time. These include realtime alerts sent to community alert brokers, newly-discovered Solar System Objects reported to the Minor Planet Center, processed visit and difference images, and source catalogs. Prompt Products are not a static single data release but continually grow throughout the ten-year LSST survey.

79 ASTRONOMY AND ASTROPHYSICS↗

Vera C. Rubin Observatory Prompt Products: alert packets data

Data products produced by prompt and daily processing of images obtained in the Legacy Survey of Space and Time. These include realtime alerts sent to community alert brokers, newly-discovered Solar System Objects reported to the Minor Planet Center, processed visit and difference images, and source catalogs. Prompt Products are not a static single data release but continually grow throughout the ten-year LSST survey. This dataset is a subset of the full data release consisting of a dataset named alert packets. This dataset contains measurements for 5-sigma sources detected in difference images that were issued to the community brokers.

79 ASTRONOMY AND ASTROPHYSICS↗