Search NASA⌕ Search

SEARCH · Search NASA

Results for “knowledge discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

122 records · Page 7

Causal Directions Matter: How Environmental Factors Drive Convective Cloud Detrainment Heights

This study investigates how environmental factors influence the level of maximum detrainment (LMD) in deep convective clouds. Through a novel application of the Linear Non‐Gaussian Acyclic Model (LiNGAM), we discover causal structures between environmental variables and LMD, observed at six tropical sites operated by the Atmospheric Radiation Measurement (ARM) user facility. LiNGAM effectively identifies causal directions among variables of interest, revealing robust relationships such as those among the lifting condensation level (LCL), level of free convection (LFC), and convective inhibition (CIN), aligning with prior knowledge. Relative humidity is shown to directly influence LMD; however, this relationship exhibits strong nonlinearity and becomes difficult to detect when the contrast between oceanic and continental environments is excluded from the analysis. This study highlights the importance of establishing causal relationships before performing statistical inference.

54 ENVIRONMENTAL SCIENCES↗

Oakland University Cybersecurity Center (Final Scientific/Technical Report)

This report summarizes the outcomes of Award DE-CR0000023, “Oakland University Cybersecurity Center,” a 31-month project funded by the U.S. Department of Energy Office of Cybersecurity, Energy Security, and Emergency Response (CESER). The project addressed cybersecurity risks facing small and medium-sized manufacturers (SMMs) transitioning to Industry 4.0. The project integrated customer discovery, applied research, and cybersecurity training development. A total of 51 cybersecurity assessments identified significant gaps in baseline practices, incident response, and workforce capability. Research efforts produced a scalable mitigation framework tailored to SMM environments, and workforce analysis identified persistent talent gaps. Eight cybersecurity training modules were developed and deployed via Oakland University’s Professional and Continuing Education (PACE) platform. All objectives were completed, with 98.93% federal budget utilization and cost share exceeding requirements. The project establishes a scalable model for strengthening cybersecurity resilience and workforce capacity across U.S. manufacturing supply chains.

24 POWER TRANSMISSION AND DISTRIBUTION↗

2024 NMDC Ambassador Training Materials [Slides]

The NMDC is a sustainable data discovery platform that promotes open science and shared-ownership across a broad and diverse community of researchers, funders, publishers, societies, and other collaborators. The NMDC aims to enable multi-omic microbiome research to accelerate scientific discovery. The NMDC is a Department of Energy funded program that is a collaboration between 3 National Laboratories: Lawrence Berkeley National Laboratory (LBNL), Los Alamos National Laboratory (LANL), and Pacific Northwest National Laboratory (PNNL).

54 ENVIRONMENTAL SCIENCES↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

Omics-Based Comparison of Fungal Virulence Genes, Biosynthetic Gene Clusters, and Small Molecules in Penicillium expansum and Penicillium chrysogenum

Penicillium expansum is a ubiquitous pathogenic fungus that causes blue mold decay of apple fruit postharvest, and another member of the genus, Penicillium chrysogenum, is a well-studied saprophyte valued for antibiotic and small molecule production. While these two fungi have been investigated individually, a recent discovery revealed that P. chrysogenum can block P. expansum-mediated decay of apple fruit. To shed light on this observation, we conducted a comparative genomic, transcriptomic, and metabolomic study of two P. chrysogenum (404 and 413) and two P. expansum (Pe21 and R19) isolates. Global transcriptional and metabolomic outputs were disparate between the species, nearly identical for P. chrysogenum isolates, and different between P. expansum isolates. Further, the two P. chrysogenum genomes revealed secondary metabolite gene clusters that varied widely from P. expansum. This included the absence of an intact patulin gene cluster in P. chrysogenum, which corroborates the metabolomic data regarding its inability to produce patulin. Additionally, a core subset of P. expansum virulence gene homologues were identified in P. chrysogenum and were similarly transcriptionally regulated in vitro. Molecules with varying biological activities, and phytohormone-like compounds were detected for the first time in P. expansum while antibiotics like penicillin G and other biologically active molecules were discovered in P. chrysogenum culture supernatants. Our findings provide a solid omics-based foundation of small molecule production in these two fungal species with implications in postharvest context and expand the current knowledge of the Penicillium-derived chemical repertoire for broader fundamental and practical applications.

Bartholomew, Holly P. (ORCID:0000000292726399)↗

Tunable Few-Layer van der Waals Crystals and Heterostructures as Emerging Energy and Quantum Materials (Final Technical Report)

2D and layered (van der Waals) semiconductors offer extraordinary opportunities for manipulating optically excited charge carriers, many-body excitations, and non-charge based quantum numbers. To date, research has focused on a limited group of materials, mostly transition metal dichalcogenides in the monolayer limit. Other van der Waals semiconductors, and especially few-layer to multilayer crystals and their heterostructures, carry large potential for the discovery of phenomena of interest for future energy and information technologies. But they remain largely unexplored, often due to a lack of access to high-quality materials and approaches for measuring their properties at the relevant scales. The goal of this project was to develop an EPSCoR-State/National Laboratory Partnership that addresses the challenges of preparing high-quality van der Waals semiconductors and of probing their structure, composition, and especially their optoelectronic and photonic properties, near the atomic scale using electron microscopy techniques. A central component of the project was the development of advanced methods for electron microscopy and electron-excited spectroscopy, taking advantage of unique samples as well as leading capabilities and expertise at the partner institutions. Efforts to advance leading-edge techniques was supported by ancillary developments, such as precision sample preparation for electron microscopy/spectroscopy, coordinated chemical imaging, and analytical electron microscopy. Experiments in materials synthesis and technique development were closely linked to theory and computation. The results obtained under this project yielded multifaceted benefits to the involved partners and their institutions, DOE-BES, the wider scientific community, and society at large, particularly in the State of Nebraska through dividends from knowledge and human capital generated under the project.

36 MATERIALS SCIENCE↗

Metadata Standards for the NSE: Extended Field Standards

This standard presents a set of optional metadata fields for managed digital objects within the Nuclear Security Enterprise (NSE) and provides a deeper look at data representation in metadata by looking at the representation of 1) Records Management required metadata, and 2) common representations of technical/scientific data. Metadata standardization is a critical enabler for effectively sharing data, documents, and other digital objects between NSE sites, and for tracing the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a complementary field standard, recommending an optional set of fields that should be uniformly built for all managed digital objects within the NSE. This document specifically focuses on extending the shared discovery layer defined in the first white paper by introducing additional descriptive and data representation fields that improve cross-site search and interpretation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Dynamic and Responsive Distributed Energy Resource Education Solutions for Building, Fire, and Safety Department Officials (Final Technical Report)

From April 2021 through March 2024, the Interstate Renewable Energy Council (IREC) led a collaborative project to develop a free online clearinghouse of educational resources about solar photovoltaics (PV), energy storage systems (ESS), electric vehicle supply equipment (EVSE), and grid-interactive efficient building (GEB) technologies. Two websites—the Clean Energy Clearinghouse and CleanEnergyTraining.org—housed over 70 educational resources. Over the course of the three-year project, 154,272 unique visitors accessed the learning materials. Learner feedback was overwhelmingly positive. Even through the end of the project, there was sustained demand for education and communication. A primary innovation of the project was to drive multiple complementary audiences to the same place. Building owners, designers, installation contractors and developers, authorities having jurisdiction (AHJs), and fire service personnel all benefit from a shared understanding of clean energy technologies, including safety and code-related requirements. When considering the impact on the target audience, the project team worked with partners and advisors to inform resource creation and delivery in such a way as to address key motivational factors of the target audience and compel each user to seek additional information on the topic and return to the Clean Energy Clearinghouse website as their central location for more information. Resources were intentionally developed to be concise—five to 15 minutes—and accessible, meaning not overly technical. Providing basic information demystified the technologies and invited the professional to explore additional learning opportunities. Awardee and partner collaboration was key to project success. IREC facilitated collaboration among the other Topic 2 awardees, Southface and New Buildings Institute (NBI). The three awardees shared relevant information gained through discovery and validation questionnaires that informed product development and reduced duplication of effort by coordinating the development of complementary, and not competing, educational resources. Inspired by this collaboration, IREC brought on additional partners even in the final year of the project. Five regional energy efficiency organizations were part of the project, which expanded the connection between efficiency and distributed energy resources. We also included resources on the Clearinghouse that were developed through other federally funded projects, such as the Buildings Energy Efficiency Frontiers & Innovation Technologies (BENEFIT) program. The website was developed with the learner in mind, and not solely the funding source. Feedback from stakeholders throughout the project, and especially in its final year, indicated the need for continued education and facilitated communication among stakeholders to further the safe and widespread adoption of clean energy.

14 SOLAR ENERGY↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

Ensuring continued operation of INSPIRE as a cornerstone of the HEP information infrastructure

The INSPIRE platform — the most widely-used discovery service specifically tailored to the needs of researchers in High Energy Physics (HEP) — has become a central component of the information infrastructure for the discipline. Despite this, INSPIRE's continued sustainability is frequently endangered by resource constraints, recently made more acute by the loss of support from historical funders changing their research priorities. If the European particle physics community wishes to ensure INSPIRE's long-term sustainability, the community should secure international support and ensure appropriate funding.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗

Neutrons in Structural Biology: Challenges and Opportunities (Workshop Report)

Gaining a thorough understanding of biological systems requires building our knowledge about biological processes from the level of atoms and electrons, and up to whole organisms. Such comprehensive knowledge will allow for a predictive understanding of complex biological systems behavior. It will guide us in the design and development of novel therapeutics and vaccines to tackle existing health threats and to prepare for future pandemics, and it will provide information necessary to create new biomaterials and bio-inspired technologies through manipulation of biological macromolecules, their assemblies, single cells and even microorganisms. Reaching these goals will require a synergistic combination of multiple experimental techniques with molecular calculations and predictive simulations, and the design and development of new techniques and capabilities that bridge current knowledge and technology gaps. Neutron scattering provides unique information about the biomacromolecular structure and function and can play a major role in achieving these goals. A workshop was held to engage the scientific community in identifying pressing challenges in biochemistry, structural biology, enzymology and structure-guided drug design not solved with the current neutron scattering technologies or utilizing other structural biology techniques such as X-ray crystallography, NMR, and cryo-EM. The workshop brought together structural biology, biochemistry and computational experts, as well as early career researchers and students, creating a forum for discussing scientific advancement and collaboration. The workshop included a one-day satellite training workshop where graduate students and postdoctoral researchers were educated in the application of neutron crystallography and small-angle scattering in structural biology. Furthermore, the Instrument Scientific Advisory Board (ISAB) for the development of a macromolecular neutron diffractometer at ORNL’s Second Target Station was introduced at the workshop. The major outcome was that neutrons can provide atomic-level understanding of biomacromolecular structure, function and dynamics which is of paramount importance for addressing the identified challenges. Neutron crystallography, in particular, can resolve long-standing biochemical issues regarding enzyme function by delineating the underlying chemistry and can have a major impact on the design of small-molecule therapeutics, especially in combination with molecular computation (quantum chemistry and molecular dynamics simulations) and the emerging artificial intelligence (AI)-assisted drug design technologies. The unique properties of neutrons, including their high sensitivity to hydrogen and their non-destructive nature, make them ideal probes of biological matter. There is a palpable need in the scientific community to expand and enhance the impact of neutron sciences on biology. Neutron crystallography is the only structural biology method capable of determining positions of all hydrogen atoms in proteins, nucleic acids and their complexes at near-physiological temperatures and of unstable species at cryogenic temperatures. Moreover, neutron analysis is non-ionizing, non-destructive and does not perturb the structure or redox chemistry of active site metal centers and clusters in proteins, which can be invaluable for studying radiation-sensitive metalloprotein complexes. Further, neutron energies used in scattering applications are similar to atomic motions, permitting neutron spectroscopies to characterize the dynamics of biomacromolecules on the picosecond to microsecond timescales. The different sensitivities of neutrons to protium (H) and deuterium (D) isotopes of hydrogen allow enhanced visibility of specific parts of biological complexes through isotopic labeling. The impact of neutrons will be most powerful when neutron scattering is combined with complementary experimental techniques that use photons and electrons, and with high-performance computing. The interconnection and mutuality of the experimental and theoretical capabilities will drive discoveries in biological and health sciences to generate more complete picture of complex biological systems. The major limitation in the field of biological neutron crystallography has been signal-to-noise, demanding large samples that are difficult to produce for the majority of biomacromolecules and limiting the applicability of this technique in biological sciences. A neutron crystallography instrument at the Second Target Station will revolutionize biological science with neutrons by engaging a large scientific community of structural biologists, enabling successful neutron diffraction experiments from radically smaller biomacromolecular crystals, resolving unanswered biochemical questions, and meaningfully contributing to rational drug design. The meeting highlighted 10 grand challenges that will be addressed with this advanced capability over the next decade and beyond, and the recommendations required to help address them are given below.

59 BASIC BIOLOGICAL SCIENCES↗

Gerischer Electrochemistry Today

Semiconductor photoelectrochemistry is a dynamic and interdisciplinary field at the forefront of research in solar fuels, energy conversion, and catalysis. Here, this Perspective captures the collective insights from the second Gerischer Electrochemistry Today Symposium, held at Colorado State University in Fort Collins, CO, in August 2024, which convened leading researchers, early-career scientists, and industry partners to define the critical next steps for the field. Through interactive sessions, technical talks, panel discussions, and training initiatives─including a Semiconductor Electrochemistry Bootcamp─the symposium emphasized three pillars of advancement: (i) facilitating the exchange of new ideas in semiconductor electrochemistry and charge separation; (ii) fostering the development of future researchers, research topics, and participation in the semiconductor workforce; and (iii) building community. This Energy Focus distills key themes from the meeting and identifies major knowledge gaps in the following areas: mechanisms of charge separation and recombination, role of defects and disorder, dynamic and operando characterization methods, interfacial chemistry and surface passivation, theoretical and modeling limitations, and standardization and benchmarking. The inclusive and collaborative structure of the symposium enabled the generation of this comprehensive report that will serve as a roadmap for fundamental and applied research in the rapidly evolving field of semiconductor electrochemistry over the next decade.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗