Search NASA⌕ Search

SEARCH · Search NASA

Results for “Advanced Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗

Making Better Use of Satellite Data: The Satellite Needs Working Group

The U.S. Group on Earth Observations (USGEO) initiated in 2016 the Satellite Needs Working Group (SNWG) to identify and communicate the Earth observation needs of U.S. federal agencies. The SNWG identifies such needs through a biennial survey followed by interviews and follow-up discussions by the satellite Earth data providers of the U.S. Government: the National Aeronautics and Space Administration (NASA), the National Oceanic and Atmospheric Administration (NOAA), and the U.S. Geological Survey (USGS). Solutions and services are identified that leverage current or upcoming satellite missions to meet the identified needs; implementation of services that are estimated to significantly increase the level of satisfaction of multiple U.S. agencies are funded by NASA. The SNWG process has resulted in the implementation of numerous services that have impacted operations of not only U.S. agencies but of academic and international institutions as well. Notable examples include the Harmonized Landsat Sentinel-2 project that leverages European Space Agency (ESA) and NASA satellite assets to generate a global, analysis-ready, surface reflectance product with a temporal resolution of two days; the Airborne Data Management Group that curates and provides access to relevant resources, information and data from existing and past NASA field campaigns; and the Dynamic Surface Water Extent product which consists of harmonized but independent water extents derived from both from optical and radar data. These and other products are hosted at the NASA Distributed Active Archive Centers (DAACs) for free and open access. The SNWG Management Office at NASA’s Interagency Implementation and Advanced Concepts Team (IMPACT) manages the implementation of selected solutions and, importantly, NASA’s response to the needs of federal agencies through a Stakeholder Engagement Program. The impact of implemented services and solutions is hampered without efforts to build capacity around the use of the services. The Stakeholder Engagement Program ensures relevant training and outreach to SNWG agencies in collaboration with each solution implementation team.

remote sensing↗

NEEMO - NASA's Extreme Environment Mission Operations: On to a NEO

During NEEMO missions, a crew of six Aquanauts lives aboard the National Oceanic and Atmospheric Administration (NOAA) Aquarius Underwater Laboratory the world's only undersea laboratory located 5.6 km off shore from Key Largo, Florida. The Aquarius habitat is anchored 62 feet deep on Conch Reef which is a research only zone for coral reef monitoring in the Florida Keys National Marine Sanctuary. The crew lives in saturation for a week to ten days and conducts a variety of undersea EVAs (Extra Vehicular Activities) to test a suite of long-duration spaceflight Engineering, Biomedical, and Geoscience objectives. The crew also tests concepts for future lunar exploration using advanced navigation and communication equipment in support of the Constellation Program planetary exploration analog studies. The Astromaterials Research and Exploration Science (ARES) Directorate and Behavioral Health and Performance (BHP) at NASA/Johnson Space Center (JSC), Houston, Texas support this effort to produce a high-fidelity test-bed for studies of human planetary exploration in extreme environments as well as to develop and test the synergy between human and robotic curation protocols including sample collection, documentation, and sample handling. The geoscience objectives for NEEMO missions reflect the requirements for Lunar Surface Science outlined by the LEAG (Lunar Exploration Analysis Group) and CAPTEM (Curation and Analysis Planning Team for Extraterrestrial Materials) white paper [1]. The BHP objectives are to investigate best meas-ures and tools for assessing decrements in cogni-tive function due to fatigue, test the feasibility study examined how teams perform and interact across two levels, use NEEMO as a testbed for the development, deployment, and evaluation of a scheduling and planning tool. A suite of Space Life Sciences studies are accomplished as well, ranging from behavioral health and performance to immunology, nutrition, and EVA suit design results of which will directly support the investigation of open questions and operational concepts that will enable NASA to continue its plan for planetary exploration.

Bell, M. S.↗

Expanding on the BRIAR Dataset: A Comprehensive Whole Body Biometric Recognition Resource at Extreme Distances and Real-World Scenarios (Collections 1-4)

The state-of-the-art in biometric recognition algorithms and operational systems has advanced quickly in recent years providing high accuracy and robustness in more challenging collection environments and consumer applications. However, the technology still suffers greatly when applied to non-conventional settings such as those seen when performing identification at extreme distances or from elevated cameras on buildings or mounted to UAVs. This paper summarizes an extension to the largest dataset currently focused on addressing these operational challenges, and describes its composition as well as methodologies of collection, curation, and annotation.

Cornett, David [ORNL] (ORCID:0000000222910860)↗

GeoLab: A Geological Workstation for Future Missions

The GeoLab glovebox was, until November 2012, fully integrated into NASA's Deep Space Habitat (DSH) Analog Testbed. The conceptual design for GeoLab came from several sources, including current research instruments (Microgravity Science Glovebox) used on the International Space Station, existing Astromaterials Curation Laboratory hardware and clean room procedures, and mission scenarios developed for earlier programs. GeoLab allowed NASA scientists to test science operations related to contained sample examination during simulated exploration missions. The team demonstrated science operations that enhance theThe GeoLab glovebox was, until November 2012, fully integrated into NASA's Deep Space Habitat (DSH) Analog Testbed. The conceptual design for GeoLab came from several sources, including current research instruments (Microgravity Science Glovebox) used on the International Space Station, existing Astromaterials Curation Laboratory hardware and clean room procedures, and mission scenarios developed for earlier programs. GeoLab allowed NASA scientists to test science operations related to contained sample examination during simulated exploration missions. The team demonstrated science operations that enhance the early scientific returns from future missions and ensure that the best samples are selected for Earth return. The facility was also designed to foster the development of instrument technology. Since 2009, when GeoLab design and construction began, the GeoLab team [a group of scientists from the Astromaterials Acquisition and Curation Office within the Astromaterials Research and Exploration Science (ARES) Directorate at JSC] has progressively developed and reconfigured the GeoLab hardware and software interfaces and developed test objectives, which were to 1) determine requirements and strategies for sample handling and prioritization for geological operations on other planetary surfaces, 2) assess the scientific contribution of selective in-situ sample characterization for mission planning, operations, and sample prioritization, 3) evaluate analytical instruments and tools for providing efficient and meaningful data in advance of sample return and 4) identify science operations that leverage human presence with robotic tools. In the first year of tests (2010), GeoLab examined basic glovebox operations performed by one and two crewmembers and science operations performed by a remote science team. The 2010 tests also examined the efficacy of basic sample characterization [descriptions, microscopic imagery, X-ray fluorescence (XRF) analyses] and feedback to the science team. In year 2 (2011), the GeoLab team tested enhanced software and interfaces for the crew and science team (including Web-based and mobile device displays) and demonstrated laboratory configurability with a new diagnostic instrument (the Multispectral Microscopic Imager from the JPL and Arizona State University). In year 3 (2012), the GeoLab team installed and tested a robotic sample manipulator and evaluated robotic-human interfaces for science operations.

Evans, Cynthia↗

Mars Sample Return Science Planning Group Phase 2 (MSPG2): Overview & Interim Report

Mars Sample Return (MSR) has been a high priority of the international planetary science community for decades. In recent years, significant programmatic advances have brought MSR closer to becoming a reality. In 2018, NASA and the European Space Agency (ESA) signed a joint Statement of Intent to continue defining respective roles and responsibilities in the flight missions required to realize MSR. In October 2020, NASA and ESA formalized this partnership with the signature of a Memorandum of Understanding for the MSR flight elements. The MSR campaign consists of M2020, two MSR flight elements and the ground-based infrastructure to receive, handle and curate the samples from Mars. In an engineering sense, MSR consists of a linked set of missions, and a concluding set of ground-based activities, that we refer to as the MSR Campaign.

G. Kminek↗

Application of Nuclear Technology to Art Identification Problems: First Annual Report

Throughout the centuries, as man has become increasingly affluent, his interest in antiquities and objects of art, and the price he has been willing to pay for them, has increased. Inevitably, one result has been to attract the unscrupulous to the production and sale of counterfeit paintings, sculptures, and other works of art. The skill of forgers varies but often is highly sophisticated. However, dealers, museum curators, and other experts have also become more skillful in the detection of forgeries. Scientific tools of increasing sensitivity and sophistication have gradually been applied to the examination of the materials of art and archaeology. Such tools frequently support and render more reliable the judgment of the experts. The quality of forgeries may improve further as their makers become acquainted with the new methods of detection and in turn learn to circumvent or confound the methods. Therefore, the development of still more advanced methods of examination has great utility in the art world. This is particularly true if the ultimate effect is to make circumvention of the methods of examination so costly that the economic incentives for producing forgeries may be substantially reduced, if not eliminated.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj↗

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES↗

LSKnowledge: Nexus for Transformative Scientific Discoveries and Enhanced Information Retrieval in NASA Life Sciences Portal

We stand at the brink of an extraordinary transformation in the field of AI, driven by the convergence of generative AI and semantic technologies (e.g., knowledge graphs). This fusion holds immense potential and could redefine the future of scientific exploration, particularly in the realm of life sciences research. In this context, we shed light on the pivotal roles that Large Language Models (LLMs) and semantic technologies will play in advancing research, unearthing and comprehending life sciences information through innovative approaches, and empowering researchers to extract insights from NASA's extensive Life Sciences Data Archive. Within the NASA Life Sciences Portal (NLSP), the integration of LLMs and semantic technologies unlocks several advanced capabilities. First and foremost, it equips scientists with sophisticated tools to manage the ever-expanding wealth of scientific literature and data. Furthermore, it facilitates the creation of knowledge graphs that visually represent intricate relationships among biological entities, enabling comprehensive systems-level analysis. Additionally, the fusion of generative AI (including LLMs) and semantic technology can significantly benefit NASA's life sciences research by enhancing information retrieval and hypothesis generation. These tools enhance natural language understanding, facilitating knowledge discovery within NLSP. The overarching vision is to establish a cohesive knowledge ecosystem within NLSP, harnessing the power of LLMs and semantic technologies to synthesize and cross-reference data from diverse missions, disciplines, and research domains. This holistic approach ultimately deepens our understanding of how space environments impact life sciences data. To advance this initiative, we have launched LSKnowledge, aimed at enhancing the information retrieval capabilities of NLSP. In the short term, our primary goal is to develop a robust semantic search system. This system will empower HRP (Human Research Program) researchers to navigate NLSP data repositories more efficiently and precisely, catalyzing the process of hypothesis formation and scientific breakthroughs. To achieve this, we have employed pre-trained LLMs as part of a semantic search tool that can rank and highlight the most relevant records for user queries. To assess the tool's performance, we have curated a set of approximately 200 queries from subject matter experts (SMEs) and manually ranked the top records retrieved by both the current search system and the new semantic search, using SME judgments as the gold standard for relevancy. Herein, we present the results of our comparative analysis and illustrate how these findings have informed the fine-tuning of the system for enhanced performance. In the long term, our objectives include 1) retrieving publicly available information and integrating it with NLSP data to provide more precise answers to user queries, and 2) incorporating non-textual information from the NLSP database into our approach. In conclusion, the fusion of LLMs and semantic technologies within NLSP represents a pioneering stride towards reshaping the landscape of scientific discovery. This synergy not only equips researchers with powerful tools to navigate the burgeoning sea of information but also facilitates a deeper understanding of complex biological relationships, all while accelerating hypothesis generation and knowledge discovery. Through our initiative, LSKnowledge, we are committed to continually refining and expanding these capabilities, with the aim of not only enhancing information retrieval but also integrating diverse data sources to provide more precise insights. In the grand vision, NLSP strives to become the cornerstone of a comprehensive knowledge ecosystem, unraveling the enigmatic intricacies of life sciences phenomena in the context of space environments.

Life Sciences↗

Advancing the Prediction of MS/MS Spectra Using Machine Learning

Tandem mass spectrometry (MS/MS) is an important tool for the identification of small molecules and metabolites where resultant spectra are most commonly identified by matching them with spectra in MS/MS reference libraries. While popular, this strategy is limited by the contents of existing reference libraries. In response to this limitation, various methods are being developed for the in silico generation of spectra to augment existing libraries. Recently, machine learning and deep learning techniques have been applied to predict spectra with greater speed and accuracy. Here, in this work, we investigate the challenges these algorithms face in achieving fast and accurate predictions on a wide range of small molecules. The challenges are often amplified by the use of generic machine learning benchmarking tactics, which lead to misleading accuracy scores. Curating data sets, only predicting spectra for sufficiently high collision energies, and working more closely with experimental mass spectrometrists are recommended strategies to improve overall prediction accuracy in this nuanced field.

47 OTHER INSTRUMENTATION↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Stardust Curation at Johnson Space Center: Photo Documentation and Sample Processing of Submicron Dust Samples from Comet Wild 2 for Meteoritics Science Community

Dust particles released from comet 81P/Wild-2 were captured in silica aerogel on-board the STARDUST spacecraft and successfully returned to the Earth on January 15, 2006. STARDUST recovered thousands of particles ranging in size from 1 to 100 micrometers. The analysis of these samples is complicated by the small total mass collected ( < 1mg), its entrainment in the aerogel collection medium, and the fact that the cometary dust is comprised of submicrometer minerals and carbonaceous material. During the six month Preliminary Examination period, 75 tracks were extracted from the aerogel cells , but only 25 cometary residues were comprehensively studied by an international consortium of 180 scientists who investigated their mineralogy/petrology, organic/inorganic chemistry, optical properties and isotopic compositions. These detailed studies were made possible by sophisticated sample preparation methods developed for the STARDUST mission and by recent major advances in the sensitivity and spatial resolution of analytical instruments.

Nakamura-Messenger, K.↗

GLORIA - A Globally Representative Hyperspectral In Situ Dataset for Optical Sensing of Water Quality

The development of algorithms for remote sensing of water quality (RSWQ) requires a large amount of in situ data to account for the bio-geo-optical diversity of inland and coastal waters. The GLObal Reflectance community dataset for Imaging and optical sensing of Aquatic environments (GLORIA) includes 7,572 curated hyperspectral remote sensing reflectance measurements at 1 nm intervals within the 350 to 900 nm wavelength range. In addition, at least one co-located water quality measurement of chlorophyll a , total suspended solids, absorption by dissolved substances, and Secchi depth, is provided. The data were contributed by researchers affiliated with 59 institutions worldwide and come from 450 different water bodies, making GLORIA the de-facto state of knowledge of in situ coastal and inland aquatic optical diversity. Each measurement is documented with comprehensive methodological details, allowing users to evaluate fitness-for-purpose, and providing a reference for practitioners planning similar measurements. We provide open and free access to this dataset with the goal of enabling scientific and technological advancement towards operational regional and global RSWQ monitoring.

remote sensing of water quality↗

Extending the Reach of IGSN Beyond Earth: Implementing IGSN Registration to Link Nasa's Apollo Lunar Samples and Their Data

The rock and soil samples returned from the Apollo missions from 1969-72 have supported 46 years of research leading to advances in our understanding of the formation and evolution of the inner Solar System. NASA has been engaged in several initiatives that aim to restore, digitize, and make available to the public existing published and unpublished research data for the Apollo samples. One of these initiatives is a collaboration with IEDA (Interdisciplinary Earth Data Alliance) to develop MoonDB, a lunar geochemical database modeled after PetDB (Petrological Database of the Ocean Floor). In support of this initiative, NASA has adopted the use of IGSN (International Geo Sample Number) to generate persistent, unique identifiers for lunar samples that scientists can use when publishing research data. To facilitate the IGSN registration of the original 2,200 samples and over 120,000 subdivided samples, NASA has developed an application that retrieves sample metadata from the Lunar Curation Database and uses the SESAR API to automate the generation of IGSNs and registration of samples into SESAR (System for Earth Sample Registration). This presentation will describe the work done by NASA to map existing sample metadata to the IGSN metadata and integrate the IGSN registration process into the sample curation workflow, the lessons learned from this effort, and how this work can be extended in the future to help deal with the registration of large numbers of samples.

Todd, Nancy S.↗

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE↗

A Curated Dataset of Regional Meteor Events with Simultaneous Optical and Infrasound Observations (2006–2011)

We present a curated, openly accessible dataset of 71 regional meteor events simultaneously recorded by optical and infrasound instrumentation between 2006 and 2011. These events were captured during an observational campaign using the all-sky cameras of the Southern Ontario Meteor Network and the co-located Elginfield Infrasound Array. Each entry provides optical trajectory measurements, infrasound waveforms, and atmospheric specification profiles. The integration of optical and acoustic data enables robust linkage between observed acoustic signals and specific points along meteor trajectories, offering new opportunities to examine shock wave generation, propagation, and energy deposition processes. This release fills a critical observational gap by providing the first validated, openly accessible archive of simultaneous optical–infrasound meteor observations that supports trajectory reconstruction, acoustic propagation modeling, and energy deposition analyses. By making these data openly available in a structured format, this work establishes a durable reference resource that advances reproducibility, fosters cross-disciplinary research, and underpins future developments in meteor physics, atmospheric acoustics, and planetary defense.

astrometry↗