Search NASASearch

SEARCH · Search NASA

Results for “text extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

The mass profiles of dwarf galaxies from Dark Energy Survey lensing

We present a novel approach to extracting dwarf galaxies from photometric data to measure their average halo mass profile with weak lensing. We characterize their stellar mass and redshift distributions with a spectroscopic calibration sample. By combining the ${\sim} 5000\,\mathrm{deg}^2$ multiband photometry from the Dark Energy Survey and redshifts from the Satellites Around Galactic Analogs Survey with an unsupervised machine learning method, we select a low-mass galaxy sample spanning redshifts $z\lt 0.3$ and divide it into three mass bins. From low to high median mass, the bins contain [146 420, 330 146, 275 028] galaxies and have median stellar masses of $\log _{10}(M_*/\text{M}_\odot)=\left[8.52\substack{+0.57 -0.76},\, 9.02\substack{+0.50 -0.64},\, 9.49\substack{+0.50 -0.58}\right]$ . We measure the stacked excess surface mass density profiles, $\Delta \Sigma (R)$, of these galaxies using galaxy–galaxy lensing with a signal-to-noise ratio of [14, 23, 28]. Through a simulation-based forward-modelling approach, we fit the measurements to constrain the stellar-to-halo mass relation and find the median halo mass of these samples to be $\log _{10}(M_{\rm halo}/\text{M}_\odot)$ = [$10.67\substack{+0.2 -0.4}$, $11.01\substack{+0.14 -0.27}$, $11.40\substack{+0.08 -0.15}$]. The cold dark matter profiles are consistent with NFW (Navarro, Frenk, and White) profiles over scales ${\lesssim} 0.15 \, {h}^{-1}$ Mpc. We find that ${\sim} 20$ per cent of the dwarf galaxy sample are satellites. This is the first measurement of the halo profiles and masses of such a comprehensive, low-mass galaxy sample. The techniques presented here pave the way for extracting and analysing even lower mass dwarf galaxies and for more finely splitting galaxies by their properties with future photometric and spectroscopic survey data.

dark matter

Verification of RESRAD-OFFSITE Code (V.4)

This report documents the verification of RESRAD-OFFSITE Version 4.0 and describes, where necessary, the verification of the following: • The data comprising the standard dose and risk coefficient libraries in the RESRAD database files Master_dcf_ICRP07.mdb and Master_dcf_2k.mdb. • The extraction and transfer of the data from the selected database file to the computational code by the RESRAD-OFFSITE 4.0 interface, ResOWin.exe. • The different processes that are modeled by the main computational code in RESRAD OFFSITE 4.0, ResOMain.exe. • The data displayed in the graphical and text reports. Many verifications were performed as part of the quality assurance quality control program associated with the development and release of RESRAD-OFFSITE 4.0, namely: • developer testing, • internal independent testing, and • release testing. Some were also performed in response to questions from users regarding the performance of the code. The main text of the report focuses on summarizing a subset of those tests, both independent and developer tests that verified the computations performed by the code. The verifications included in this report served as the basis for the development of the release tests of the computational executables and provided the quantitative results to be compared with the code output. The input and output interfaces and the data transfers between the various executables of the code were tested while performing the verification testing. They were tested intentionally during release testing. This report also provides some basic information to help in understanding the activities that were verified. The report: • outlines the components of RESRAD-OFFSITE 4.0 and the interconnections between these components, • outlines the processes modeled by the computational code, • provides summary figures and tables to offer confirmation of the verification of the computational components of the code, • reproduces the verifiers’ reports, if available, in individual appendices, • refers to the previous verification report (Yu et al. 2011) for more details about some of the verifications, and • reproduces the test cases and the testers’ reports from the release testing in individual appendices, when possible.

54 ENVIRONMENTAL SCIENCES

Orbitrap LC-MS Analysis of Nanoparticle Composition at the EPCAPE Mount Soledad site between 04 18 2023 and 06 14 2023

Weekly peak lists containing m/z, intensity, and assigned formula for filter samples, size selected for sub-100 nm particles. Filters were collected daily between 4/18/23 and 6/14/23, grouped based on calendar week for extraction, and analyzed via Thermo Scientific Q Exactive Plus Orbitrap LC-MS. Formulas were assigned to background-corrected peak lists and restricted to CHONS/CHONSNa atoms for the negative and positive modes respectively. Filters were grouped into calendar weeks 0-8 with dates provided in README text file.

54 ENVIRONMENTAL SCIENCES

Rotorcraft Optimization Tools: Incorporating Rotorcraft Design Codes into Multi-Disciplinary Design, Analysis, and Optimization

One of the goals of NASA's Revolutionary Vertical Lift Technology Project (RVLT) is to provide validated tools for multidisciplinary design, analysis and optimization (MDAO) of vertical lift vehicles. As part of this effort, the software package, RotorCraft Optimization Tools (RCOTOOLS), is being developed to facilitate incorporating key rotorcraft conceptual design codes into optimizations using the OpenMDAO multi-disciplinary optimization framework written in Python. RCOTOOLS, also written in Python, currently supports the incorporation of the NASA Design and Analysis of RotorCraft (NDARC) vehicle sizing tool and the Comprehensive Analytical Model of Rotorcraft Aerodynamics and Dynamics II (CAMRAD II) analysis tool into OpenMDAO-driven optimizations. Both of these tools use detailed, file-based inputs and outputs, so RCOTOOLS provides software wrappers to update input files with new design variable values, execute these codes and then extract specific response variable values from the file outputs. These wrappers are designed to be flexible and easy to use. RCOTOOLS also provides several utilities to aid in optimization model development, including Graphical User Interface (GUI) tools for browsing input and output files in order to identify text strings that are used to identify specific variables as optimization input and response variables. This paper provides an overview of RCOTOOLS and its use

Analysi

NASA Indexing Benchmarks: Evaluating Text Search Engines

The current proliferation of on-line information resources underscores the requirement for the ability to index collections of information and search and retrieve them in a convenient manner. This study develops criteria for analytically comparing the index and search engines and presents results for a number of freely available search engines. A product of this research is a toolkit capable of automatically indexing, searching, and extracting performance statistics from each of the focused search engines. This toolkit is highly configurable and has the ability to run these benchmark tests against other engines as well. Results demonstrate that the tested search engines can be grouped into two levels. Level one engines are efficient on small to medium sized data collections, but show weaknesses when used for collections 100MB or larger. Level two search engines are recommended for data collections up to and beyond 100MB.

Esler, Sandra L.

Artificial Intelligence (AI) Methods for Automating the Impact Tool Evidence Library

INTRODUCTION: The development of the Evidence Library for use with the IMPACT (Informing Mission Planning via Analysis of Complex Tradespaces) probability risk assessment tool involved a multilayered, time intensive process of data collection and analysis by subject matter experts from the Exploration Medical Capability (ExMC) Element Clinical and Science Team to produce clinical findings forms (CliFFs) for 120 medical conditions. Artificial Intelligence Large Language Models (LLMs) can be leveraged to facilitate this process, thus reducing labor and time. TOPIC: CliFFs contain information about medical conditions as they pertain to spaceflight. This includes condition definitions, incidence data, crew task impairment estimates caused by conditions, treatment protocols and references to literature used for gathering condition evidence. Guided by the Evidence Library Methods document and the CliFF development instructions, a team has leveraged Microsoft Azure AI services and open-source documentation to construct an AI-assisted automated pipeline for CliFF development. This process is designed to search, retrieve, and evaluate the applicable data, and ultimately generate a completed CliFF. The LLM evaluates the relevance of each of the source materials to spaceflight, either as direct evidence or as an analog. The model extracts keywords and generates brief summaries to enhance search and retrieval in later stages of CliFF development. For instance, it can calculate epidemiological statistical data, such as incidence rates and the likelihood of best or worst-case scenarios. APPLICATION: Large Language Models (LLMs) can efficiently summarize large amounts of text. Leveraging this technology will automate data retrieval and evidence gathering for medical databases, like the IMPACT tool, by aiding in the labor-intensive process of analyzing large bodies of literature and organizing it into a formatted document like a CliFF. This added efficiency will enable expeditious expansion of the Evidence Library with additional medical conditions and update previous CLiFFs as new technology becomes available.

Ali Al

Knowledge Preservation for Design of Rocket Systems

An engineer at NASA Lewis RC presented a challenge to us at Southern University. Our response to that challenge, stated circa 1993, has evolved into the Knowledge Preservation Project which is here reported. The stated problem was to capture some of the knowledge of retiring NASA engineers and make it useful to younger engineers via computers. We evolved that initial challenge to this - design a system of tools such that, with this system, people might efficiently capture and make available via commonplace computers, deep knowledge of retiring NASA engineers. In the process of proving some of the concepts of this system, we would (and did) capture knowledge from some specific engineers and, so, meet the original challenge along the way to meeting the new. Some of the specific knowledge acquired, particularly that on the RL- 10 engine, was directly relevant to design of rocket engines. We considered and rejected some of the techniques popular in the days we began - specifically "expert systems" and "oral histories". We judged that these old methods had too high a cost per sentence preserved. That cost could be measured in hours of labor of a "knowledge professional". We did spend, particularly in the grant preceding this one, some time creating a couple of "concept maps", one of the latest ideas of the day, but judged this also to be costly in time of a specially trained knowledge-professional. We reasoned that the cost in specialized labor could be lowered if less time were spent being selective about sentences from the engineers and in crafting replacements for those sentences. The trade-off would seem to be that our set of sentences would be less dense in information, but we found a computer-based way around this seeming defect. Our plan, details of which we have been carrying out, was to find methods of extracting information from experts which would be capable of gaining cooperation, and interest, of senior engineers and using their time in a way they would find worthy (and, so, they would give more of their time and recruit time of other engineers as well). We studied these four ways of creating text: 1) the old way, via interviews and discussions - one of our team working with one expert, 2) a group-discussion led by one of the experts themselves and on a topic which inspires interaction of the experts, 3) a spoken dissertation by one expert practiced in giving talks, 4) expropriating, and modifying for our system, some existing reports (such as "oral histories" from the Smithsonian Institution).

Moreman, Douglas

Enhanced Reporting of Mars Exploration Rover Telemetry

Mars Exploration Rover Enhanced Telemetry Extraction and Reporting System (METERS) is software that generates a human-readable representation of the state of the mobility and arm-related systems of the Mars Exploration Rover (MER) vehicles on each Martian solar day (sol). Data are received from the MER spacecraft in multiple streams having various formats including text messages, sparsely-sampled engineering quantities, images, and individual motor-command histories.

Maimone, Mark W.

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING

Artificial Intelligence (AI) Methods for Augmenting the IMPACT Tool Evidence Library

Development of the Evidence Library for use with the IMPACT probability risk assessment tool took several years and involved a staggering amount of effort from a multi-disciplinary team. A very significant amount of the labor effort to collect, assess and finalize the Clinical Finding Form (CliFF) for each of the 119 medical conditions was provided by physician subject matter experts from the Exploration Medical Capability (ExMC) Element Clinical and Science Team. Many AI tools such as ChatGPT are excellent at summarizing large amounts of information and the current project was initiated to determine how such tools might streamline laborious processes, e.g., review and summarization of many scientific research publications, to execute key steps more efficiently in the process of developing CliFFs. The process for collecting the evidence which is found in the CliFFs is well documented in the Evidence Library Methods document (ELM; HRP-48036*). Using ELM and the CliFF development instructions as a guideline, a team of developers is leveraging Microsoft Azure AI tools and services along with open-source frameworks, to construct an AI-assisted automated pipeline. This pipeline is designed to search, retrieve, and process the necessary data sources, and ultimately help generate the final version of a CliFF. Currently, the large language model evaluates the relevance of each source material to spaceflights, either as direct evidence or as an analog. Additionally, the model assists in extracting keywords and generating brief summaries to enhance augmented retrieval and search processes in later stages of CliFF development. Once the data is ready, the model can perform semantic search and retrieval, generating and extracting valuable information for the CliFF. For instance, it can handle epidemiological statistical data, such as incidence rates and the likelihood of best or worst-case scenarios. The steps that required reading and summarizing articles were viewed as providing the greatest return on investment since large language models are very efficient and accurate in summarizing large amounts of text. Since labor effort to complete the original CliFF was not recorded with sufficient granularity, comparisons with an AI tool-generated CliFF will provide merely an approximation of time saved. Upon completion of the process, the CliFF for the medical condition “appendicitis” generated with the support of AI-based methods will serve as a proof-of-concept and will be compared to the original appendicitis CliFF to determine if use of the tools resulted in content and conclusory similarity. Based upon the results from face validation of the two CliFFs, modifications to the process will be made if necessary and additional condition CliFFs will be evaluated. Ultimately, CliFFs for the entire set of medical conditions will be created with the assistance of AI tools. Depending on the cost savings realized, CliFFs for additional medical conditions can be created to expand the Evidence Library. Future direction includes specifying the characteristics of the reviewer (prompting the AI tools to generate output assuming the reviewer is a sub-specialist physician, or nurse or EMT/medic) to determine if the effects on AI-generated output are different based on knowledge, skills and abilities. *Exploration Medical Capability Evidence Library Methods, HRP-48036 Rev A, July 2022.

Ali Al

Forensic Analysis of Compromised Computers

Directory Tree Analysis File Generator is a Practical Extraction and Reporting Language (PERL) script that simplifies and automates the collection of information for forensic analysis of compromised computer systems. During such an analysis, it is sometimes necessary to collect and analyze information about files on a specific directory tree. Directory Tree Analysis File Generator collects information of this type (except information about directories) and writes it to a text file. In particular, the script asks the user for the root of the directory tree to be processed, the name of the output file, and the number of subtree levels to process. The script then processes the directory tree and puts out the aforementioned text file. The format of the text file is designed to enable the submission of the file as input to a spreadsheet program, wherein the forensic analysis is performed. The analysis usually consists of sorting files and examination of such characteristics of files as ownership, time of creation, and time of most recent access, all of which characteristics are among the data included in the text file.

Wolfe, Thomas

Disorder-induced magnetoelastic behaviors of MnTexSbyBi1-x-y alloys

This dataset contains input and output files from density functional theory (DFT) simulations used to study the disorder-induced magnetoelastic behaviors of MnTexSbyBi1-x-y (0 ≤ x + y ≤ 1) alloys and their binary end members MnTe, MnSb, and MnBi. The alloys adopt the hexagonal NiAs-type (nickeline) structure and span ternary (MnTexSb1-x, MnTexBi1-x, MnBixSb1-x), and quaternary compositions across the full MnTe–MnSb–MnBi composition triangle. For each alloy composition, the dataset provides DFT calculations in three magnetic configurations: A-type antiferromagnetic (AFM), C-type AFM, and ferromagnetic (FM). Every magnetic configuration folder contains the fully relaxed crystal structure (CONTCAR), VASP input parameters (INCAR), and the main VASP output file (OUTCAR), from which total electronic energies, Mn magnetic moments, lattice parameters, and percent volume changes between magnetic states are extracted. These data are used to construct compositional phase diagrams, evaluate thermodynamic stability (formability), and map magnetoelastic responses across the alloy space. For A-type AFM and FM configurations, additional data are provided as follows: (i) FORCE_CONSTANTS and thermal_properties.yaml files at the top level of A-type_AFM/ and FM/ folders — present only for compositions marked with an asterisk (*) in Table I of the main text. These are derived from Phonopy finite-displacement calculations on full disordered 128-atom supercells and provide vibrational free energy, entropy (Svib)contribution from explicit disorder calculations. (Table I of the associated main manuscript) (ii) A VCA/ subfolder within A-type_AFM/ and FM/, containing FORCE_CONSTANTS and thermal_properties.yaml from Virtual Crystal Approximation phonon calculations (without spin-orbit coupling). VCA data are available for all compositions and are used to estimate vibrational contributions to the Gibbs free energy across the full composition space. (iii) A SOC/ subfolder containing CONTCAR, INCAR, and OUTCAR from spin-orbit coupling calculations, providing relativistic corrections to electronic energies and lattice parameters (Tables S2–S3 of the SM, and Table I of the main manuscript). (iv) A SOC/VCA/ subfolder containing FORCE_CONSTANTS and thermal_properties.yaml from VCA phonon calculations performed within the SOC framework, combining relativistic and vibrational thermodynamic corrections. The computed properties are used to map the AFM–FM magnetic crossover near MnTe0.75Sb0.25, demonstrate disorder- and spin-induced phonon broadening, identify a semiconductor-to-metal crossover, and quantify the pronounced magnetoelastic volume response near the magnetic phase boundary.

36 MATERIALS SCIENCE

Leveraging Large Language Models for Understanding Fundamental Principles of Catalysis

Heterogeneous catalysis presents a distinct challenge for artificial intelligence (AI). Data sets are often small and inconsistently reported, catalyst representations are not standardized, and extracting fundamental knowledge requires integrating performance data, spectroscopic characterizations, and mechanistic models across multiple scales. Language offers a unifying representation across these modalities, making catalysis well suited for leveraging large language models (LLMs). By standardizing how catalytic data is represented, LLMs make dispersed experimental results more accessible to downstream statistical modeling. In this perspective, we focus our discussion around three opportunities where LLMs can significantly contribute to catalysis: (1) text to properties; (2) text to structure; and (3) text to mechanistic models. The discussion is followed by a perspective section on LLM-readiness of data, aligning LLM outputs with scientific correctness, and bridging lab-scale discovery to industrial deployment. Across each area, the most productive applications couple dispersed chemical knowledge with physics-grounded validation to produce verifiable hypotheses and actionable representations.

Catalysts

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john