Search NASA⌕ Search

DOE OSTI · code-80722

Name Normalizer

Abstract

Our name normalizer and enrichment micro-service address the shortcomings of standard text normalization. Initially the text input is split into tokens. This process takes into account some conventions of formatting. Any dashes present between two tokens preserves the relationship of those two tokens. Conversely a comma between two tokens ensures that the separation between the two tokens is maintained. The order of tokens is also preserved. Standard text normalization (conversion to lowercase, trimming extra white space, and canonicalization) is then applied to the tokens. Once normalized every unique and sequence of tokens is given confidence values by comparing the normalized values to publicly available data.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bleeker, Amelia, Larimer, Curtis, Avila, Andrew, Chin, Jr, George. 2022-09-07. Name Normalizer. https://doi.org/10.11578/dc.20240614.240

Cite the original work for its findings. Save a collection to share your selection of sources.