Comparative Evaluation of String-Similarity Algorithms for Keyword Classification in Mobile Library OCR Systems
DOI:
https://doi.org/10.52158/javict.v1i3.1653Keywords:
Fuzzy String Matching, Jaro-Winkler,, Levenshtein Distance,, TF-IDF, Keyword ClassificationAbstract
Automated keyword classification is a critical step in OCR-based library metadata systems, where OCR-extracted synopsis terms must ultimately be mapped to standardized subject keywords despite character-level recognition noise. While prior mobile OCR pipelines for library metadata extraction have adopted Levenshtein-distance-based fuzzy matching for this task, no systematic comparison exists to justify this choice against alternative string-similarity algorithms. This study isolates the keyword-classification sub-task and empirically compares three algorithms — Levenshtein-based similarity ratio, Jaro-Winkler similarity, and TF-IDF with cosine similarity — using noisy keyword-level queries as a controlled proxy for OCR-corrupted terms, under simulated OCR-noise conditions (clean, light, and heavy character-level corruption). Evaluation uses a curated 20-keyword library subject database and 900 test queries per algorithm across five random seeds. Jaro-Winkler achieves the highest overall classification accuracy (97.8%) and the fastest average matching time (22.68 microseconds per query), compared with Levenshtein-based matching (96.4% accuracy, 34.45 microseconds) and TF-IDF with cosine similarity (74.2% accuracy, 581.01 microseconds). A paired t-test across seeds confirmed that Jaro-Winkler's accuracy advantage over TF-IDF was statistically significant, while its advantage over Levenshtein was not and instead rests on its markedly lower and more stable matching latency. The gap against TF-IDF widened sharply under heavy noise (95.3% vs. 34.0% accuracy), showing that TF-IDF-based matching, despite its popularity for document-level text similarity, is poorly suited to short-keyword matching under character-level noise. These findings, based on a 20-keyword benchmark, offer preliminary algorithm-selection guidance for developers of OCR-based metadata classification pipelines, pending validation on larger, real-world OCR-noise data.