Comparative Evaluation of String-Similarity Algorithms for Keyword Classification in Mobile Library OCR Systems

Authors

DOI:

https://doi.org/10.52158/javict.v1i3.1653

Keywords:

Fuzzy String Matching, Jaro-Winkler,, Levenshtein Distance,, TF-IDF, Keyword Classification

Abstract

Automated keyword classification is a critical step in OCR-based library metadata systems, where OCR-extracted synopsis terms must ultimately be mapped to standardized subject keywords despite character-level recognition noise. While prior mobile OCR pipelines for library metadata extraction have adopted Levenshtein-distance-based fuzzy matching for this task, no systematic comparison exists to justify this choice against alternative string-similarity algorithms. This study isolates the keyword-classification sub-task and empirically compares three algorithms — Levenshtein-based similarity ratio, Jaro-Winkler similarity, and TF-IDF with cosine similarity — using noisy keyword-level queries as a controlled proxy for OCR-corrupted terms, under simulated OCR-noise conditions (clean, light, and heavy character-level corruption). Evaluation uses a curated 20-keyword library subject database and 900 test queries per algorithm across five random seeds. Jaro-Winkler achieves the highest overall classification accuracy (97.8%) and the fastest average matching time (22.68 microseconds per query), compared with Levenshtein-based matching (96.4% accuracy, 34.45 microseconds) and TF-IDF with cosine similarity (74.2% accuracy, 581.01 microseconds). A paired t-test across seeds confirmed that Jaro-Winkler's accuracy advantage over TF-IDF was statistically significant, while its advantage over Levenshtein was not and instead rests on its markedly lower and more stable matching latency. The gap against TF-IDF widened sharply under heavy noise (95.3% vs. 34.0% accuracy), showing that TF-IDF-based matching, despite its popularity for document-level text similarity, is poorly suited to short-keyword matching under character-level noise. These findings, based on a 20-keyword benchmark, offer preliminary algorithm-selection guidance for developers of OCR-based metadata classification pipelines, pending validation on larger, real-world OCR-noise data.

Author Biographies

  • Damar Galih AJi Pradana, Politeknik Elektronika Negeri Surabaya (PENS)

    Department of Informatics and Computer Engineering

  • Sritrusta Sukaridhoto, PoliteknikElektronika Negeri Surabaya (PENS)

    Department of Creative Multimedia Technology

  • Nirwana Haidar Hari, Politeknik Elektronika Negeri Surabaya (PENS)

    Department of Informatics and Computer Engineering

Downloads

Published

2026-09-30

How to Cite

Pradana, D. G. A., Sukaridhoto, S., & Hari, N. H. (2026). Comparative Evaluation of String-Similarity Algorithms for Keyword Classification in Mobile Library OCR Systems. Journal of Advanced Vocational Information and Communication Technology, 1(3), 14-28. https://doi.org/10.52158/javict.v1i3.1653

Similar Articles

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)