-
-
Notifications
You must be signed in to change notification settings - Fork 2
Quick Start
This guide provides a concise introduction to CmpStr. It covers the essential steps required to install the library, understand its core concepts, and perform common tasks such as string comparison, normalization, and phonetic search.
The examples are intentionally minimal and practical, allowing you to get productive within a few minutes. No prior knowledge of the internal architecture is required; the focus is on applying the API in real-world scenarios.
For detailed and complete documentation, refer to the API Reference.
Install CmpStr in your project using npm:
npm install cmpstrAlternatively, you can load the prebuilt browser bundles via jsDelivr or directly from the dist/ directory. For a full overview of supported environments and module formats (ESM, CommonJS, UMD), see Installation & Setup.
The primary entry point of the library is the CmpStr class. To perform a basic similarity comparison, create an instance using the .create() factory method and configure it with a similarity metric (for example, levenshtein) and optional normalization flags, such as i for case-insensitive matching.
import { CmpStr } from 'cmpstr';
const cmp = CmpStr.create( {
metric: 'levenshtein', // use Levenshtein distance
flags: 'i' // case-insensitive
} );
console.log( cmp.test( 'hello', 'Hallo' ) );
// { source: 'hello', target: 'Hallo', match: 0.8 }
console.log( cmp.batchTest( [ 'hello', 'hola' ], 'Hallo', { raw: true } ) );
/**
* [
* { metric: 'levenshtein', a: 'hello', b: 'Hallo', res: 0.8, raw: { dist: 1, maxLen: 5 } },
* { metric: 'levenshtein', a: 'hola', b: 'Hallo', res: 0.4, raw: { dist: 3, maxLen: 5 } }
* ]
*/
console.log( cmp.compare( 'hello', 'Hallo' ) );
// 0.8In this example, input strings are normalized using case-insensitive preprocessing and then compared using .test() for a single comparison and .batchTest() for multiple inputs. The .compare() method returns only the normalized similarity score in the range 0..1.
For a complete overview of available methods and options, see the API Reference.
For larger datasets or I/O-heavy workloads (for example, API backends or database-driven search), CmpStr provides a fully asynchronous API:
import { CmpStrAsync } from 'cmpstr';
const cmp = CmpStrAsync.create().setProcessors( {
phonetic: { algo: 'soundex' }
} );
const result = await cmp.searchAsync( 'Maier', [
'Meyer', 'Müller', 'Miller', 'Meyers', 'Meier'
] );
console.log( result );
// [ 'Meyer', 'Meier' ]The asynchronous variant mirrors the synchronous API but exposes await-able methods for scalable, non-blocking workflows. Available methods include .searchAsync(), .testAsync(), .batchTestAsync(), .phoneticIndexAsync(), and others.
For a complete reference, see the Asynchronous API.
CmpStr follows a modular and layered design that emphasizes clarity and flexibility in text comparison workflows. The following concepts are central to effective usage.
Before any comparison is performed, input strings typically pass through a normalization and filtering pipeline. This may include trimming whitespace, converting to lowercase, or removing diacritics. The pipeline is fully configurable, and custom filters can be registered to accommodate domain-specific requirements.
Learn more about this process in Normalization & Filtering.
CmpStr includes a variety of built-in similarity metrics, each suited to different use cases:
- Levenshtein – edit-distance based metric
- Dice-Sørensen – bigram-based similarity
- Jaro-Winkler – optimized for short strings and names
- and many more
All metrics return normalized scores between 0 (no similarity) and 1 (identical). If required, you can also register custom metrics.
For details, see Similarity Metrics.
CmpStr provides a built-in profiler for collecting performance and diagnostic information. This is useful in performance-sensitive environments such as search systems, data matching pipelines, or high-volume comparison tasks.
The profiler is memory-efficient and designed to introduce minimal runtime overhead. It collects timing data in milliseconds (ms) and memory usage in bytes. In environments where memory information is unavailable (such as some browsers), memory usage defaults to 0.
The profiler is disabled by default. To activate it:
CmpStr.profiler.enable();Diagnostic data can be accessed using the following methods:
CmpStr.profiler.report(); // full report (array of entries)
CmpStr.profiler.last(); // most recent entry
CmpStr.profiler.clear(); // clear profiler data
CmpStr.profiler.total(); // aggregated time and memory
// Example: { time: 0.38, mem: 25696 }Profiling remains active until explicitly disabled.
CmpStr also includes utilities for text analysis and difference inspection. While not part of the core comparison workflow, these modules can be useful without introducing external dependencies:
- TextAnalyzer – token statistics, frequency analysis, reading metrics, histograms, and more
- DiffChecker – line-based and word-based text diffs using a unified diff format
For details, see Diff & Text Analysis.
CmpStr 3.2 / API Reference • FAQ • npm Package • jsDelivr • DevDocs
CmpStr is a lightweight, fast and well performing package for calculating string similarity.
Getting Started
Installation & Setup
Quick Start
API Reference
Documentation
Similarity Metrics
Phonetic Algorithms
Normalization & Filtering
Comparison Modes
Structured Data
Asynchronous API
Diff & Text Analysis
Extending CmpStr
Project Management