Skip to content

Quick Start

Paul Köhler edited this page Jan 22, 2026 · 10 revisions

This guide provides a concise introduction to CmpStr. It covers the essential steps required to install the library, understand its core concepts, and perform common tasks such as string comparison, normalization, and phonetic search.

The examples are intentionally minimal and practical, allowing you to get productive within a few minutes. No prior knowledge of the internal architecture is required; the focus is on applying the API in real-world scenarios.

For detailed and complete documentation, refer to the API Reference.

Installation

Install CmpStr in your project using npm:

npm install cmpstr

Alternatively, you can load the prebuilt browser bundles via jsDelivr or directly from the dist/ directory. For a full overview of supported environments and module formats (ESM, CommonJS, UMD), see Installation & Setup.

Basic Usage

The primary entry point of the library is the CmpStr class. To perform a basic similarity comparison, create an instance using the .create() factory method and configure it with a similarity metric (for example, levenshtein) and optional normalization flags, such as i for case-insensitive matching.

Minimal Synchronous Usage

import { CmpStr } from 'cmpstr';

const cmp = CmpStr.create( {
  metric: 'levenshtein', // use Levenshtein distance
  flags: 'i' // case-insensitive
} );

console.log( cmp.test( 'hello', 'Hallo' ) );
// { source: 'hello', target: 'Hallo', match: 0.8 }

console.log( cmp.batchTest( [ 'hello', 'hola' ], 'Hallo', { raw: true } ) );
/**
 * [
 *   { metric: 'levenshtein', a: 'hello', b: 'Hallo', res: 0.8, raw: { dist: 1, maxLen: 5 } },
 *   { metric: 'levenshtein', a: 'hola', b: 'Hallo', res: 0.4, raw: { dist: 3, maxLen: 5 } }
 * ]
 */

console.log( cmp.compare( 'hello', 'Hallo' ) );
// 0.8

In this example, input strings are normalized using case-insensitive preprocessing and then compared using .test() for a single comparison and .batchTest() for multiple inputs. The .compare() method returns only the normalized similarity score in the range 0..1.

For a complete overview of available methods and options, see the API Reference.

Asynchronous Usage

For larger datasets or I/O-heavy workloads (for example, API backends or database-driven search), CmpStr provides a fully asynchronous API:

import { CmpStrAsync } from 'cmpstr';

const cmp = CmpStrAsync.create().setProcessors( {
  phonetic: { algo: 'soundex' }
} );

const result = await cmp.searchAsync( 'Maier', [
  'Meyer', 'Müller', 'Miller', 'Meyers', 'Meier'
] );

console.log( result );
// [ 'Meyer', 'Meier' ]

The asynchronous variant mirrors the synchronous API but exposes await-able methods for scalable, non-blocking workflows. Available methods include .searchAsync(), .testAsync(), .batchTestAsync(), .phoneticIndexAsync(), and others.

For a complete reference, see the Asynchronous API.

Core Concepts & Usage Tips

CmpStr follows a modular and layered design that emphasizes clarity and flexibility in text comparison workflows. The following concepts are central to effective usage.

Normalization & Filtering (Input Preprocessing)

Before any comparison is performed, input strings typically pass through a normalization and filtering pipeline. This may include trimming whitespace, converting to lowercase, or removing diacritics. The pipeline is fully configurable, and custom filters can be registered to accommodate domain-specific requirements.

Learn more about this process in Normalization & Filtering.

Built-in Algorithms

CmpStr includes a variety of built-in similarity metrics, each suited to different use cases:

All metrics return normalized scores between 0 (no similarity) and 1 (identical). If required, you can also register custom metrics.

For details, see Similarity Metrics.

Profiling & Diagnostics

CmpStr provides a built-in profiler for collecting performance and diagnostic information. This is useful in performance-sensitive environments such as search systems, data matching pipelines, or high-volume comparison tasks.

The profiler is memory-efficient and designed to introduce minimal runtime overhead. It collects timing data in milliseconds (ms) and memory usage in bytes. In environments where memory information is unavailable (such as some browsers), memory usage defaults to 0.

The profiler is disabled by default. To activate it:

CmpStr.profiler.enable();

Diagnostic data can be accessed using the following methods:

CmpStr.profiler.report(); // full report (array of entries)
CmpStr.profiler.last();   // most recent entry
CmpStr.profiler.clear();  // clear profiler data
CmpStr.profiler.total();  // aggregated time and memory
// Example: { time: 0.38, mem: 25696 }

Profiling remains active until explicitly disabled.

Diffing & Text Analysis

CmpStr also includes utilities for text analysis and difference inspection. While not part of the core comparison workflow, these modules can be useful without introducing external dependencies:

  • TextAnalyzer – token statistics, frequency analysis, reading metrics, histograms, and more
  • DiffChecker – line-based and word-based text diffs using a unified diff format

For details, see Diff & Text Analysis.

Clone this wiki locally