Skip to content

feat: add OpenAI-compatible embedding provider - #111

Closed
sk7n4k3d wants to merge 2 commits into
qdrant:masterfrom
sk7n4k3d:feat/openai-embedding-provider
Closed

sk7n4k3d wants to merge 2 commits into
qdrant:masterfrom
sk7n4k3d:feat/openai-embedding-provider

Conversation

@sk7n4k3d

Copy link
Copy Markdown

Summary

Adds a new openai embedding provider that supports any OpenAI-compatible embedding API endpoint (/v1/embeddings). This enables using local inference servers, Azure OpenAI, LiteLLM, Ollama, vLLM, or any other service exposing the standard OpenAI Embeddings API.

Motivation

Currently mcp-server-qdrant only supports FastEmbed for embeddings. Many users already run dedicated embedding servers (for hardware acceleration, model flexibility, or centralized deployment) and would benefit from a generic provider that can call any OpenAI-compatible endpoint.

Configuration

EMBEDDING_PROVIDER=openai
EMBEDDING_MODEL=your-model-name
OPENAI_BASE_URL=http://your-server:8080/v1
OPENAI_API_KEY=optional-api-key  # defaults to "not-needed"

Changes

  • NEW embeddings/openai.py — OpenAIEmbeddingProvider class
  • embeddings/factory.py — Added OPENAI branch
  • embeddings/types.py — Added OPENAI enum value
  • settings.py — Added openai_base_url and openai_api_key fields

Design Decisions

  • Zero additional dependencies — uses only Python stdlib (urllib.request)
  • Async via run_in_executor — wraps synchronous urllib calls to avoid blocking the event loop
  • Auto-detect vector dimensions — probes the API on first call
  • Namespaced vector names — uses openai-{model-slug} prefix to avoid collisions with FastEmbed vectors

@castlenthesky

castlenthesky commented Sep 5, 2026 •

Copy link
Copy Markdown

@sk7n4k3d, heads up that I've opened #187 as a draft doing the same thing, so a duplicate doesn't catch you by surprise.

#111 has gone stale against master and hasn't had a review since February. My read is that the blocker on all of these is a mix of review bandwidth and scope. In #55 @kacperlukawski asked for a PR with "just OpenAI embeddings, but nothing else," and the 417 deleted lines here are probably more than he wants to work through in one pass.

Mine is additive only for that reason. Not trying to get ahead of you though. If you'd rather rebase #111 and slim it down, I'll close mine and back yours instead.

@sk7n4k3d

Copy link
Copy Markdown
Author

Thanks for the heads-up and for writing this up properly — the #55 quote is the context I was missing, and you're right on both counts.

#111 is stale against master (mergeable_state is dirty), it's +135/-417 when the maintainer asked for "just OpenAI embeddings, but nothing else", and it's been sitting since February. I'd rather this feature land than my branch land.

Go ahead with #187. I'm closing #111.

Two notes that might save you a round-trip:

  1. I went with stdlib urllib.request to keep dependencies at zero. Your approach — depending on the official openai client — is almost certainly what the maintainer will prefer, since it matches how FastEmbed is already wired. Don't let the "no new deps" angle pull you back toward hand-rolled HTTP.
  2. EMBEDDING_BASE_URL covering Ollama/OpenRouter/Gemini in one provider is the right call. Splitting those into three classes is what fragmented this feature across five PRs in the first place.

Happy to be a second pair of eyes on the tests or the settings table if that helps. Good luck with the review.

@sk7n4k3d sk7n4k3d closed this Sep 16, 2026
@sk7n4k3d

Copy link
Copy Markdown
Author

Closed in favour of #187 — @castlenthesky's version is additive, matches the #55 brief, and isn't stale. Piling a rebase on top of four existing PRs wasn't going to get the feature merged faster.

For anyone landing here looking for an OpenAI-compatible embedding provider: follow #187. It covers OpenAI, Ollama, vLLM, LM Studio and OpenRouter via EMBEDDING_BASE_URL.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants