Skip to content

feat: add OpenAI-compatible embedding provider - #187

Draft
castlenthesky wants to merge 1 commit into
qdrant:masterfrom
castlenthesky:feat/openai-embeddings-provider
Draft

castlenthesky wants to merge 1 commit into
qdrant:masterfrom
castlenthesky:feat/openai-embeddings-provider

Conversation

@castlenthesky

@castlenthesky castlenthesky commented Sep 5, 2026 •

Copy link
Copy Markdown

Adds EMBEDDING_PROVIDER=openai, so you can embed through anything that speaks the OpenAI /v1/embeddings API.

I've left it as a draft on purpose. There are already three open PRs doing roughly this, and I'd rather help get one of them finished than pile on a fourth. More on that below.

What it's for

Two open issues asking for exactly this:

What's been tried already

Still open, none of them reviewed:

PR Author Size
#92 @chaserhkj +1074/-900, 6 files
#111 @sk7n4k3d +135/-417, 5 files
#158 @HosseinShahabadi +178/-144, 5 files
#118 @pablomichelettii OpenRouter provider
#114 @donbowman Gemini provider

Closed:

Why I think they stalled

Two different things going on here, and only one of them is about the code.

#55 was too big, and the maintainers said so. From @kacperlukawski in that thread:

It would be great to add support for OpenAI embeddings, but this PR modified too many server parts. Happy to review another PR that will introduce just OpenAI embeddings, but nothing else.

That's a clear brief and as far as I can tell nobody has matched it since. All three of the open OpenAI PRs bundle refactors or unrelated deletions in with the feature, and all three have gone stale against master. The files they touch are the ones that keep moving.

Everything after that just looks like review bandwidth. #143 was small, mergeable and green, and the author pinged four times over two months before giving up and closing it himself. No amount of rewriting fixes that one.

There's a fragmentation problem too. OpenAI, OpenRouter, Ollama and Gemini all showed up as separate provider classes, but three of those four speak the same wire format. Reviewing them one at a time is a lot of work for not much payoff.

What I did differently

  • Nothing gets refactored. No existing source line is deleted or rewritten. The change to existing files is 26 lines: one enum value, one elif in the factory, six settings fields, one dependency.
  • One provider, lots of backends. Because you point EMBEDDING_BASE_URL wherever you like, OpenAI, Ollama, vLLM, LM Studio and OpenRouter all become config instead of new code. That should cover feat: add OpenRouter embedding provider support #118 and both Ollama PRs.
  • FastEmbed doesn't move. Still the default, still behaves the same. Nobody's existing setup changes.
  • It's small enough to review in one go, which seems to be the actual bottleneck here.

The diff

File What changed
embeddings/openai.py new provider, 78 lines
embeddings/types.py one enum value
embeddings/factory.py one elif branch
settings.py six optional fields
pyproject.toml adds openai
tests/ 142 lines, all mocked
README.md env var table plus an example

Three bits I'd flag for whoever reviews it:

  1. EMBEDDING_API_KEY falls back to OPENAI_API_KEY and then to a placeholder. Local servers don't check the key, but AsyncOpenAI won't construct without one, so a keyless endpoint would otherwise blow up with a confusing credentials error.
  2. EMBEDDING_VECTOR_SIZE gets detected by embedding a short probe string if you don't set it. There's no reliable way to read dimensionality off /v1/models for arbitrary backends. Set it explicitly and the probe is skipped.
  3. EMBEDDING_QUERY_PREFIX and EMBEDDING_DOCUMENT_PREFIX are there for asymmetric models like Qwen3-Embedding, E5 and BGE, which want an instruction prefix on queries but not on documents. Both default to empty, so they're invisible if you don't need them.

Testing

  • 34 passing, of which 10 are new. Existing tests untouched.
  • The provider tests are mocked, so CI needs no network and no API key.
  • ruff, ruff-format and isort are clean. mypy is clean on the new file, and the 9 pre-existing errors elsewhere are unchanged from master.
  • Ran it for real against a self-hosted vLLM serving Qwen3-Embedding-8B at 4096 dimensions, into a self-hosted Qdrant 1.18.2. Store and search both behave.

What I'm actually asking

I don't want to jump the queue on four people who got here first. So, for whoever picks this up (see the comment below on who's been active lately):

  1. Is this the shape you were asking for in feat: Add OpenAI embedding provider support #55? If it's still too much, dropping the prefix options or the size detection would each make it smaller. Just say which.
  2. Would you rather land one of the existing ones? If Add OpenAI-compatible embedding provider support #92, feat: add OpenAI-compatible embedding provider #111 or feat: add OpenAI-compatible embedding provider #158 is closer to what you want, I'll close this and do the rebase and conflict cleanup on that branch instead. I care that this exists, not that it's mine.
  3. Does folding OpenRouter, Ollama and Gemini into one provider make sense to you, or do you want a class per vendor?

Happy to mark it ready, cut it down, or close it in favour of someone else's branch.

@chaserhkj @sk7n4k3d @HosseinShahabadi @zsxh1990 @pablomichelettii, tagging you since this overlaps what you already wrote. Not trying to step on anyone, I'd just rather we combined efforts than kept filing the same PR.

Adds `EMBEDDING_PROVIDER=openai`, which embeds through any service exposing an
OpenAI-compatible `/v1/embeddings` endpoint. This covers OpenAI itself as well as
locally hosted servers such as Ollama, vLLM and LM Studio.

The provider is additive: FastEmbed remains the default and its behaviour is
unchanged. No existing server code is refactored.

- `EMBEDDING_BASE_URL` points at the endpoint, so one provider serves every
  OpenAI-compatible backend rather than one provider per vendor.
- `EMBEDDING_API_KEY` falls back to `OPENAI_API_KEY` and then to a placeholder,
  because local servers do not check it but the client refuses to start without one.
- `EMBEDDING_VECTOR_SIZE` is detected from the API when unset, since the
  dimensionality is not otherwise discoverable for arbitrary models.
- `EMBEDDING_QUERY_PREFIX` / `EMBEDDING_DOCUMENT_PREFIX` support asymmetric models
  that expect an instruction prefix on queries but not on documents.

Tests are mocked and require no network access.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EnVZkTXryuPFcRqe8R9XAa
@castlenthesky
castlenthesky force-pushed the feat/openai-embeddings-provider branch from 2ccdf6b to 66f18b3 Compare September 5, 2026 13:56
@castlenthesky
castlenthesky marked this pull request as ready for review September 5, 2026 14:15
@castlenthesky
castlenthesky marked this pull request as draft September 5, 2026 14:17
@castlenthesky

castlenthesky commented Sep 5, 2026 •

Copy link
Copy Markdown
Author

Sorry for the tag list, I'll keep this brief.

@anush008 @kanungle @timvisee @tbung, you've all been active here recently so you're who I landed on. @joein, pulling you in as well since this sits right next to fastembed and that's more your area than anyone's. @generall, same, if you have a minute.

No pressure to read the code. What I'm really after is a steer on whether this is wanted and roughly the right shape. Any one of you saying so would unblock it.

Bit of context for why I'm asking instead of just waiting. There are five open PRs doing versions of this, the oldest going back about a year, and none have had a review yet. The Ollama one (#143) was mergeable with CI passing and its author ended up closing it himself after a couple of months of not hearing anything back. I'd rather not quietly add a sixth to that pile, so if the answer is "no thanks" or "not right now", that's genuinely useful and I'll close this out. It doesn't have to be my branch, but the desire for this feature is clear, and the community is effectively begging to help add it - myself included.

If it is wanted, the things I'd like a steer on:

  1. Is this roughly the right size? I took the feat: Add OpenAI embedding provider support #55 comment asking for "just OpenAI embeddings, but nothing else" as the brief, so it's additive only and doesn't touch anything that already exists. If the query and document prefixes or the vector size detection are more than you want to take on, I'm happy to drop either.

  2. Would you rather pick up Add OpenAI-compatible embedding provider support #92, feat: add OpenAI-compatible embedding provider #111 or feat: add OpenAI-compatible embedding provider #158? Any of them could work with a rebase, and I'll gladly do that on someone else's branch instead. I don't mind whose version lands.

  3. Should this be one OpenAI-compatible provider covering Ollama, vLLM, LM Studio and OpenRouter, or a separate one per vendor? A vendor-agnostic approach feels right, but there are also other open PRs with different approaches. Whichever way you go would also settle feat: add OpenRouter embedding provider support #118, feat: Add Google Gemini embedding provider #114, Add Ollama embedding provider support #108 and feat: add Ollama embedding provider for local models #143.

One small practical thing: CI hasn't run here because the workflows need a maintainer to approve them for a first-time contributor. It all passes locally, 34 tests, and the provider tests are mocked so there's no network or API key involved. You just can't see any of that from here until someone clicks approve.

@pablomichelettii

Copy link
Copy Markdown

Hi @castlenthesky,

Please don't worry about "jumping the queue". I don't mind at all. Personally, my PR was born out of a specific need for a project that is still up and running, and it perfectly serves its purpose for my company.

I opened my PR with the exact same spirit as your comment: to help push the development forward. I am more than happy to support any solution that brings this functionality into the main codebase, whether it's my implementation, yours, or an even better one.

I completely agree with you, though: it is definitely time for this issue to move forward.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants