Skip to content

feat: add Ollama embedding provider for local models - #143

Closed
zsxh1990 wants to merge 1 commit into
qdrant:masterfrom
zsxh1990:feat/ollama-provider
Closed

zsxh1990 wants to merge 1 commit into
qdrant:masterfrom
zsxh1990:feat/ollama-provider

Conversation

@zsxh1990

@zsxh1990 zsxh1990 commented Jun 4, 2026

Copy link
Copy Markdown

What

Adds Ollama as an embedding provider alongside FastEmbed, enabling local embedding models without external API keys.

Why

From #62: "Being stuck with fastembed only means this mcp server won't work at all for a lot of us. I'm using Ollama/bge-m3:latest" — @magnus919

@kacperlukawski confirmed: "We have a pretty flexible interface, so adding a new embedding model should not be a big deal. Please feel free to open a PR."

How

New file: src/mcp_server_qdrant/embeddings/ollama.py

class OllamaProvider(EmbeddingProvider):
    """Ollama implementation using /api/embed endpoint."""
    
    def __init__(self, model_name: str, base_url: str = "http://localhost:11434"):
        ...
    
    async def embed_documents(self, documents: list[str]) -> list[list[float]]:
        return await self._embed(documents, input_type="search_document")
    
    async def embed_query(self, query: str) -> list[float]:
        return (await self._embed([query], input_type="search_query"))[0]

Updated files:

  • types.py — Add OLLAMA = "ollama" to EmbeddingProviderType
  • factory.py — Add OLLAMA case to create_embedding_provider()
  • settings.py — Add ollama_base_url field (env: OLLAMA_BASE_URL)

Usage

# .env
EMBEDDING_PROVIDER=ollama
EMBEDDING_MODEL=bge-m3
OLLAMA_BASE_URL=http://localhost:11434

Or with any Ollama-compatible model:

EMBEDDING_PROVIDER=ollama
EMBEDDING_MODEL=nomic-embed-text

Testing

# Start Ollama
ollama serve
ollama pull bge-m3

# Run MCP server
EMBEDDING_PROVIDER=ollama EMBEDDING_MODEL=bge-m3 python -m mcp_server_qdrant

Design decisions

  1. Uses /api/embed endpoint — Ollama's native embedding API, supports batch embedding
  2. Auto-detects vector size — Embeds a dummy text on first call to get dimensions
  3. Separate base_url setting — Allows connecting to remote Ollama instances
  4. Async with httpx — Non-blocking I/O, consistent with FastEmbed provider

Ref

Closes #62

Adds Ollama as an embedding provider alongside FastEmbed.
Supports any model available in Ollama (bge-m3, nomic-embed-text, etc.).

Changes:
- Add OllamaProvider class using Ollama /api/embed endpoint
- Add OLLAMA to EmbeddingProviderType enum
- Update factory to create OllamaProvider
- Add OLLAMA_BASE_URL setting (default: http://localhost:11434)

Usage:
  EMBEDDING_PROVIDER=ollama
  EMBEDDING_MODEL=bge-m3
  OLLAMA_BASE_URL=http://localhost:11434

Ref: qdrant#62
@zsxh1990

zsxh1990 commented Jul 1, 2026

Copy link
Copy Markdown
Author

Hi @kacperlukawski — friendly check-in on this PR 🙂 It's been about 4 weeks since the Ollama embedding provider was added. The implementation follows the existing pattern and enables local embedding models (like ) without external API keys, as requested in #62.

Let me know if there's anything that needs adjusting — happy to rebase or make changes as needed!

@zsxh1990

zsxh1990 commented Jul 2, 2026

Copy link
Copy Markdown
Author

Hi @kacperlukawski — second friendly check-in on this PR 🙂

Status:

If the maintainer team is happy with the direction, I am happy to:

  • Rebase to main if master → main migration is planned
  • Add tests if needed
  • Adjust naming/API to match qdrant-client conventions

If there is no bandwidth to review right now, no problem — happy to keep it open as a reference. If closing is preferred, I can close it cleanly with a summary of what was implemented for future contributors.

Thanks for your time.

@zsxh1990

zsxh1990 commented Jul 7, 2026

Copy link
Copy Markdown
Author

Hi @kacperlukawski — third check-in on this PR 🙂

Status update:

  • 30 days since open (4+ weeks)
  • 2 previous check-ins, no maintainer response
  • PR is mergeable, no conflicts, base=master
  • All CI checks pass

The Ollama embedding provider enables local model usage without API keys — a common request from the community. Would love to get this merged or receive feedback. Thanks!

@zsxh1990

Copy link
Copy Markdown
Author

Hi @qdrant/mcp-server-qdrant maintainers! 👋

Just a friendly bump on this PR. It's been open for about 5 weeks.

Summary: Adds Ollama as an embedding provider for local model support. This enables users to run embeddings locally without cloud API dependencies.

Changes:

  • New class
  • Support for any Ollama-compatible model
  • Auto-detection of available models

Is there anything I can do to help move this forward? Happy to address any feedback or make adjustments.

Thanks! 🙏

@zsxh1990

zsxh1990 commented Aug 3, 2026

Copy link
Copy Markdown
Author

Closing this PR as it has been open for 60 days with no review decision.

This was a contribution to add Ollama embedding provider support for local models. The PR includes tests and documentation. If there's renewed interest, happy to reopen with updated code.

Thanks for maintaining mcp-server-qdrant!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Need to support custom embedding models

1 participant