Skip to content

feat: Add ArcadeDB backend driver - #1310

Open
lvca wants to merge 22 commits into
getzep:mainfrom
ArcadeData:feat/arcadedb-backend
Open

lvca wants to merge 22 commits into
getzep:mainfrom
ArcadeData:feat/arcadedb-backend

Conversation

@lvca

@lvca lvca commented Mar 9, 2026

Copy link
Copy Markdown

Summary

Adds ArcadeDB as a new graph database backend for Graphiti, as requested in #1259.

ArcadeDB is an open-source (Apache 2.0) multi-model DBMS that natively supports graph, document, vector (HNSW), and full-text search (Lucene) — all in a single engine. ArcadeDB 26.2.1+ ships the Neo4j Bolt wire protocol, allowing the existing neo4j Python async driver to connect directly with zero additional dependencies.

What's included

  • ArcadeDBDriver — main driver using AsyncGraphDatabase (Bolt transport)
  • 11 operation implementations — full coverage of all CRUD and search operations:
    • Entity, Episode, Community, Saga node operations
    • Entity, Episodic, Community, HasEpisode, NextEpisode edge operations
    • Search operations (fulltext, similarity, BFS, rerankers)
    • Graph maintenance operations
  • ArcadeDB-specific SQL DDL for index creation (range indexes + Lucene fulltext indexes)
  • GraphProvider.ARCADEDB enum value
  • ArcadeDB branches in shared query builders (node_db_queries, edge_db_queries, graph_queries, search_filters)
  • arcadedb optional dependency in pyproject.toml (empty — reuses neo4j core dependency)

Key design decisions

Concern Approach
Transport Reuses neo4j async driver via Bolt protocol — no new dependency
Embeddings Stored as regular list properties (no db.create.setNodeVectorProperty())
Labels Stored as node property (ArcadeDB has single-type-per-vertex)
Batch deletes Direct DETACH DELETE (no Neo4j IN TRANSACTIONS syntax)
Vector search Cosine similarity computed in Python via numpy; can be optimized with ArcadeDB's native vectorNeighbors()
Fulltext search Uses CONTAINS predicates; can be enhanced with native Lucene index queries
Multi-label MATCH Label-less MATCH (node {uuid: ...}) for polymorphic queries (like FalkorDB)

Usage

from graphiti_core.driver.arcadedb_driver import ArcadeDBDriver

driver = ArcadeDBDriver(
    uri="bolt://localhost:7687",
    user="root",
    password="arcadedb",
    database="graphiti",
)
await driver.build_indices_and_constraints()
# Docker quickstart
docker run -d -p 2480:2480 -p 7687:7687 \
  -e JAVA_OPTS="-Darcadedb.server.plugins=Bolt:com.arcadedb.server.bolt.BoltPlugin" \
  arcadedata/arcadedb

Why ArcadeDB for Graphiti

With Kùzu archived after the Apple acquisition (#1132), users looking for a self-contained, local-first graph + vector + fulltext backend now have ArcadeDB as an option:

  • Graph traversal — native graph engine with O(1) traversal
  • Vector similarity — built-in HNSW vector indexes
  • Full-text search — Lucene-based fulltext indexes
  • License — Apache 2.0 (permissive)
  • Deployment — embedded or server mode, single JAR

Closes #1259

Test plan

  • Unit tests for ArcadeDB driver class (session, query execution, index ops)
  • Integration tests with ArcadeDB Docker container (arcadedata/arcadedb)
  • Verify CRUD operations for all node and edge types
  • Verify search operations (fulltext, similarity, BFS)
  • Verify index creation via SQL DDL over Bolt
  • Run make check (ruff + pyright)

🤖 Generated with Claude Code

ArcadeDB is an open-source multi-model DBMS that natively supports
graph, document, vector (HNSW), and full-text search (Lucene) in a
single engine. ArcadeDB 26.2.1+ ships the Neo4j Bolt wire protocol,
allowing the existing neo4j Python async driver to connect directly.

This adds:
- ArcadeDBDriver using AsyncGraphDatabase (Bolt transport)
- 11 operation implementations (entity/episode/community/saga nodes,
  entity/episodic/community/has-episode/next-episode edges, search,
  graph maintenance)
- ArcadeDB-specific SQL DDL for index creation (range + Lucene fulltext)
- GraphProvider.ARCADEDB enum value
- ArcadeDB branches in shared query builders (node_db_queries,
  edge_db_queries, graph_queries, search_filters)
- arcadedb optional dependency in pyproject.toml

Key design decisions:
- Reuses neo4j async driver (no new dependency needed)
- Embeddings stored as regular list properties (no vector property API)
- Labels stored as node property (single-type-per-vertex constraint)
- Batch deletes without Neo4j's IN TRANSACTIONS syntax
- Vector similarity computed in Python via numpy (can be optimized
  with ArcadeDB's native vectorNeighbors() in future iterations)
- Fulltext search via CONTAINS predicates (can be enhanced with
  native Lucene index queries in future iterations)

Closes getzep#1259

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@danielchalef

danielchalef commented Mar 9, 2026 •

Copy link
Copy Markdown
Member


Thank you for your submission, we really appreciate it. Like many open-source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution. For privacy information, see our Privacy Notice. You can sign the CLA by just posting a Pull Request Comment same as the below format.


I have read the CLA Document and I hereby sign the CLA behalf on myself, e-mail: example@example.com

or

I have read the CLA Document and I hereby sign the CLA behalf of my company, e-mail: example@example.com

Signature is valid for 6 months.


1 out of 3 committers have signed the CLA.
✅ (agc-63)[https://github.com/agc-63]
❌ @lvca
❌ @g33kroid
g33kroid seems not to be a GitHub user. You need a GitHub account to be able to sign the CLA. If you have already a GitHub account, please add the email address used for this commit to your account.
This bot will be retriggered when the Contributor License Agreement comment has been provided. Posted by the CLA Assistant Lite bot.

Add unit tests for ArcadeDBDriver (driver init, query execution,
sessions, health check, transactions, operations properties).
Add quickstart example following the FalkorDB pattern. Update
helpers_test.py with ArcadeDB driver discovery for integration
tests. Update README with ArcadeDB installation, configuration,
and architecture sections.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@lvca

lvca commented Mar 9, 2026

Copy link
Copy Markdown
Author

I have read the CLA Document and I hereby sign the CLA

danielchalef added a commit that referenced this pull request Mar 9, 2026
ehfazrezwan added a commit to ehfazrezwan/neuralscape that referenced this pull request Apr 2, 2026
feafc422 Bump the uv group across 2 directories with 2 updates (#1363)
c4e6923b Upstream Zep internal improvements (#1361)
e88c09ca @VictorECDSA has signed the CLA in getzep/graphiti#1356
91fe7e0e @majiayu000 has signed the CLA in getzep/graphiti#1351
c52786d2 @dudo has signed the CLA in getzep/graphiti#1350
d631437d @Ker102 has signed the CLA in getzep/graphiti#1339
73cff2cb @chengjon has signed the CLA in getzep/graphiti#1340
8c617639 @rhlsthrm has signed the CLA in getzep/graphiti#1335
e6424bae @pratyush618 has signed the CLA in getzep/graphiti#1332
6f05647c @bsolomon1124 has signed the CLA in getzep/graphiti#1330
10d91394 @spencer2211 has signed the CLA in getzep/graphiti#1326
1ca14686 Add hiring promotion section to README (#1323)
19e44a97 Bump mcp-server to 1.0.2 and require graphiti-core>=0.28.2 (#1317)
77b16096 Bump graphiti-core version to 0.28.2 (#1315)
7d65d5e7 Harden search filters against Cypher injection (#1312)
b10b4889 Restore README title and subtitle (#1314)
a9065fa9 Refresh README content and fix image refs (#1313)
5a334ec5 @lvca has signed the CLA in getzep/graphiti#1310
45c8040e @jawherkh has signed the CLA in getzep/graphiti#1309
9eb2c9e8 @kraft87 has signed the CLA in getzep/graphiti#1305
334c8faa @adsharma has signed the CLA in getzep/graphiti#1296
b6f9d874 @StephenBadger has signed the CLA in getzep/graphiti#1295
4b91076a feat: Add GLiNER2 hybrid LLM client (#1284)
db54ce09 chore: update Docker images to graphiti-core 0.28.1 (#1292)
edc71e8e @devmao has signed the CLA in getzep/graphiti#1289
b4ddc55a @carlos-alm has signed the CLA in getzep/graphiti#1288
aa8e81e3 @giulio-leone has signed the CLA in getzep/graphiti#1280
6fdb352f @aelhajj has signed the CLA in getzep/graphiti#1281
2099603d @avianion has signed the CLA in getzep/graphiti#1278
9eb59f7f @themavik has signed the CLA in getzep/graphiti#1214
98f5b5ff fix: replace edge name with uuid in debug log (#1261)
510bd50d @hanxiao has signed the CLA in getzep/graphiti#1257
17a8ea9e @sprotasovitsky has signed the CLA in getzep/graphiti#1254
9d509a2a @Yifan-233-max has signed the CLA in getzep/graphiti#1245
ef52a2ad chore: regenerate lockfiles to drop diskcache (#1244)
76053036 chore: bump version to 0.28.1 (#1243)
bde2f797 fix: replace diskcache with sqlite-based cache to resolve CVE (#1238)

git-subtree-dir: graphiti
git-subtree-split: feafc422c739f0da166241d4804a9830a294d366
@lvca

lvca commented Apr 8, 2026

Copy link
Copy Markdown
Author

Please let me know if there is anything I can do.

verveguy pushed a commit to verveguy/graphiti that referenced this pull request Apr 15, 2026
verveguy pushed a commit to verveguy/graphiti that referenced this pull request Apr 16, 2026
verveguy pushed a commit to verveguy/graphiti that referenced this pull request Apr 16, 2026
verveguy pushed a commit to verveguy/graphiti that referenced this pull request Apr 16, 2026
verveguy added a commit to verveguy/graphiti that referenced this pull request Apr 17, 2026
Brings in the 20 commits liminis added since this branch forked, including the
Kuzu → LadybugDB rename refactor (GraphProvider.LADYBUG, ladybug_driver.py,
ladybug/operations/) and the LadybugDriver.close() hardening for clean
file-backed teardown.

Conflicts resolved in 6 files:

- README.md: kept liminis's expanded "Graph Driver Architecture" section
  (full 11-ABC walkthrough + adding-a-driver guide), and updated the telemetry
  bullet to read "LadybugDB" instead of upstream's stale "Kuzu".

- signatures/version1/cla.json: kept upstream's six new CLA entries appended
  after PR getzep#1310.

- graphiti_core/search/search_filters.py: combined upstream's Cypher injection
  hardening (validate_node_labels) with liminis's GraphProvider.LADYBUG rename
  in both node and edge filter constructors.

- graphiti_core/search/search_utils.py: combined upstream's validate_group_ids
  call with liminis's GraphProvider.LADYBUG rename in fulltext_query.

- tests/helpers_test.py: kept liminis's three additions — DISABLE_LADYBUG
  driver registration, LADYBUG_DB env var, and LADYBUG branch in get_driver.

- tests/test_graphiti_drivers_int.py: kept liminis's six pytest.skip guards
  for LadybugDB on tests it doesn't yet support (fulltext indexing, etc.).
  Note the file rename from test_graphiti_mock.py we did earlier; resolution
  markers reflected the rename via "liminis:tests/test_graphiti_mock.py".

433 unit tests pass (~10s). Zero new failures versus liminis baseline.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@lvca

lvca commented Jun 9, 2026

Copy link
Copy Markdown
Author

Hi guys, anew news on this?

@agc-63

agc-63 commented Jul 3, 2026 •

Copy link
Copy Markdown

Hi @lvca — first, thank you for this driver, it's a solid foundation. We cloned feat/arcadedb-backend (commit e9e4b47) and ran it end-to-end against a real ArcadeDB 26.7.1 instance (multi-tenant writes/reads with real LLM entity extraction, build_communities() under real concurrent writes, isolation checks across tenants). We found and fixed 6 bugs, all reproduced independently with minimal repros, with regression tests added. Opening a PR against feat/arcadedb-backend (not main) with the fixes + tests.

  1. Cross-tenant data leak in ArcadeDBDriver.execute_query() (arcadedb_driver.py:167) — database_ is set inside the query params dict instead of passed as a real kwarg to the underlying neo4j driver call, so per-tenant routing is silently ignored. Repro: query tenant B's with_database()-cloned driver, get tenant A's data back.
  2. No retry on ArcadeDB's optimistic concurrency conflicts — concurrent writes to the same bucket (e.g. build_communities()'s semaphore_gather() of node.save()/edge.save()) fail reproducibly with Neo.ClientError.Transaction.TransactionNotFound / "Concurrent modification on page ... Please retry the operation" — the driver never retries despite ArcadeDB explicitly asking to.
  3. build_indices_and_constraints() DDL never succeeds — get_range_indices/get_fulltext_indices (graph_queries.py) generate ArcadeDB SQL syntax (CREATE VERTEX TYPE ... IF NOT EXISTS, CREATE INDEX ... NOTUNIQUE) but send it over the Cypher/Bolt channel, which rejects all 28 statements with syntax errors. Fix: standard Cypher CREATE INDEX ... FOR (n:Label) ON (n.prop) syntax (ArcadeDB supports standard/RANGE indexes via Bolt, just not FULLTEXT).
  4. Single edge save silently no-ops — get_entity_edge_save_query (edge_db_queries.py, ARCADEDB case) uses MATCH (source:Entity {uuid: $edge_data.source_uuid}) — ArcadeDB doesn't resolve nested parameter field access inside a MATCH pattern. The MERGE runs without error but never creates the relationship. Fix: wrap in UNWIND [$edge_data] AS edge first (same pattern the bulk query already uses, which is why add_episode() never hit this).
  5. retrieve_episodes() temporal filter never matches — WHERE e.valid_at <= $reference_time compares a value ArcadeDB auto-promotes to its native datetime type on write against a plain string parameter — always false, no error. Fix: wrap both sides in datetime(...).
  6. ArcadeDB's Bolt plugin silently drops native Python datetime parameters (not a graphiti-core bug, but blocks everything upstream of it) — confirmed reading ArcadeDB's own PackStreamReader.java: it doesn't decode Bolt's temporal PackStream structs on the input side (output/read direction works fine). Workaround at the neo4j driver level: monkeypatch dehydrate_datetime to emit ISO-8601 strings instead. Might be worth a heads-up to the ArcadeDB team too — happy to open an issue there if useful.

All 6 verified with real reproduction (not just code reading), plus 8 new/updated unit tests in tests/driver/test_arcadedb_driver.py (confirmed each one fails against the unfixed code and passes against the fix). PR: ArcadeData#1

@lvca

lvca commented Jul 3, 2026

Copy link
Copy Markdown
Author

@agc-63 amazing job, I'm going to review your PR asap, thanks.

@lvca

lvca commented Jul 3, 2026

Copy link
Copy Markdown
Author

Commented the PR ArcadeData#1

agc-63 and others added 4 commits July 4, 2026 16:17
database_ was set inside the Cypher query params dict instead of
passed as a real kwarg to the underlying neo4j driver call, so
per-tenant database routing was silently ignored -- a real
cross-tenant data leak (verified: querying tenant B's driver
returned tenant A's data).

Verified: 30 tenants x 15 rounds concurrent isolation check (450
verifications, 0 leaks).
…ported fulltext DDL

get_range_indices()/get_fulltext_indices() generated ArcadeDB SQL DDL
(CREATE VERTEX TYPE ... IF NOT EXISTS, CREATE INDEX ... NOTUNIQUE,
FULL_TEXT ENGINE LUCENE) but sent it over the Cypher/Bolt channel,
which rejects it entirely -- verified all 28 statements failed with
'mismatched input TYPE/ON' syntax errors.

- Vertex/edge type declarations are unnecessary: ArcadeDB creates
  types implicitly on first MERGE/CREATE (schemaless behaviour,
  verified throughout this driver).
- Range indexes rewritten with standard Cypher syntax
  (CREATE INDEX <name> IF NOT EXISTS FOR (n:Label) ON (n.prop)) --
  ArcadeDB does support standard/RANGE indexes via Bolt, just not
  FULLTEXT.
- Fulltext indexes dropped entirely (return []): not supported by
  ArcadeDB via Bolt regardless of syntax, and graphiti-core already
  runs without BM25/fulltext search enabled.

Verified: build_indices_and_constraints() completes with zero syntax
errors (previously all 28 statements failed).
- test_execute_query_success / test_execute_query_with_params: updated to
  assert the correct behaviour (database_ passed as a kwarg, not inside
  parameters_) -- the previous versions asserted the buggy behaviour.
- test_execute_query_routes_to_cloned_database: with_database() must route
  execute_query() to the cloned database.
- test_range_indices_use_valid_cypher_syntax / test_fulltext_indices_empty_for_arcadedb:
  DDL generation.
fix: 6 bugs found running the ArcadeDB driver end-to-end (cross-tenant leak, MVCC retry, DDL syntax, edge save, temporal filter)
@agc-63

agc-63 commented Jul 5, 2026

Copy link
Copy Markdown

I have read the CLA Document and I hereby sign the CLA behalf on myself, e-mail: angelgarcia@codimatic.com

zep-cla-assistant Bot added a commit that referenced this pull request Jul 5, 2026
@lvca

lvca commented Jul 5, 2026

Copy link
Copy Markdown
Author

Hi @paul-paliychuk @prasmussen15 @danielchalef - this adds a native ArcadeDB backend driver and is ready for review.

Some context on where it stands: @agc-63 ran the driver end-to-end against a live ArcadeDB instance and found 6 issues. Rather than paper over them in the driver, we fixed the root causes properly:

  • 4 were engine/Bolt bugs in ArcadeDB itself - all now fixed and merged upstream (arcadedb #4906, #4910, #4921, #4923), shipping in the next ArcadeDB release (26.7.2). These cover inbound/outbound native temporal handling, offset-datetime write persistence, retryable MVCC conflicts, and nested $param.field resolution inside MATCH.
  • 2 are genuine driver-side fixes kept in this PR: a cross-tenant routing fix (database_ was being passed as a Cypher parameter instead of a real driver kwarg, so per-tenant DB routing was silently ignored) and a Cypher CREATE INDEX DDL rewrite (the previous code sent ArcadeDB SQL DDL over the Bolt/Cypher channel, which rejects it).

Everything is covered by regression tests - 21/21 driver unit tests plus 339 non-integration tests pass, and ruff check/format are clean. CI here is green except the triage job, which fails on the token/permission model for fork-based pull_request_target runs (infra, not the code).

Happy to make any adjustments you'd like. Thanks for taking a look.

@g33kroid

g33kroid commented Sep 3, 2026

Copy link
Copy Markdown

Independent validation of this PR from a prospective user, plus one thing that has changed since it was written.

We evaluated ArcadeDB as a graphiti backend to replace FalkorDB (SSPL is a problem for a commercial multi-tenant deployment; Neo4j Community turned out unable to isolate tenants by credential at all — RBAC is Enterprise-only). We built a smaller driver independently — subclassing Neo4jDriver rather than reimplementing the ops interfaces — purely to answer "does this work", and then ran it against a real workload.

The fulltext limitation this PR works around no longer exists.

test_fulltext_indices_empty_for_arcadedb documents ArcadeDB not exposing fulltext over Bolt. That was true when this PR was written. On 2026-08-25 ArcadeDB added Neo4j-compatible Cypher procedures — db.index.fulltext.queryNodes, db.index.fulltext.queryRelationships, db.index.vector.queryNodes — registered in CypherProcedureRegistry and reachable over Bolt. Verified working:

CALL db.index.fulltext.queryNodes('Entity[name,summary,group_id]', 'log4shell')
YIELD node, score

Two differences from the Neo4j form worth knowing: the procedure takes 2 arguments, so Neo4j's {limit: $limit} config map is a hard error (callers already append ORDER BY score DESC LIMIT $limit, so dropping it is safe); and an index is addressed as Type[prop1,prop2] rather than by a logical name, so a mapping constant like the existing NEO4J_TO_FALKORDB_MAPPING is needed. FULL_TEXT index creation still has to go over HTTP/SQL — Cypher rejects it with "Only standard, RANGE and TEXT index types are supported", and CREATE TEXT INDEX builds a non-BM25 index the procedure won't accept.

Two things that bit us, in case they apply here too.

  1. get_entity_node_save_bulk_query emits SET n:$(node.labels). ArcadeDB doesn't interpolate dynamic labels — it creates a vertex type named $(node.labels) and attaches it to the nodes. Nothing fails at write time; the next episode's dedup reads those nodes back, pydantic rejects the label, and ingest stops. Filed as Cypher dynamic labels: SET n:$(expr) creates a literal vertex type named '$(expr)' ArcadeData/arcadedb#7059. Emitting labels literally per node — as the FalkorDB and Neptune branches in that same function already do — fixes it.

  2. The four fulltext search methods in Neo4jSearchOperations call the module-level _build_neo4j_fulltext_query directly rather than self.build_fulltext_query, even though that method is declared on the SearchOperations ABC as an override point. Any Bolt-compatible subclass silently gets Neo4j's Lucene field-scoped syntax. Small fix, and independent of this PR.

Results. Same suite against ArcadeDB, Neo4j 5.26 CE and FalkorDB, identical data and embedder:

  • 15/15 functional (all retrieval legs, tenancy, bi-temporal persistence) — same as FalkorDB; Neo4j CE fails only on physical per-tenant stores
  • Real two-episode ingest with gpt-4o-mini: entity extraction and edge invalidation correct, invalid_at taken from the episode text
  • build_communities: identical output on all three backends
  • FalkorDB → ArcadeDB migration: 19/19 (counts, uuids, edge endpoints, expired edges, valid_at instants)
  • Backup/restore, in-place upgrade, Raft HA failover (~12s), network partition (no split-brain), disk exhaustion, TLS on Bolt — all pass
  • Tenant access control: cross-tenant reads denied on both HTTP and Bolt, and no database enumeration — stricter than FalkorDB, which leaks tenant names via GRAPH.LIST unless you revoke it

Performance, 384-dim, one engine at a time on 4 cores:

nodes FalkorDB ArcadeDB (LSM_VECTOR ANN)
10k 7.7 ms 45.5 ms
50k 32.1 ms 59.8 ms
200k 129.5 ms 78.2 ms

FalkorDB's inline cosine scans linearly; ArcadeDB's ANN is near-flat. They cross around 100k nodes. The vector index costs ~3.7× on write throughput, and types need BUCKETS 16 at creation for concurrent writes (worth 3× — ALTER TYPE ... BUCKETS is rejected).

Reporting the four defects we hit produced fixes in under 12 hours (ArcadeData/arcadedb#7056, #7058), which is its own signal about the backend's maintenance.

Happy to contribute the fulltext/vector paths, the bulk-save label fix, or the test coverage to this PR rather than opening a competing one — whichever is most useful. Either way this looks ready to us, and we'd like to see it land.

Retested the driver end-to-end against a live ArcadeDB 26.9.1 server.
26.9.1 and 26.7.3 behaved identically, so nothing regressed with the
version bump, but the run surfaced several pre-existing defects.

Connection details were wrong. Bolt was documented on port 2480, which
is the HTTP port; Bolt ships as a server plugin that must be enlisted at
startup and listens on 7687. Nothing could connect as documented.

Correctness fixes:

- search_filters used Kuzu's list_has_all() for ArcadeDB, which has no
  such function. Replaced with portable all(l IN $labels WHERE l IN
  n.labels).
- EntityNode.save() never placed labels in entity_data, and ArcadeDB's
  save query has no SET n:{labels} clause, so entity labels were silently
  dropped on every write and read back as None.
- ArcadeDBGraphMaintenanceOperations.remove_communities() overrode the
  base method with an incompatible signature and ignored group_ids
  entirely, so it raised TypeError and, called directly, would have
  deleted every tenant's communities. Removed; the base implementation is
  correct and portable.

Full-text search now works natively. 26.9.1 added Neo4j-compatible
db.index.fulltext.queryNodes/queryRelationships, which did not exist when
this driver was written. They take 2 arguments (Neo4j's {limit: $limit}
config map is an error) and address an index as Type[prop1,prop2]. Index
creation still requires ArcadeDB SQL, which the Cypher channel rejects,
so the driver creates them over the HTTP API; the endpoint is
configurable via http_uri and failure degrades to a warning rather than
breaking the driver.

Edge full-text search also needed an explicit WITH between CALL ... YIELD
and MATCH; without it ArcadeDB silently returns no rows (ArcadeDB #7165).

CI never started an ArcadeDB, even though the provider is in the default
test matrix, so none of this was covered. Added the service and wired the
ArcadeDB driver tests into the integration job.

Known limitation: creating a range index on a property that already holds
data makes ArcadeDB declare it as millisecond DATETIME, truncating
sub-millisecond precision on subsequent writes. This affects the
bi-temporal timestamps on pre-existing databases and is being fixed
engine-side (ArcadeDB #7164) rather than worked around here.

Against 26.9.1: 90 passed, 0 failed. Neo4j and FalkorDB unaffected
(47 passed). Unit suite 371 passed; ruff and pyright clean.
@lvca

lvca commented Sep 5, 2026

Copy link
Copy Markdown
Author

Pushed an update (92ab811). I retested the driver end-to-end against a live ArcadeDB 26.9.1 server, and ran the same suite against 26.7.3 side by side to separate version regressions from pre-existing problems. The two releases behaved identically, so the version bump changed nothing — but the retest did surface several real defects, which are fixed here.

@g33kroid — thanks, your report was accurate on every point I was able to check. The full-text half is implemented below.

Connection details were wrong

Bolt was documented on port 2480. That is ArcadeDB's HTTP port; Bolt ships as a server plugin that has to be enlisted at startup and listens on 7687. As documented, nothing could connect. Fixed in the README, the quickstart, the driver docstring and the test defaults, with the startup flag spelled out:

-Darcadedb.server.plugins=Bolt:com.arcadedb.bolt.BoltProtocolPlugin

Correctness fixes

  • list_has_all() was Kuzu-only. The ArcadeDB branch of search_filters reused Kuzu's function, which ArcadeDB does not have, so every label-filtered search failed with UnknownFunction. Replaced with portable all(l IN $labels WHERE l IN n.labels).
  • Entity labels were silently dropped on write. EntityNode.save() never puts labels into entity_data, because for Neo4j and FalkorDB the labels travel in the query text as SET n:{labels}. ArcadeDB's save query is a plain SET n = $entity_data with no such clause, so labels were never persisted and read back as None. Nothing failed at write time.
  • remove_communities() was broken and unsafe. The ArcadeDB override took an incompatible signature (TypeError on group_ids) and ignored group_ids entirely, so if it had been reachable it would have deleted every tenant's communities rather than the requested ones. Removed — the base implementation is already correct and portable.

Full-text search now works natively

26.9.1 added Neo4j-compatible db.index.fulltext.queryNodes / queryRelationships, which did not exist when this PR was written; the old get_fulltext_indices(ARCADEDB) == [] behaviour is gone. Two differences from the Neo4j form, both handled:

  • the procedure takes 2 arguments — Neo4j's {limit: $limit} config map is a hard error, and callers already append ORDER BY score DESC LIMIT $limit, so dropping it is safe;
  • an index is addressed as Type[prop1,prop2] rather than by a logical name, so there is a mapping constant kept adjacent to the DDL it has to agree with, plus a test asserting the two cannot drift.

Index creation still requires ArcadeDB SQL — Cypher rejects it with "Only standard, RANGE and TEXT index types are supported", and CREATE TEXT INDEX builds a non-BM25 index the procedure will not accept. The driver therefore creates these over the HTTP API. The endpoint is configurable (http_uri=, defaulting to the Bolt host on 2480), and if it is unreachable the indexes are skipped with a warning instead of breaking the driver; everything else keeps working.

Edge full-text additionally needed an explicit WITH between CALL ... YIELD and MATCH — without it ArcadeDB silently returns no rows.

CI never actually ran ArcadeDB

ArcadeDB is in the default test matrix in helpers_test.py, but the integration job started only Neo4j and FalkorDB and never set DISABLE_ARCADEDB, so none of this was covered. Added the service to unit_tests.yml, disabled the provider in the no-database job, and wired the ArcadeDB driver tests into the integration job so this stays honest.

Results

  • ArcadeDB 26.9.1: 90 passed, 0 failed
  • Neo4j 5.26 and FalkorDB: 47 passed, unchanged — the shared-file edits are all provider-gated
  • Unit suite: 371 passed; ruff check, ruff format and pyright clean

One known limitation, tracked upstream

Two engine-level bugs came out of this, both filed with standalone reproducers:

  • ArcadeData/arcadedb#7164 — creating an index on a property that already holds data makes ArcadeDB declare it as millisecond DATETIME, truncating sub-millisecond precision on every subsequent write. This is the one to be aware of: on a fresh database the indexes are created before any data exists, so precision survives and the suite is green, but an existing database adopting this driver will lose microseconds on the bi-temporal fields from the moment build_indices_and_constraints() runs. Seeding a single row before the run reproduces it.
  • ArcadeData/arcadedb#7165 — the CALL ... YIELD / MATCH composition above, worked around driver-side here.

Both are being fixed in the engine for the next ArcadeDB release rather than papered over in this driver, so no change here depends on an unreleased build — everything above is against released 26.9.1.

Happy to adjust anything. Thanks for taking a look.

… the wrong tenant

GraphDriver.clone() returns self. Every other driver overrides it; the
ArcadeDB driver did not, so clone() was a no-op here.

That default is wrong on ArcadeDB specifically, because a database is the
tenant boundary on this backend rather than a filter over one shared store:
per-tenant databases are the reason to pick it. graphiti calls
clone(group_id) to switch tenant, so with clone() returning self every
subsequent query stayed on the previous tenant's database. Nothing raises —
the query is valid, the database simply does not hold that tenant's rows —
so the caller sees an empty tenant instead of an error, and any code that
treats "no results" as "nothing to do" writes a second copy into the wrong
database.

Found running graphiti's search suite across ArcadeDB, Neo4j and FalkorDB on
ArcadeDB 26.9.1: the per-tenant switch returned 0 rows on ArcadeDB and the
expected rows on the other two.

The Bolt client is shared with the original driver. It is not bound to a
database — session() and execute_query() both take database_ per call — so
sharing it matches the other drivers, and closing one clone closes them all
as it does elsewhere.

Two regression tests: clone() returns a distinct driver on the requested
database without disturbing the original, and a cloned driver's session()
opens on the cloned database.
@g33kroid

g33kroid commented Sep 8, 2026

Copy link
Copy Markdown

Ran this branch end-to-end on the tagged ArcadeDB 26.9.1 release (arcadedata/arcadedb@sha256:02a1a74f…, driver at 92ab811), against Neo4j 5.26 CE and FalkorDB as controls — same data, same checks, three engines. Evaluating it as a graph store for a multi-tenant security platform, so the tenancy and temporal behaviour got most of the attention.

The good news first, since it contradicts what this PR looked like earlier in the year: BM25 now works. Full-text indexes created over HTTP/SQL and queried through the two-argument db.index.fulltext.queryNodes('Type[props]', …) is a sound answer to ArcadeDB not exposing Lucene over Bolt, and node/edge/episode full-text all pass. Bi-temporal edge invalidation was observed end-to-end through a real LLM ingestion run. A FalkorDB → ArcadeDB migration of two tenants verified clean at 19/19, including expired edges and valid_at values preserved, and tenant isolation holding through the migration. ArcadeDB is also the only one of the three that passes the physical per-tenant store check — Neo4j CE fails it, since one user database means tenancy is a filter.

Two defects, both in the driver rather than in ArcadeDB:

1. Similarity search raises a syntax error whenever a filter is present. The group/uuid filter is emitted as WHERE …, then the embedding guard is concatenated as a second WHERE:

MATCH (n:Entity) WHERE n.group_id IN $group_ids
            WHERE n.name_embedding IS NOT NULL

CypherSyntaxError: Unexpected input 'WHERE' at line 2, column 12. Three sites — node, edge and community similarity. It fires only when a filter exists, which in a multi-tenant deployment is always, so the dense leg of hybrid search is effectively dead. ArcadeData#2 already fixes exactly this and has been open since July; it just needs merging, so I have not duplicated it.

2. clone() is a no-op, so a tenant switch reads the wrong tenant. GraphDriver.clone() returns self and ArcadeDBDriver does not override it. On this backend a database is the tenant boundary, so clone(group_id) leaves every subsequent query on the previous tenant's database — no error, just zero rows, which any "nothing there yet" branch will happily act on. Fix plus two regression tests: ArcadeData#4.

With both applied, the functional suite is 15/15 on ArcadeDB (Neo4j CE 14/15, failing only the physical-store check).

One gap worth naming, not a bug: the driver creates no vector index — no LSM_VECTOR, nothing dimension-aware — so the dense leg is a brute-force scan. Measured over 1800 nodes, all three engines on the same data and embedder:

arcadedb neo4j falkordb
writes/sec 199 73 926
fulltext p50 (ms) 8.0 10.4 7.8
vector p50 (ms) 330.3 18.2 11.9

Full-text is competitive and writes beat Neo4j. Vector is 18–27× slower and scales linearly with node count, so at production sizes it is the blocker rather than BM25. ArcadeDB does have LSM_VECTOR with a dimensions/similarity metadata block; wiring it into build_indices_and_constraints() (dimension must match the embedder) would close the gap. Happy to send that as a separate PR if it is wanted here rather than in the ArcadeData fork.

🤖 Generated with Claude Code

The three similarity searches fetched every candidate row *with its
embedding* over Bolt, then computed cosine in Python with numpy, sorted, and
truncated to the limit. For a 384-dimension embedder that ships ~1.5KB per
candidate row across the wire to discard almost all of it, and it grows with
the tenant's node count rather than with the limit.

ArcadeDB supports the same `vector.similarity.cosine()` that the Neo4j path
uses, so `get_vector_cosine_func_query()` already returns the correct
expression for this provider through its fallthrough — nothing there needed
changing. Scoring, `min_score` filtering, ordering and the limit all move
into the query, which is what the Neo4j, FalkorDB and Kuzu drivers already
do. The Python-side `_cosine_similarity()` helper and the numpy import go
with it.

Measured on ArcadeDB 26.9.1 (digest 02a1a74f), 1800 nodes, fastembed
BGE-small (384-dim), graphiti's own load test:

    vector search p50   330.3ms -> 79.9ms
    vector search p95   354.9ms -> 133.5ms

Results are unchanged: the same exact cosine over the same candidate set,
computed one hop earlier. The functional suite is 15/15 on ArcadeDB with
this applied.

This also folds the embedding guard into the filter list rather than
emitting it as a second WHERE, because the rewritten queries build one
WHERE. That happens to fix the `WHERE ... WHERE ...` syntax error that made
every filtered similarity search fail, which #2 fixes separately and for
which #2 should get the credit.
@g33kroid

g33kroid commented Sep 8, 2026

Copy link
Copy Markdown

Follow-up on the vector gap I mentioned above — I benchmarked it rather than guessing, and the answer turned out not to be "add an index".

The actual cost is the round trip, not the missing index. The three similarity searches fetch every candidate row with its embedding over Bolt and compute cosine in Python. At 384 dimensions that is ~1.5KB per candidate shipped to be thrown away, scaling with the tenant's node count rather than with the limit. ArcadeDB supports the same vector.similarity.cosine() the Neo4j path uses, so get_vector_cosine_func_query() already returns the right expression for this provider through its fallthrough — the driver simply never used it. Moving scoring, min_score, ordering and the limit into the query, exactly as the Neo4j/FalkorDB/Kuzu drivers do:

before after
vector p50 330.3 ms 79.9 ms
vector p95 354.9 ms 133.5 ms

Same exact cosine, same candidate set, unchanged results. Sent as ArcadeData#5.

On the LSM_VECTOR index — worth knowing before anyone reaches for it. It does work over Bolt on 26.9.1, and it is faster again, but it returns the global top-k before the tenant filter is applied. 1800 nodes across 3 tenants, querying one:

in-query cosine, tenant-filtered : p50  35.1ms  rows=10
index ANN k=10   + tenant filter : p50   7.3ms  rows=2   <- want 10
index ANN k=30   + tenant filter : p50   6.5ms  rows=10

k = limit silently returns 2 rows where 10 were asked for. No error, just a short result — the failure mode hardest to spot inside a hybrid search that fuses three legs, and it gets worse as tenant count rises, since the required over-fetch tracks selectivity. So I have deliberately left the index out of #5 and kept exact semantics. If graphiti wants the ANN path, the over-fetch factor (or a fallback to the exact query when the filtered result comes up short) is a design call for the driver owners, and I am happy to implement whichever shape you prefer.

That is the last item from my list. Recap of where the ArcadeDB backend stands after this: full-text works, bi-temporal invalidation verified end-to-end, per-tenant physical isolation is real (Neo4j CE cannot match it), migration from FalkorDB verified 19/19, and the functional suite is 15/15 with #2, #4 and #5 applied.

🤖 Generated with Claude Code

@agc-63

agc-63 commented Sep 30, 2026

Copy link
Copy Markdown

Sharing our production experience with this driver, plus one finding for reviewers and anyone deploying it today.

We've run a staging deployment on e9e4b47 + our fixes (ArcadeData#1, merged) since July — one ArcadeDB database per tenant, 23 databases, real LLM ingestion — and today moved it to ArcadeDB 26.9.1 + 92ab811, re-validating end-to-end (reads, writes, hybrid search). No issues with the bump itself.

One finding that will bite anyone using the generic search path on any currently released engine (26.9.1 is still the latest release): the edge_fulltext_search caller in search_utils.py silently returns 0 rows. The root cause is the engine, not the driver — ArcadeData/arcadedb#7165 (fixed on the 26.10.1 milestone, still unreleased): a CALL db.index.fulltext.queryRelationships(...) followed by a MATCH that consumes the YIELD variable returns empty unless a WITH separates the two. Reproduced on 26.9.1:

  • YIELD relationship, score alone → 4 matching rows
  • same query + MATCH (n:Entity)-[e:RELATES_TO {uuid: rel.uuid}]->(m:Entity) → 0 rows
  • same query with WITH rel, score in between → 4 rows

(Node-side fulltext is unaffected — no MATCH after its YIELD.)

Until 26.10.1 ships, deployments need to either disable the edge bm25 leg in the hybrid search config (what we do) or route search through search_interface — @g33kroid's note above already points at ArcadeData#2 for that, open since July.

Minor deployment note: the fulltext queries in 92ab811 target the two-argument db.index.fulltext.query* procedures and Type[props] index naming, so this branch needs a 26.9.x engine — older releases reject the syntax.

With three independent validations now on this thread (your retest, @g33kroid's three-engine comparison, our multi-tenant deployment), it would be great to see this get a review from the Zep side — happy to help with anything missing.

@lvca

lvca commented Oct 1, 2026

Copy link
Copy Markdown
Author

Updated this PR with the latest contributions.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] ArcadeDB backend support

4 participants