Repository navigation
Conversation
ArcadeDB is an open-source multi-model DBMS that natively supports graph, document, vector (HNSW), and full-text search (Lucene) in a single engine. ArcadeDB 26.2.1+ ships the Neo4j Bolt wire protocol, allowing the existing neo4j Python async driver to connect directly. This adds: - ArcadeDBDriver using AsyncGraphDatabase (Bolt transport) - 11 operation implementations (entity/episode/community/saga nodes, entity/episodic/community/has-episode/next-episode edges, search, graph maintenance) - ArcadeDB-specific SQL DDL for index creation (range + Lucene fulltext) - GraphProvider.ARCADEDB enum value - ArcadeDB branches in shared query builders (node_db_queries, edge_db_queries, graph_queries, search_filters) - arcadedb optional dependency in pyproject.toml Key design decisions: - Reuses neo4j async driver (no new dependency needed) - Embeddings stored as regular list properties (no vector property API) - Labels stored as node property (single-type-per-vertex constraint) - Batch deletes without Neo4j's IN TRANSACTIONS syntax - Vector similarity computed in Python via numpy (can be optimized with ArcadeDB's native vectorNeighbors() in future iterations) - Fulltext search via CONTAINS predicates (can be enhanced with native Lucene index queries in future iterations) Closes getzep#1259 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
I have read the CLA Document and I hereby sign the CLA behalf on myself, e-mail: example@example.com or I have read the CLA Document and I hereby sign the CLA behalf of my company, e-mail: example@example.com Signature is valid for 6 months. 1 out of 3 committers have signed the CLA. |
Add unit tests for ArcadeDBDriver (driver init, query execution, sessions, health check, transactions, operations properties). Add quickstart example following the FalkorDB pattern. Update helpers_test.py with ArcadeDB driver discovery for integration tests. Update README with ArcadeDB installation, configuration, and architecture sections. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
I have read the CLA Document and I hereby sign the CLA |
feafc422 Bump the uv group across 2 directories with 2 updates (#1363) c4e6923b Upstream Zep internal improvements (#1361) e88c09ca @VictorECDSA has signed the CLA in getzep/graphiti#1356 91fe7e0e @majiayu000 has signed the CLA in getzep/graphiti#1351 c52786d2 @dudo has signed the CLA in getzep/graphiti#1350 d631437d @Ker102 has signed the CLA in getzep/graphiti#1339 73cff2cb @chengjon has signed the CLA in getzep/graphiti#1340 8c617639 @rhlsthrm has signed the CLA in getzep/graphiti#1335 e6424bae @pratyush618 has signed the CLA in getzep/graphiti#1332 6f05647c @bsolomon1124 has signed the CLA in getzep/graphiti#1330 10d91394 @spencer2211 has signed the CLA in getzep/graphiti#1326 1ca14686 Add hiring promotion section to README (#1323) 19e44a97 Bump mcp-server to 1.0.2 and require graphiti-core>=0.28.2 (#1317) 77b16096 Bump graphiti-core version to 0.28.2 (#1315) 7d65d5e7 Harden search filters against Cypher injection (#1312) b10b4889 Restore README title and subtitle (#1314) a9065fa9 Refresh README content and fix image refs (#1313) 5a334ec5 @lvca has signed the CLA in getzep/graphiti#1310 45c8040e @jawherkh has signed the CLA in getzep/graphiti#1309 9eb2c9e8 @kraft87 has signed the CLA in getzep/graphiti#1305 334c8faa @adsharma has signed the CLA in getzep/graphiti#1296 b6f9d874 @StephenBadger has signed the CLA in getzep/graphiti#1295 4b91076a feat: Add GLiNER2 hybrid LLM client (#1284) db54ce09 chore: update Docker images to graphiti-core 0.28.1 (#1292) edc71e8e @devmao has signed the CLA in getzep/graphiti#1289 b4ddc55a @carlos-alm has signed the CLA in getzep/graphiti#1288 aa8e81e3 @giulio-leone has signed the CLA in getzep/graphiti#1280 6fdb352f @aelhajj has signed the CLA in getzep/graphiti#1281 2099603d @avianion has signed the CLA in getzep/graphiti#1278 9eb59f7f @themavik has signed the CLA in getzep/graphiti#1214 98f5b5ff fix: replace edge name with uuid in debug log (#1261) 510bd50d @hanxiao has signed the CLA in getzep/graphiti#1257 17a8ea9e @sprotasovitsky has signed the CLA in getzep/graphiti#1254 9d509a2a @Yifan-233-max has signed the CLA in getzep/graphiti#1245 ef52a2ad chore: regenerate lockfiles to drop diskcache (#1244) 76053036 chore: bump version to 0.28.1 (#1243) bde2f797 fix: replace diskcache with sqlite-based cache to resolve CVE (#1238) git-subtree-dir: graphiti git-subtree-split: feafc422c739f0da166241d4804a9830a294d366
|
Please let me know if there is anything I can do. |
Brings in the 20 commits liminis added since this branch forked, including the Kuzu → LadybugDB rename refactor (GraphProvider.LADYBUG, ladybug_driver.py, ladybug/operations/) and the LadybugDriver.close() hardening for clean file-backed teardown. Conflicts resolved in 6 files: - README.md: kept liminis's expanded "Graph Driver Architecture" section (full 11-ABC walkthrough + adding-a-driver guide), and updated the telemetry bullet to read "LadybugDB" instead of upstream's stale "Kuzu". - signatures/version1/cla.json: kept upstream's six new CLA entries appended after PR getzep#1310. - graphiti_core/search/search_filters.py: combined upstream's Cypher injection hardening (validate_node_labels) with liminis's GraphProvider.LADYBUG rename in both node and edge filter constructors. - graphiti_core/search/search_utils.py: combined upstream's validate_group_ids call with liminis's GraphProvider.LADYBUG rename in fulltext_query. - tests/helpers_test.py: kept liminis's three additions — DISABLE_LADYBUG driver registration, LADYBUG_DB env var, and LADYBUG branch in get_driver. - tests/test_graphiti_drivers_int.py: kept liminis's six pytest.skip guards for LadybugDB on tests it doesn't yet support (fulltext indexing, etc.). Note the file rename from test_graphiti_mock.py we did earlier; resolution markers reflected the rename via "liminis:tests/test_graphiti_mock.py". 433 unit tests pass (~10s). Zero new failures versus liminis baseline. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
Hi guys, anew news on this? |
|
Hi @lvca — first, thank you for this driver, it's a solid foundation. We cloned
All 6 verified with real reproduction (not just code reading), plus 8 new/updated unit tests in |
|
@agc-63 amazing job, I'm going to review your PR asap, thanks. |
|
Commented the PR ArcadeData#1 |
database_ was set inside the Cypher query params dict instead of passed as a real kwarg to the underlying neo4j driver call, so per-tenant database routing was silently ignored -- a real cross-tenant data leak (verified: querying tenant B's driver returned tenant A's data). Verified: 30 tenants x 15 rounds concurrent isolation check (450 verifications, 0 leaks).
…ported fulltext DDL get_range_indices()/get_fulltext_indices() generated ArcadeDB SQL DDL (CREATE VERTEX TYPE ... IF NOT EXISTS, CREATE INDEX ... NOTUNIQUE, FULL_TEXT ENGINE LUCENE) but sent it over the Cypher/Bolt channel, which rejects it entirely -- verified all 28 statements failed with 'mismatched input TYPE/ON' syntax errors. - Vertex/edge type declarations are unnecessary: ArcadeDB creates types implicitly on first MERGE/CREATE (schemaless behaviour, verified throughout this driver). - Range indexes rewritten with standard Cypher syntax (CREATE INDEX <name> IF NOT EXISTS FOR (n:Label) ON (n.prop)) -- ArcadeDB does support standard/RANGE indexes via Bolt, just not FULLTEXT. - Fulltext indexes dropped entirely (return []): not supported by ArcadeDB via Bolt regardless of syntax, and graphiti-core already runs without BM25/fulltext search enabled. Verified: build_indices_and_constraints() completes with zero syntax errors (previously all 28 statements failed).
- test_execute_query_success / test_execute_query_with_params: updated to assert the correct behaviour (database_ passed as a kwarg, not inside parameters_) -- the previous versions asserted the buggy behaviour. - test_execute_query_routes_to_cloned_database: with_database() must route execute_query() to the cloned database. - test_range_indices_use_valid_cypher_syntax / test_fulltext_indices_empty_for_arcadedb: DDL generation.
fix: 6 bugs found running the ArcadeDB driver end-to-end (cross-tenant leak, MVCC retry, DDL syntax, edge save, temporal filter)
|
I have read the CLA Document and I hereby sign the CLA behalf on myself, e-mail: angelgarcia@codimatic.com |
|
Hi @paul-paliychuk @prasmussen15 @danielchalef - this adds a native ArcadeDB backend driver and is ready for review. Some context on where it stands: @agc-63 ran the driver end-to-end against a live ArcadeDB instance and found 6 issues. Rather than paper over them in the driver, we fixed the root causes properly:
Everything is covered by regression tests - 21/21 driver unit tests plus 339 non-integration tests pass, and Happy to make any adjustments you'd like. Thanks for taking a look. |
|
Independent validation of this PR from a prospective user, plus one thing that has changed since it was written. We evaluated ArcadeDB as a graphiti backend to replace FalkorDB (SSPL is a problem for a commercial multi-tenant deployment; Neo4j Community turned out unable to isolate tenants by credential at all — RBAC is Enterprise-only). We built a smaller driver independently — subclassing The fulltext limitation this PR works around no longer exists.
CALL db.index.fulltext.queryNodes('Entity[name,summary,group_id]', 'log4shell')
YIELD node, scoreTwo differences from the Neo4j form worth knowing: the procedure takes 2 arguments, so Neo4j's Two things that bit us, in case they apply here too.
Results. Same suite against ArcadeDB, Neo4j 5.26 CE and FalkorDB, identical data and embedder:
Performance, 384-dim, one engine at a time on 4 cores:
FalkorDB's inline cosine scans linearly; ArcadeDB's ANN is near-flat. They cross around 100k nodes. The vector index costs ~3.7× on write throughput, and types need Reporting the four defects we hit produced fixes in under 12 hours (ArcadeData/arcadedb#7056, #7058), which is its own signal about the backend's maintenance. Happy to contribute the fulltext/vector paths, the bulk-save label fix, or the test coverage to this PR rather than opening a competing one — whichever is most useful. Either way this looks ready to us, and we'd like to see it land. |
Retested the driver end-to-end against a live ArcadeDB 26.9.1 server.
26.9.1 and 26.7.3 behaved identically, so nothing regressed with the
version bump, but the run surfaced several pre-existing defects.
Connection details were wrong. Bolt was documented on port 2480, which
is the HTTP port; Bolt ships as a server plugin that must be enlisted at
startup and listens on 7687. Nothing could connect as documented.
Correctness fixes:
- search_filters used Kuzu's list_has_all() for ArcadeDB, which has no
such function. Replaced with portable all(l IN $labels WHERE l IN
n.labels).
- EntityNode.save() never placed labels in entity_data, and ArcadeDB's
save query has no SET n:{labels} clause, so entity labels were silently
dropped on every write and read back as None.
- ArcadeDBGraphMaintenanceOperations.remove_communities() overrode the
base method with an incompatible signature and ignored group_ids
entirely, so it raised TypeError and, called directly, would have
deleted every tenant's communities. Removed; the base implementation is
correct and portable.
Full-text search now works natively. 26.9.1 added Neo4j-compatible
db.index.fulltext.queryNodes/queryRelationships, which did not exist when
this driver was written. They take 2 arguments (Neo4j's {limit: $limit}
config map is an error) and address an index as Type[prop1,prop2]. Index
creation still requires ArcadeDB SQL, which the Cypher channel rejects,
so the driver creates them over the HTTP API; the endpoint is
configurable via http_uri and failure degrades to a warning rather than
breaking the driver.
Edge full-text search also needed an explicit WITH between CALL ... YIELD
and MATCH; without it ArcadeDB silently returns no rows (ArcadeDB #7165).
CI never started an ArcadeDB, even though the provider is in the default
test matrix, so none of this was covered. Added the service and wired the
ArcadeDB driver tests into the integration job.
Known limitation: creating a range index on a property that already holds
data makes ArcadeDB declare it as millisecond DATETIME, truncating
sub-millisecond precision on subsequent writes. This affects the
bi-temporal timestamps on pre-existing databases and is being fixed
engine-side (ArcadeDB #7164) rather than worked around here.
Against 26.9.1: 90 passed, 0 failed. Neo4j and FalkorDB unaffected
(47 passed). Unit suite 371 passed; ruff and pyright clean.
|
Pushed an update (92ab811). I retested the driver end-to-end against a live ArcadeDB 26.9.1 server, and ran the same suite against 26.7.3 side by side to separate version regressions from pre-existing problems. The two releases behaved identically, so the version bump changed nothing — but the retest did surface several real defects, which are fixed here. @g33kroid — thanks, your report was accurate on every point I was able to check. The full-text half is implemented below. Connection details were wrongBolt was documented on port 2480. That is ArcadeDB's HTTP port; Bolt ships as a server plugin that has to be enlisted at startup and listens on 7687. As documented, nothing could connect. Fixed in the README, the quickstart, the driver docstring and the test defaults, with the startup flag spelled out: Correctness fixes
Full-text search now works natively26.9.1 added Neo4j-compatible
Index creation still requires ArcadeDB SQL — Cypher rejects it with "Only standard, RANGE and TEXT index types are supported", and Edge full-text additionally needed an explicit CI never actually ran ArcadeDBArcadeDB is in the default test matrix in Results
One known limitation, tracked upstreamTwo engine-level bugs came out of this, both filed with standalone reproducers:
Both are being fixed in the engine for the next ArcadeDB release rather than papered over in this driver, so no change here depends on an unreleased build — everything above is against released 26.9.1. Happy to adjust anything. Thanks for taking a look. |
… the wrong tenant GraphDriver.clone() returns self. Every other driver overrides it; the ArcadeDB driver did not, so clone() was a no-op here. That default is wrong on ArcadeDB specifically, because a database is the tenant boundary on this backend rather than a filter over one shared store: per-tenant databases are the reason to pick it. graphiti calls clone(group_id) to switch tenant, so with clone() returning self every subsequent query stayed on the previous tenant's database. Nothing raises — the query is valid, the database simply does not hold that tenant's rows — so the caller sees an empty tenant instead of an error, and any code that treats "no results" as "nothing to do" writes a second copy into the wrong database. Found running graphiti's search suite across ArcadeDB, Neo4j and FalkorDB on ArcadeDB 26.9.1: the per-tenant switch returned 0 rows on ArcadeDB and the expected rows on the other two. The Bolt client is shared with the original driver. It is not bound to a database — session() and execute_query() both take database_ per call — so sharing it matches the other drivers, and closing one clone closes them all as it does elsewhere. Two regression tests: clone() returns a distinct driver on the requested database without disturbing the original, and a cloned driver's session() opens on the cloned database.
|
Ran this branch end-to-end on the tagged ArcadeDB 26.9.1 release ( The good news first, since it contradicts what this PR looked like earlier in the year: BM25 now works. Full-text indexes created over HTTP/SQL and queried through the two-argument Two defects, both in the driver rather than in ArcadeDB: 1. Similarity search raises a syntax error whenever a filter is present. The group/uuid filter is emitted as
2. With both applied, the functional suite is 15/15 on ArcadeDB (Neo4j CE 14/15, failing only the physical-store check). One gap worth naming, not a bug: the driver creates no vector index — no
Full-text is competitive and writes beat Neo4j. Vector is 18–27× slower and scales linearly with node count, so at production sizes it is the blocker rather than BM25. ArcadeDB does have 🤖 Generated with Claude Code |
The three similarity searches fetched every candidate row *with its
embedding* over Bolt, then computed cosine in Python with numpy, sorted, and
truncated to the limit. For a 384-dimension embedder that ships ~1.5KB per
candidate row across the wire to discard almost all of it, and it grows with
the tenant's node count rather than with the limit.
ArcadeDB supports the same `vector.similarity.cosine()` that the Neo4j path
uses, so `get_vector_cosine_func_query()` already returns the correct
expression for this provider through its fallthrough — nothing there needed
changing. Scoring, `min_score` filtering, ordering and the limit all move
into the query, which is what the Neo4j, FalkorDB and Kuzu drivers already
do. The Python-side `_cosine_similarity()` helper and the numpy import go
with it.
Measured on ArcadeDB 26.9.1 (digest 02a1a74f), 1800 nodes, fastembed
BGE-small (384-dim), graphiti's own load test:
vector search p50 330.3ms -> 79.9ms
vector search p95 354.9ms -> 133.5ms
Results are unchanged: the same exact cosine over the same candidate set,
computed one hop earlier. The functional suite is 15/15 on ArcadeDB with
this applied.
This also folds the embedding guard into the filter list rather than
emitting it as a second WHERE, because the rewritten queries build one
WHERE. That happens to fix the `WHERE ... WHERE ...` syntax error that made
every filtered similarity search fail, which #2 fixes separately and for
which #2 should get the credit.
|
Follow-up on the vector gap I mentioned above — I benchmarked it rather than guessing, and the answer turned out not to be "add an index". The actual cost is the round trip, not the missing index. The three similarity searches fetch every candidate row with its embedding over Bolt and compute cosine in Python. At 384 dimensions that is ~1.5KB per candidate shipped to be thrown away, scaling with the tenant's node count rather than with the limit. ArcadeDB supports the same
Same exact cosine, same candidate set, unchanged results. Sent as ArcadeData#5. On the LSM_VECTOR index — worth knowing before anyone reaches for it. It does work over Bolt on 26.9.1, and it is faster again, but it returns the global top-k before the tenant filter is applied. 1800 nodes across 3 tenants, querying one:
That is the last item from my list. Recap of where the ArcadeDB backend stands after this: full-text works, bi-temporal invalidation verified end-to-end, per-tenant physical isolation is real (Neo4j CE cannot match it), migration from FalkorDB verified 19/19, and the functional suite is 15/15 with #2, #4 and #5 applied. 🤖 Generated with Claude Code |
|
Sharing our production experience with this driver, plus one finding for reviewers and anyone deploying it today. We've run a staging deployment on One finding that will bite anyone using the generic search path on any currently released engine (26.9.1 is still the latest release): the
(Node-side fulltext is unaffected — no Until 26.10.1 ships, deployments need to either disable the edge bm25 leg in the hybrid search config (what we do) or route search through Minor deployment note: the fulltext queries in With three independent validations now on this thread (your retest, @g33kroid's three-engine comparison, our multi-tenant deployment), it would be great to see this get a review from the Zep side — happy to help with anything missing. |
fix(arcadedb): clone() must switch database, or a tenant switch reads the wrong tenant
perf(arcadedb): score vectors in the engine, not in Python
|
Updated this PR with the latest contributions. |
Summary
Adds ArcadeDB as a new graph database backend for Graphiti, as requested in #1259.
ArcadeDB is an open-source (Apache 2.0) multi-model DBMS that natively supports graph, document, vector (HNSW), and full-text search (Lucene) — all in a single engine. ArcadeDB 26.2.1+ ships the Neo4j Bolt wire protocol, allowing the existing
neo4jPython async driver to connect directly with zero additional dependencies.What's included
ArcadeDBDriver— main driver usingAsyncGraphDatabase(Bolt transport)GraphProvider.ARCADEDBenum valuenode_db_queries,edge_db_queries,graph_queries,search_filters)arcadedboptional dependency inpyproject.toml(empty — reusesneo4jcore dependency)Key design decisions
neo4jasync driver via Bolt protocol — no new dependencydb.create.setNodeVectorProperty())DETACH DELETE(no Neo4jIN TRANSACTIONSsyntax)vectorNeighbors()CONTAINSpredicates; can be enhanced with native Lucene index queriesMATCH (node {uuid: ...})for polymorphic queries (like FalkorDB)Usage
Why ArcadeDB for Graphiti
With Kùzu archived after the Apple acquisition (#1132), users looking for a self-contained, local-first graph + vector + fulltext backend now have ArcadeDB as an option:
Closes #1259
Test plan
arcadedata/arcadedb)make check(ruff + pyright)🤖 Generated with Claude Code