Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docusaurus-docs/docs/cli/bulk.md
Original file line number Diff line number Diff line change
Expand Up @@ -164,6 +164,7 @@ Flags:
--skip_map_phase Skip the map phase (assumes that map output files already exist).
--skip_reduce_phase Skip the reduce phase (stops after map phase completion).
--store_xids Generate an xid edge for each node.
--tablet_placement string Path to a JSON file pinning predicates to groups, as an array of {"predicate", "group", "namespace"} entries. Pinned predicates are written to the output shard of their group (group N is out/<N-1>); everything else is packed as usual. Placement is applied at load time only; until the cluster enforces pins, Zero's rebalancer may later move tablets. Use map_shards > reduce_shards so unpinned predicates still balance by size.
--tls string TLS Client options
ca-cert=; The CA cert file used to verify server certificates. Required for enabling TLS.
client-cert=; (Optional) The Cert file provided by the client to the server.
Expand Down
52 changes: 52 additions & 0 deletions docusaurus-docs/docs/migration/bulk-loader.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,6 +153,57 @@ Copy `p` directories directly for faster deployment:
3. Start all Alphas simultaneously
4. Verify all Alphas create snapshots with matching index values

## Predicate Placement

By default, Bulk Loader packs predicates into reduce shards by size, so which group serves a given predicate varies between runs. The `--tablet_placement` flag pins chosen predicates to specific groups, so the cluster starts with a deterministic, operator-chosen layout.

Create a placement file containing a JSON array of entries:

```json
[
{"predicate": "payload", "group": 2},
{"predicate": "friend", "group": 3, "namespace": 0}
]
```

| Field | Description |
|-------|-------------|
| `predicate` | Predicate name, without a namespace prefix |
| `group` | Target Alpha group, from 1 to `--reduce_shards`. Group N's data is written to `out/<N-1>/p` |
| `namespace` | Optional namespace the predicate belongs to (default: `0`) |

Pass the file to the loader:

```sh
dgraph bulk \
--files data.rdf.gz \
--schema schema.txt \
--zero localhost:5080 \
--map_shards 6 \
--reduce_shards 3 \
--tablet_placement placement.json
```

### Behavior

- A pinned predicate's data, index, and schema keys are all written to its group's output directory. This includes predicates that appear only in the schema and carry no data in the load.
- Predicates not listed in the file keep the default size-balanced packing. Use `--map_shards` greater than `--reduce_shards` for this: with equal values every map shard is dedicated to a group and size balancing is disabled (the loader logs a notice).
- The loader logs the routing of every pinned predicate, for example `pinned: 0-payload -> map shard 1 -> out/1/p (group 2)`, and warns after the map phase about pinned predicates that matched nothing in the schema or data — usually a typo in the placement file.

### Validation

The loader exits with an error before loading any data if the placement file contains:

- a group below 1 or above `--reduce_shards`
- duplicate entries for the same namespace and predicate
- reserved (`dgraph.*`) predicates — these are always served by group 1
- unknown fields, or a document that isn't a JSON array
- entries for namespaces that can never match when `--force-namespace` is also set

:::note
Placement is applied at load time only. Once the cluster is running, Zero's automatic rebalancer (`--rebalance_interval`, default 8 minutes) may move tablets between groups; set `--rebalance_interval=0` on every Zero to disable automatic rebalancing and preserve the layout. `/moveTablet` can still move tablets at any time.
:::

## Multi-tenancy

By default, Bulk Loader preserves namespace information from data files. Without namespace info, data loads into the default namespace.
Expand Down Expand Up @@ -267,6 +318,7 @@ Increase if you have RAM to spare:
| `--xidmap` | Directory for XID→UID mappings |
| `--format` | Force format (`rdf` or `json`) |
| `--force-namespace` | Load into specific namespace |
| `--tablet_placement` | JSON file pinning predicates to groups |
| `--encryption` | Encryption key file |
| `--encrypted` | Input files are encrypted |
| `--encrypted_out` | Encrypt output (default: true if key provided) |
Expand Down
Loading