Research intake for gated APIs, JS pages and scanned PDFs.
pip install gather-enginegather pulls research out of the places most tools break on: arXiv papers, authenticated JSON APIs, JavaScript-rendered pages via a real headless browser, scanned images through OCR, and audio through transcription, alongside video, web, feeds, and local docs. The core runs with zero third-party runtime dependencies, and the same engine is reachable from the CLI, MCP tools, and plain Python. Every run writes a receipt you can re-check.
Project Telos | gather | crucible | index | forum | telos | learn | emet | buildlang
gather-engine 2.3.0 is the current source version. Core intake,
content-addressed corpus storage, exact UTF-8 source-byte receipts, readable
context selection, descriptor handoff, pilot monitoring, CLI, Python API, and MCP
surfaces are present in this checkout. Browser rendering, OCR, audio, and fast
parsing remain explicit optional capabilities reported by gather caps.
Use gather status --json and gather doctor --json to inspect the installed
surface, gather caps to see which optional backends are available, and
gather mcp to expose status, documentation, source, context, run, and pilot
operations to an MCP host. A run or pilot that would start a command, reach the
network or send a credential needs a grant you set at launch
(--allow-exec, --allow-network, --auth-env NAME@HOST); without it the call
returns GRANT_REQUIRED. File sources and every MCP path argument refuse Windows
network and device paths (\\host\share, \\?\, CON) with NON_LOCAL_PATH,
so a tool call cannot make the machine sign in to someone else's share. External
tools start from an absolute path in a private empty folder.
USAGE.md has the details.
gather pilot takes a closed list of sources and produces something a third
party can check without trusting the run that made it.
The edge worth reading twice is the one that leaves the pipeline. Every value
the extractor proposes has to be found on the page that was actually fetched,
matched on word boundaries so 100 does not match inside 1000. A source that
fails that check is recorded as an error and nothing from it is stored, so a
fabricated field cannot reach the corpus by being plausible.
- Extract any page to LLM-ready Markdown.
gather extractturns a URL or local HTML file into structured Markdown plus a per-block record binding every block to its source node path and content hash.gather markdownprints the Markdown alone. - Crawl whole sites. A concurrent, resumable crawler with robots.txt, sitemap discovery, URL dedup, and per-host throttling.
gather crawl <url> --depth 2 --max-pages 50emits an append-only, hash-chained ledger of the crawl as JSON. - Track elements across redesigns. Fingerprint a scraped element once, then relocate it in any later version of the page and get a typed MATCH, RELOCATED, DRIFT, or GONE verdict (
gather.track, Python API). - Structured extraction with a hallucination check.
gather.schema_extractbinds schema fields to source nodes, andverify_recordrejects any LLM-proposed field value not grounded in the fetched content. - Streaming extraction.
gather.streamparses an HTML stream chunk by chunk and emits partial-update commits as blocks complete, each folded into a hash chain, so a streamed extraction is replayable. - Hard-source adapters behind one shape. Video with captions and comments (
yt-dlpfirst, with the official YouTube Data API as a fallback when you supply your own key; each item records which path served it and how fast), Reddit through its official Data API with your own app credentials, static web, RSS/Atom feeds, local docs, arXiv, PDFs (pdftotext), authenticated JSON APIs (token from env, never logged), JS-rendered pages (headless Chromium), scanned images (tesseract), and audio (whisper). Each external tool is optional and only needed for its own adapter. - Cited reports, checked in code.
gather reportasks a model on your own machine to answer from fixed excerpts with exact quotes, then checks every quote against the excerpt it cites and marks any it cannot verify. An answer that cites an excerpt without quoting it is sent back for a rewrite, up to twice. - Filter ledgers.
--scope TERMS --ledger DIRwrites one row per dropped item with its reason code, plus the input, sogather ledger verifycan recompute every count. - Channel and playlist intake.
gather channel <url> --store DIRlists a channel's videos, shorts, and streams tabs (or one playlist) and gathers every entry with bounded concurrency, resumable from a per-pass ledger. Metadata and comments run as one pass and captions as another (--no-captions,--captions-only), with polite pacing, bounded backoff on HTTP 429 and session rate limits, a stop at the first bot check (never retried or answered), one caption track per video, and a run summary that counts every outcome by reason. - Scholarly-graph federation.
gather scholarqueries OpenAlex, Semantic Scholar, and Crossref in one call, dedupes results by normalized DOI (never a fuzzy title match), and can capture citation edges as first-class records with--edges. - A durable local corpus. Any fetch command takes
--store DIR: bodies are content-addressed and deduped by hash, new writes preserve exact UTF-8 source bytes, andgather corpus list|verify|digest|runs|search|stats|prune|availability|contextinspects, re-checks, and queries what you stored. - Readable context selection.
gather corpus context DIRshows bounded, re-hashed source/comment excerpts with row refs and missing/corrupt/unsafe/oversized-body reasons;--select ROW_REF[:START[:LIMIT]] --expect-digest SHA256exports a private context payload with a deterministic selection digest, refusing stale digests and selected text over budget. Python callers that already hold a retained corpus root fd/HANDLE can pass a same-processCorpusRootDescriptorso Gather does not reopen the root path. On Linux/WSL filesystems where retained directory fds cannot supply stable confined reads, currently including WSL Windows-drive 9p/v9fs mounts, Gather refuses opened corpus, descendant directory, and catalog/body file descriptors before reading corpus metadata or bodies. This is acquisition context, not a truth, claim-support, completeness, caller workspace-parent, or secret-free verdict. - Multi-source runs.
gather run config.jsonorchestrates many sources, a scope filter, and optional synthesis into one recorded session kept in the corpus history. - Accountable pilot evidence engine.
gather pilot run|refresh|verify|bundledrives a closed manifest through a source-isolated capture into a content-addressed corpus, writes a redacted report and a hash-chained receipt, monitors sources for change (NEW/CHANGED/UNCHANGED with archived history), and packages deterministic shared or full bundles any third party re-verifies offline. See docs/PILOT.md. - Three surfaces, one engine. The full CLI, an MCP stdio server (
gather mcp, toolsgather.status,gather.doctor,gather.docs,gather.arxiv,gather.federation,gather.run,gather.context,gather.pilot), and a plain Python API. - Zero-dependency core, opt-in speed. The core is pure standard library.
gather-engine[fast]adds lxml parsing (roughly 2x on large documents in our own informal timing, unpublished),[browser]adds Playwright JS rendering.gather capsreports what your install can actually do; a missing capability is reported as such, never faked.
The animated explainer walks through a local page extracted into hashed blocks, the grounding check on three proposed records, storage by content hash, corpus verify on corrupt and missing bodies, a scoped run with its digest, and the demo's tampered receipt. Every value on it is output from this repository. Its source is docs/explainer/index.html.
Claiming got cheap. Checking did not. (2 min 57 s, narrated, captioned). Gather keeps a receipt for every source it pulls, so checking a citation stays cheap. The film page carries the transcript, the sources and recall questions.
Video walkthrough: coming with the next release.
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
-
Install. Install from a checkout to run the demo. Python 3.11 or newer; none of this needs the network.
$ git clone https://github.com/HarperZ9/gather && cd gather $ pip install -e . -
First run: the demo. The demo builds a sealed digest of three receipts, then tampers with one and shows the digest no longer verifies.
$ python examples/demo.py witnessed digest: 3 receipts, seal 7da7dc456b11..., verified True after tampering one receipt, digest verifies: False <- caught -
Extract a page. Pull a saved page into structured blocks, each with its source position.
$ gather extract article.html html[1]/body[1]/h1[1] h1 sha256 b09abb7b64e1424e... html[1]/body[1]/p[1] p sha256 1a8f9554b2960300... html[1]/body[1]/p[2] p sha256 95355aaf2715a43d... content_sha256 4c907012d520ed44... markdown_sha256 94812767e4d6dd59... method html-extract -
Store your notes and re-check them. Store a folder by content hash, then verify the corpus later. A clean corpus exits 0.
$ gather docs ./notes --store ./corpus $ gather corpus verify ./corpus
pip install gather-engineThe distribution is gather-engine; it installs the gather command and the gather package (import gather). Python 3.11+. From a source checkout, python -m gather runs the same CLI.
Extract a page to Markdown with a per-block receipt:
gather extract https://example.com/article{
"blocks": [
{"path": "html[1]/body[1]/h1[1]", "sha256": "7e8cd2056da7...", "tag": "h1"},
{"path": "html[1]/body[1]/p[1]", "sha256": "64ec88ca00b2...", "tag": "p"}
],
"content_sha256": "ef4430f5f70f...",
"markdown_sha256": "752cce8836dd...",
"method": "html-extract",
"url": "https://example.com/article"
}Then try the rest of the surface:
gather caps # what this install can do (fast / browser)
gather crawl https://example.com --depth 2 # a hash-chained crawl ledger as JSON
gather arxiv "aperiodic monotile" --store ./corpus
gather scholar "10.1234/monotile" --edges --json
gather docs ./research-notes --scope "rubik,group theory"
gather corpus verify ./corpus # re-hash every stored body; non-zero exit on corruptionOffline demo, no install of extra tools, nothing downloaded:
python examples/demo.py # one video parsed, scoped, digested, then a tampered receipt caught
python examples/pipeline.py # the whole pipeline: run -> store -> verify -> recall, offline
python examples/context_selection.py # inspect verified excerpts and build a private context payloaddemo.py prints each item with its hash and verify=True, then flips one receipt and shows the digest verification fail. The hash prefixes vary; the verify results are pinned by the test suite. For a browser-viewable version of the same proof, open examples/gather-demo.html.
Gather from three different source types into one corpus, then re-check it:
gather web "https://example.com/article" --store ./corpus
gather arxiv "2301.12345" --store ./corpus
gather docs ./notes --scope "monotile,tiling" --store ./corpus
gather corpus list ./corpus # every item with source, method, and hash
gather corpus search ./corpus --terms tiling --method http-get --json
gather corpus verify ./corpus # MATCH per body, non-zero exit if anything is corrupt
gather corpus availability ./corpus # per-record availability with typed outcomes
gather corpus context ./corpus --json # bounded readable excerpts + row refsEvery item carries its source, ref, method, timestamp, and a sha256 of the exact source text. New corpus rows also carry a versioned storage witness for the object bytes written by Gather. verify re-checks the body against the source receipt and, when present, that storage witness; search filters by scope terms, source, kind, or method. context reads and re-hashes bounded inspected or selected bodies through a confined corpus-layout reader, returns bounded excerpts by default, and exports selected ranges only when the caller supplies the current corpus digest.
Readable context verifies the exact source text first. For display and range selection, CRLF or CR line endings are then normalized to LF. START and LIMIT are Python string character offsets over that readable view. sha256, verified_sha256, and source_sha256 bind the exact source text. view_sha256 binds the full LF-normalized readable view, and view_codec names that transformation. The selected slice, range, source refs, source/view hashes, storage status, and omissions are bound into selection_digest.
For rows written before exact UTF-8 storage witnesses existed, Gather attempts a bounded legacy reconstruction: exact UTF-8 bytes or the inverse of the old Windows text writer's LF-to-CRLF expansion. That can reconstruct the source text for an existing receipt, but it does not prove historical raw object-byte integrity. A storage witness makes witness stripping or codec changes detectable only to consumers that pin or independently verify the prior corpus digest with --expect-digest / expected_corpus_digest; recomputing a digest after local catalog mutation is an explicit acceptance of the current catalog state.
The context reader pins the opened corpus root for catalog and body reads, but it does not prove a caller-resolved workspace parent relationship. If a host derives DIR from an approved workspace plus a relative path, that host must bind the workspace-to-corpus resolution before calling Gather.
Plan a fleet of sources before probing any of them:
gather federation validate registry.json --json # check rows against a closed contract
gather federation plan registry.json --json # one deterministic capture plan per source
gather federation policy policy.json --json # audit retry/backoff rules as typed verdicts
gather federation entity entities.json --json # audit entity-resolution matchesAll four run offline; no probe fires. A registry row is a catalog fact and is never reported as coverage, availability, or content.
pip install 'gather-engine[fast]' # lxml, faster parsing on large docs (informal ~2x, unpublished)
pip install 'gather-engine[browser]' # Playwright JS render (then: playwright install chromium)The core never grows a hard dependency. Install a backend and it registers; skip it and the stdlib path or an explicit UNVERIFIABLE result stands in.
from gather.web import WebSource
from gather.digest import digest, verify_digest
items = WebSource().fetch("https://example.com/article")
d = digest(items) # a sealed digest over every item's receipt
assert verify_digest(d) # re-derive the seal; False if anything was alteredThe same seams the CLI uses are importable: gather.extract, gather.crawl, gather.track, gather.schema_extract, gather.stream, gather.search, gather.store, gather.run, gather.context, and the source adapters. gather.context.CorpusRootDescriptor is a Python-only same-process handoff for hosts that already retained and identity-bound a corpus root descriptor; CLI and MCP continue to accept corpus directory strings. Linux/WSL mounts that cannot preserve retained directory authority for confined child opens are rejected fail-closed on each opened corpus, descendant directory, or catalog/body file descriptor rather than read by pathname fallback. ARCHITECTURE.md maps the modules and seams.
The web adapter reads static HTML and does not run JavaScript; a client-rendered page yields its shell, and the method says so. The browser adapter runs a real headless browser and is the most exposed edge: its host guard covers only the first navigation, so do not point it at untrusted URLs where internal services are reachable. The threat model is in ARCHITECTURE.md. Credentials enter only through gather.credentials, read from the environment by name, never logged, never written into a receipt or URL.
- docs/INTRODUCTION.md: what gather is, core concepts, and a first-ten-minutes walkthrough.
- USAGE.md: the full command reference, command by command.
- ARCHITECTURE.md: the design map, seams, and threat model.
- docs/WEB-ENGINE-UPLIFT.md: the web-data engine roadmap and benchmarks.
- docs/ENTERPRISE-READINESS.md: context envelopes, action receipts, and host-neutral operation for unattended agents.
- docs/PILOT.md: the accountable pilot evidence engine, its manifest boundary, and the private/shared evidence split.
- CHANGELOG.md: version history. Current release: 2.0.0.
Peer projects: crucible (judgment), index (code maps), forum (orchestration), telos (the engine).
Research breaks when sources become a blur. One rule sits underneath all of it: every item records how it was obtained. A quote fetched over HTTP, a page rendered in a browser, text recognized from a scan, and a statement synthesized from fragments are all valid items, but they are not equally direct, and the method field plus the digest seal keep that difference on the record and re-checkable. If that discipline matters to your workflow, gather was built for it; if you just need the intake, it never gets in your way.
This tool is one tool in a family that holds a single belief steady across every surface: knowledge open to anyone who can attain the means; acceptance decided by external checks, never reputation; every result re-runnable; honest nulls first-class; ownership earned by comprehension; learning woven into the work. The full text lives in CREDO.md. The long form of this belief: The Unbundling.
Gather is fair-source: open to read, run, and build on, with commercial use reserved so the project can fund its own development. See LICENSE.
Bring papers, transcripts, local docs, or awkward public materials that need provenance before synthesis; source-adapter testing and research-lab feedback are the most useful pressure right now. To develop against a checkout:
python -m pip install -e ".[dev]"
python -m pytest # 650+ tests
python -m ruff check src tests examples
python -m mypyKeep the README, package metadata, and examples aligned with current behavior before opening a PR.
Built by Zain Dana Harper in Seattle: evidence-first tools that leave a re-checkable artifact behind. The full workbench is at Project Telos.
The optional client bundle adds a bounded read-only MCP profile with an explicit workspace selected at launch. See client package setup. Network retrieval is disabled until an exact origin is granted at launch. Allowed GET requests return source text and byte-hash receipts; redirects, ambient proxies, credentials and process execution remain disabled. Windows x64 MCPB and ZIP bundles include their runtime; source plugins require Python 3.11+. The full CLI/MCP retains advanced operations. Client-specific installation and marketplace acceptance require separate verification.
