The door
One URL-addressed entrance for every page.
GET /r/<any-url> serves a page's stored content, or fetches and stores it on demand. One code path for every tier: authentication is optional — a key raises your rate limits and unlocks your own planes — and everything served is metered on the content ledger.
- Crawl once, store once — no duplicated content
- Conditional reads:
If-None-Match→304 Not Modified - Robots.txt and crawl-delay enforced; domains can opt out entirely
- Provenance shown on every response
Your own corpus
Connect domains. Upload documents. Stay in charge.
Point Mark at your own domains — set the crawl scope (all pages or explicit paths), choose visibility, schedule refreshes, subscribe webhooks. Upload files or sync a GitHub repo and your documents become a private corpus under the same contract.
- Visibility per site: shared with the public corpus, or exclusive
- Weekly refresh sweeps; on-demand refreshes when you need them
- GitHub sync for docs that live in repos
- Signed webhooks when content changes
Estates
Many domains, one manifest, one diff.
Group your domains into collections — one estate with a single manifest, a single diff since your last sync, and a single search across all of it. Membership is grouping, not claiming: connecting a domain someone else connected keeps you reading the shared corpus with your own keys.
- Estate-wide manifest and per-estate search
- Diffs tell you exactly what changed: created, updated, deleted
- Value comes from the service layer, not from hoarding content
Search & answers
Retrieval with its manners on.
Search runs full-text and vector retrieval across the plane — license-filtered, scoped to a source, a site, or your whole estate. And when a search isn't enough, ask: a question in, a grounded answer out — one deterministic call over retrieved excerpts, with numbered citations and their provenance.
- License-filtered:
cc,public_domain,licensed - Temperature 0 — the same question, the same answer
- If the corpus doesn't contain it, the answer says so plainly
The reading layer is live.
Every request metered, every page provenanced, every answer cited. Get a key and pull your first page in under a minute.