// 01 — Reading is a transaction, not a scrape.

Live — mark.anywaye.com

The web, in clean markdown, for machines.

Crawl any domain once — Mark stores it as provenanced, licensed markdown and serves it to AI agents, chatbots, and RAG pipelines with a sync contract. No HTML soup, no wasted tokens, no parsing errors.

the door — mark.anywaye.com
❯ curl -s mark.anywaye.com/r/https://www.gov.uk/uk-visa
# 200 · text/markdown · 1,842 words
# license: ogl · provenance: verified · fresh 3m ago
---
title: "UK visa: guide for applicants"
source: www.gov.uk/uk-visa
words: 1842 · chunks: 9 · license: ogl
---
# Apply for a UK visa to work, study, or
# join family. Check which visa you need…
❯ ▍

The door

One URL-addressed entrance for every page.

GET /r/<any-url> serves a page's stored content, or fetches and stores it on demand. One code path for every tier: authentication is optional — a key raises your rate limits and unlocks your own planes — and everything served is metered on the content ledger.

  • Crawl once, store once — no duplicated content
  • Conditional reads: If-None-Match → 304 Not Modified
  • Robots.txt and crawl-delay enforced; domains can opt out entirely
  • Provenance shown on every response
// Conditional read — nothing changed?
GET /r/https://www.gov.uk/uk-visa
If-None-Match: "9f2c1e"
304 Not Modified
# zero tokens spent, sync stays honest

Your own corpus

Connect domains. Upload documents. Stay in charge.

Point Mark at your own domains — set the crawl scope (all pages or explicit paths), choose visibility, schedule refreshes, subscribe webhooks. Upload files or sync a GitHub repo and your documents become a private corpus under the same contract.

  • Visibility per site: shared with the public corpus, or exclusive
  • Weekly refresh sweeps; on-demand refreshes when you need them
  • GitHub sync for docs that live in repos
  • Signed webhooks when content changes
// Connect a domain, your terms
POST /api/v1/sites
{ "url": "https://docs.example.co.uk/" }
PUT /sites/docs.example.co.uk/pages
{ "scope": ["/guides", "/reference"] }
PUT /sites/docs.example.co.uk/visibility
{ "exclusive": true }

Estates

Many domains, one manifest, one diff.

Group your domains into collections — one estate with a single manifest, a single diff since your last sync, and a single search across all of it. Membership is grouping, not claiming: connecting a domain someone else connected keeps you reading the shared corpus with your own keys.

  • Estate-wide manifest and per-estate search
  • Diffs tell you exactly what changed: created, updated, deleted
  • Value comes from the service layer, not from hoarding content
// The estate diff
GET /api/v1/collections/12/manifest
{
  "generated_at": "2026-09-05T09:14:00Z",
  "created": 2, "updated": 14,
  "deleted": 1
}

Search & answers

Retrieval with its manners on.

Search runs full-text and vector retrieval across the plane — license-filtered, scoped to a source, a site, or your whole estate. And when a search isn't enough, ask: a question in, a grounded answer out — one deterministic call over retrieved excerpts, with numbered citations and their provenance.

  • License-filtered: cc, public_domain, licensed
  • Temperature 0 — the same question, the same answer
  • If the corpus doesn't contain it, the answer says so plainly
// A question in, citations out
POST /api/v1/answers
{ "question": "Do I need a visa to work in the UK?",
&nbsp;&nbsp;"collection_id": 12, "license": "ogl" }
→ "Most work visas require sponsorship…"
  [1] gov.uk/uk-visa · ogl
  [2] gov.uk/skilled-worker-visa · ogl

The reading layer is live.

Every request metered, every page provenanced, every answer cited. Get a key and pull your first page in under a minute.

anywaye