PRD — AgentLand (agentland.world)

"AgentLand is a microscaler — the opposite bet from a hyperscaler." Microscaler (n., category we are coining): where a hyperscaler sells infinite anonymous capacity, a microscaler provisions small, governed, fully-metered units of intelligent work — agents — at a scale a human can see, afford, and hold accountable. Hyperscalers sell you the ocean; a microscaler sells you an aquarium where you can watch every fish. AgentLand is the first. Version: Gen-1.4 — microscaler category rebrand (Gen-1.2 critique loops in §16; Gen-1.3 additions in §17) Author: Claude Fable 5 (all planning through Fable per Byron's directive), 2026-07-11 Executor: written for one-shot overnight execution by a Claude Opus 4.8 session with no other context. SPEC ONLY — nothing is built. Domain: agentland.world (owned at GoDaddy, verified 2026-07-11; picked per Byron: "I like the world one").


0. The one-paragraph pitch

AgentLand is a microscaler — a new category — where the primitive isn't a VM, it's an agent. Visitors sign in with Google, get free credits ($5–$25, or redeem a $100 coupon Byron sends), and provision agents the way you'd provision EC2: pick a model mix, set dials for compute / storage / network / context, attach Responsible-AI agents (traffic, police, courts, judges, councils) as first-class infrastructure, wire agents into workflows that execute monetizable transactions, and watch it all live in Mission Control. Everything runs small-by-design inside Byron's Mac Mini with hard cost fences (per-account caps, idle purges, acknowledgment gates), because the product being demonstrated is the shape of the future utility — modularized, metered, governed agents — not raw horsepower. It is simultaneously a working test platform, a Responsible-AI showcase, and Byron's sharpest demo asset.

1. Why (Byron's intent, distilled from his notes)

2. Vocabulary (used consistently everywhere)

Term Meaning
Agent The provisionable unit. Has a model mix, four dials, a role card, credits meter, lifecycle policy.
Model Mix An agent's primary + secondary (+arbiter, see §5) model config from the platform's cheap-model allowlist.
Dials Compute · Storage · Network · Context. Simulated capacity settings with real enforced budgets (§6).
RAI Family Governance agents: Traffic, Police, Court, Judge (LLM-as-judge), Council. Attachable to any org or workflow (§7).
Workflow A wiring of agents (DAG) that processes Transactions.
Transaction A monetizable unit of work entering a workflow (e.g. "what's the best stock for me?" → advice pipeline). The billing atom (§8).
Org A tenant. Google-auth owner + delegated members (§9).
Mission Control The live observability surface for everything (§10).
LandCredits The currency. $1 = 100 LC. All caps, meters, coupons denominated in LC.

3. Non-goals (Gen-1 — say no loudly)

4. Cast of screens (each gets a mockup — the mockup site mirrors this list)

  1. Landing (agentland.world): pitch, 1-minute promo, "Sign in with Google", coupon field.
  2. Onboarding: first-login credit grant ($25 default; $5 configurable; $100 coupon path), pick-your-first-agent wizard.
  3. Provision an Agent: the marquee screen. Role card, model-mix picker, the four dials, lifecycle policy, RAI attachments, price-per-hour estimate updating live.
  4. Agent detail: live meters (LC burn, calls, context fill), logs, pause/destroy, dial adjustments.
  5. Workflow Studio: drag agents onto a canvas, wire them, drop a Transaction template in, run it, watch tokens flow.
  6. Marketplace: launch catalog of agent templates, RAI family front-and-center, install-to-org.
  7. Mission Control: platform-wide live board — every agent, org, transaction, verdicts feed, spend heat, purge queue.
  8. Org & Members: members, delegated sub-logins, roles, coupon management (admin), caps.
  9. Transactions ledger: every transaction with cost breakdown per contributing agent.
  10. Demo mode: the narrated click-through (§12).

5. Model Mix — the feature Byron called out first

6. The Dials — microscale capacity with real enforcement

Four dials per agent, 1–10 scale, each mapping to a REAL enforced budget (simulation with teeth, not theater):

Dial Simulates Actually enforces (Gen-1 mapping)
Compute vCPU class max LLM calls/min + max concurrent calls (1→2/min, 10→30/min)
Storage disk max bytes of memory/artifacts persisted per agent (1→256KB, 10→25MB, SQLite-backed)
Network bandwidth/egress max outbound tool calls (web fetches, inter-agent messages)/hour + payload size cap
Context RAM max context tokens assembled per call (1→2k, 10→24k) + the context organizer setting: what gets packed first (recency / pins / semantic recall), visible as "organizing principles" per Byron's note

7. Responsible AI family — governance as infrastructure (the differentiator)

Attachable governance agents, modeled on Byron's existing Glass Court (policy + Ed25519-signed verdicts, preserved at ~/court-public/ + OB1 postgres):

8. Transactions & monetization

9. Identity, orgs, credits, admin

10. Mission Control

One screen that makes the whole simulated world legible (Byron: "the ability to basically see how anything is working"):

11. Lifecycle & cost gates (the "keep this cheap" machinery)

12. Demo mode, use cases, marketing (launch-blocking, not afterthoughts)

13. Architecture (Gen-1, sized to reality)

14. Build phases with acceptance criteria (for the overnight executor)

15. Risks & honest limits

Risk Position
Small scale mistaken for fakery The category IS the answer: a microscaler is small on purpose. Dials enforce real budgets, every call is really metered, and nothing claims to rent servers.
Abuse of free credits (LLM proxy farming) Cheap-models-only allowlist, org caps, coupon gating, per-IP org-creation limit, global real-dollar ceiling. Worst case exposure = ceiling.
One Mac Mini Fleet ceilings + queueing; it's a demo platform, backpressure is acceptable and even pedagogical (Traffic agents!).
Email deliverability of ack gates Purge never fires on email failure alone (log + Mission Control warning instead); Gmail-API sender is a known-good path.
Scope explosion Non-goals (§3) + the phase gates. The overnight executor builds P0–P1 minimum, P2–P3 only if green.
Fable sunset All planning was Fable (this doc); execution targets Opus 4.8; model strings in code are config, never literals.

16. Critique log (two loops, as ordered)

Loop 1 (found and fixed):

Loop 2 (what I forgot, per Byron's "see what you forgot" order):


17. Gen-1.3 additions — third forget-loop + competitive steal list (2026-07-11, pre-build)

17.A Positioning upgrade (from the landscape scan)

The July-2026 scan validates the premise — AWS AgentCore provisions agents from {model, prompt, tools} via CreateHarness; Anthropic, Microsoft, Google all ship hosted agent runtimes — and reveals the gap AgentLand owns: every real platform buries governance in settings pages and economics in invoices; the only watchable agent worlds (AI Village, Project Vend) have no provisioning console. AgentLand is the fusion: "the microscaler: SimCity for the agent economy" — a legible toy civilization where every token spent, verdict signed, and purge executed is a visible public event, which is simultaneously the sharpest RAI demo in the room. This sentence goes on the landing page.

17.B Invites & People-First communications (Byron's explicit ask, now first-class)

17.C Steal list — 12 features from the landscape, adapted to a $5 microscaler (each cited)

  1. Spectator Mission Control with public reasoning (AI Village, theaidigest.org/village): opt-in "observatory" mode streams agents' live action feed + reasoning summaries + inter-agent chat to viewers. Costs a websocket; makes the sim a show. → P3.
  2. Economy drama: balances, insolvency, public purges (Anthropic Project Vend): agents carry visible balances; hitting zero is a Mission Control EVENT (insolvency → freeze → purge countdown), not a silent state. The purge queue becomes narrative. → P1 (already specced; upgrade: make events public/animated).
  3. Agent Cards at /.well-known/agent-card.json + a registry (A2A protocol / MCP registry): every AgentLand agent serves a real A2A-style agent card; the marketplace doubles as a registry API. Static JSON — nearly free — and makes "A2A-discoverable, MCP-registered" literally true. → P2.
  4. Agent citizenship (Microsoft Entra Agent ID): every agent has a directory identity with an owner/sponsor and an org-chart slot. Gives Police/Court someone to subpoena; lifecycle policies ARE purge gates, now with a name on them. → P1 (fold into agent schema).
  5. Split-line billing: tokens + session-hours (Claude Managed Agents, $0.08/session-hour precedent): the cost waterfall itemizes tokens + runtime-hold + storage + network as separate meters, so the $5 decays visibly across four lines. → P0 gateway (meter design).
  6. 402 payments + signed mandates (Google AP2, Coinbase x402): sim services reply HTTP 402 "Payment Required" in LC; agents pay and retry; every agent purchase requires a signed spending mandate the Court audits against the cap. Teaches the real emerging standards with zero real money. → P2.
  7. Command-center IA + OTel traces (Salesforce Agentforce): Mission Control = health tiles → drill into per-session span traces (every reasoning step, tool call, guardrail check as spans). Emit actual OTel-shaped JSON for standards credibility. → P2/P3.
  8. CreateHarness-shaped provisioning + marketplace listing schema (AWS AgentCore + AWS Marketplace agents category): the provision API takes {model_mix, charter, tools, dials} exactly harness-style; marketplace listings carry the AWS-style schema (pricing model, endpoint, category). → P1/P2, pure CRUD.
  9. Model compare grid with lanes (OpenRouter): the model-mix picker shows $/Mtok, latency, context per allowlisted model, with a "commodity vs premium lane" toggle — budget as a strategic choice. → P1 UI.
  10. Online judge + precedent promotion (Braintrust): Judge agents score live traces; a conviction promotes the trace to a regression dataset ("precedent") that future agent versions are tested against; repeat offenders escalate to Court → purge. Courts with mechanics, not vibes. → P2.
  11. Inspector-General meta-agent (LangSmith Insights Agent): one cheap scheduled agent reads the fleet's ledgers/traces daily and publishes "The State of AgentLand" brief (Mission Control panel + email digest + social-ready snippet). Governance content that markets itself. → P3.
  12. Guardrails as visible pipeline nodes (OpenAI AgentKit UX — platform sunsetting, insight stands): render Traffic/Police checks as literal wired nodes in each agent's pipeline view, so trust is something you can SEE. → P2 Workflow Studio.

17.D Loop-3 items not covered above

17.E The AWS double-bonus (all $0, verified 2026-07-11)

17.F Workflow Studio interaction model + Agent Definition v1 (Byron's design review, 2026-07-11)

Designer (Workflow Studio) — committed interaction model:

Agent Definition v1 — the quality schema (replaces the thin charter-only definition):

  1. Instruction stack (layered, separately editable, composed in order): platform base → org policy → agent charter → task prompt. Inspector shows all four layers.
  2. Exemplars: 2–5 input→ideal-output pairs stored on the definition; injected as few-shot. The single cheapest quality lever; the UI nags if empty.
  3. Knowledge attachments: docs/URLs linked to the agent; chunked and stored under the Storage dial's byte cap; the Context dial's organizer (recency/pins/semantic) decides packing. Attachment list visible on the agent card.
  4. Output contract: JSON schema / format spec; Judge auto-validates every output against it; violations are RAI events.
  5. Attached evals: test cases incl. court-precedent promotions (17.C-10); any definition edit re-runs them; failing evals block deployment of that version to workflows.
  6. Versioning: immutable definition versions; workflows pin; diffs visible; rollback one click. Marketplace templates ship with all six fields populated — that's what makes a template worth installing.

17.G Memory layer (Byron's design review round 2)

Three tiers per agent, all governed by the existing dials — memory is where agent quality compounds:

17.H Professional roster + palette taxonomy (the toolbox)

Roster — Gen-1 marketplace "Staff" templates, every one shipped with ALL SIX Agent Definition fields populated (instruction stack incl. profession-baseline knowledge + norms, exemplars, starter knowledge attachments, output contract, evals, v1): Legal: Lawyer (contract/risk review; output contract = structured risk memo; hard rule: flags, never approves — a Human Gate follows it by default). Engineering: Developer (feasibility, plan, code review), CTO (architecture/tradeoff review), CISO (threat model, security review; evals include an injection-resistance case). Product: Product Manager (requirements, prioritization), Technical Writer/Documentarian (docs from artifacts), Designer (UX critique). People: HR Manager (policy-aware, escalates rather than decides on protected matters — Police-attached by default). Finance: Financial Analyst (models, budgets; advice-banner enforced). GTM: Marketer, Sales/SDR, Support Agent. Ops: Researcher/Analyst, Executive Assistant. Each template names its default model lane (most = commodity/haiku-flash; CTO/Lawyer default second-opinion mode). Palette taxonomy (easy-to-use is the requirement): left palette groups = My Agents · Departments (Legal / Engineering / Product / People / Finance / GTM / Ops — expandable) · RAI Family · Primitives, with search-as-you-type and favorites. Department bundle drops: dragging a whole department chip drops a pre-wired mini-workflow (e.g. Engineering = Developer→CTO review chain with a Judge attached). Hovering any palette item shows its definition card (model lane, LC/hr, output contract) before you commit.

17.I Definition-schema smoke test (executed pre-build — results recorded here)

Byron's order: prove a complete workflow actually works before building the platform. A standalone harness (~/agentland/smoke/) composes real Agent Definitions (layered instructions + exemplars + output contracts) for PM → Developer → Lawyer → CISO on the real cheap-model allowlist (haiku / gemini-flash), runs a launch-review transaction through them with a Judge validating each output contract, and produces the final memo + a token/LC cost waterfall. Pass = every hop honors its output contract and the Judge scores ≥ threshold; the run's artifacts and costs are committed alongside the seeds as the platform's first regression fixture. RESULTS (executed 2026-07-11, real API calls): PASS. PM → Developer → Lawyer → CISO on alternating claude-haiku-4-5 / gemini-2.5-flash; Judge scores 9/8/9/9, overall 8.75/10 (threshold 7); all contracts honored ≤1 retry; total transaction cost 11.56 LC ($0.1156) with a full per-hop waterfall. Quality proof: the haiku-powered Lawyer flagged that the PM's "warn + log" MNPI feature manufactures negligence evidence (critical severity) — charter-true reasoning from a commodity model. Schema lessons folded back into this spec: (1) output contracts MUST carry upper bounds (maxItems/maxLength) as a platform norm — an unbounded findings array made run 1 blow the token cap and truncate JSON; (2) enum restrictions must be restated in the task prompt (cross-provider enum drift) and linted roster-wide in seed CI; (3) judges need schema-valid scores for null artifacts. Artifacts: ~/agentland/seeds/*.json (all 14 roles, exemplars machine-validated), ~/agentland/smoke/run-workflow.mjs, results-2026-07-11.json (the platform's first regression fixture), archived failing run, and REPORT.md with the full memo chain.

17.J Reference integration cases (launch-blocking demo content — Byron, 2026-07-11)

Two presentation-grade reference cases, each an animated auto-advancing walkthrough (same treatment as Demo mode §12: chapters, canned data, optional cloned-voice narration, click-through or autopilot), shipped as chapters in the demo player + an "Integrations" row in the Marketplace + a docs page each. These exist so AgentLand slots directly into Byron's hero project and Claude Code project as reusable set-pieces.

RC-1 — Claude Code in the workflow ("the terminal joins the org chart"):

RC-2 — Kiro in the workflow ("spec in, software out, governed"):

Acceptance (P3): both cases run hands-free in the demo player using genuinely recorded sessions (record one real Claude Code run and one real Kiro run; replay — never fake terminal output); each has a "wire this yourself" docs page (RC-1: the MCP endpoint config; RC-2: the spec handoff format); Marketplace lists both as installable integration templates.

17.J-bis — AgentLand-as-a-Service: the consumption model (Byron's key framing: "we are the service — consume us like AWS/Azure/GCP"). The reference cases are two-way; the durable architecture is that ANY harness consumes AgentLand's agents, outputs, and workflows as a cloud service:

17.K Labs: The Observatory — 3D deck (EXPERIMENT, demoted by Byron 2026-07-12)

Status: Labs experiment, not a core screen or launch requirement. Byron's verdict after using it: fun, not usable as a working surface. It stays reachable via a small "🌐 labs" globe on the landing page; the flat Mission Control (§10) is the real observability surface. Everything below is retained as the experiment's design record; build effort on it only ever comes after P3, if at all. A cinematic 3D rendering of Mission Control (new mockup screen #12, observatory.html; platform feature in P3 as the premium view over the same SSE feed):

17.L Phase impact

P0 unchanged + meter design (17.C-5) + Strands as the agent-loop library (17.E). P1 gains citizenship fields, compare-grid picker, public insolvency events. P2 gains agent cards/registry, 402+mandates, span traces, precedent system, pipeline-node guardrails, listing schema. P3 gains observatory mode, Inspector-General, playground, status/docs/pricing pages, the credibility strip, the two reference integration cases (17.J), the as-a-service surface + Connect screen (17.J-bis; API keys land in P0 with auth), and the Observatory stays a Labs experiment (17.K), not a launch item. Nothing here changes the Gen-1 cost posture: every addition is CRUD, static JSON, websockets, replayed recordings, or one cheap scheduled agent.

17.M Hyperscaler-parity check (Byron's "are we missing essentials?" — 2026-07-12)

Parity review against AWS/Azure/GCP surface areas. Covered already: IAM→roles+scoped keys, billing→LC ledger, quotas→caps/dials, marketplace, observability→Mission Control+OTel, compliance→RAI verdicts+export/delete, status/docs (17.D). Deliberately absent BY CATEGORY (not gaps): regions/AZs, GPU scale, enterprise SLAs — "one Mac Mini" is the thesis; the status page pledges honesty, not nines. THREE REAL GAPS, all adopted as requirements:

  1. Backup/DR (P1, ops-blocking): nightly SQLite snapshot (sqlite3 .backup) to a dated file + weekly restore test; retention 14 days; a platform holding other people's orgs/credits without backups is negligent.
  2. Admin audit trail (P1): every admin action (cap change, coupon mint/revoke, allowlist edit, key revoke, freeze/unfreeze) writes an audit row (who/what/before/after) — visible to the org affected and in Mission Control. A governance-first platform governs its own admins.
  3. Budget threshold alerts (P2): email at 50/80/100% of org cap via the People-First comms register — warn before the wall, don't just freeze at it. Nice-to-have noted, not required: declarative org config export ("agentland.json" as toy IaC).

Companion mockups: agentland.world (static, deployed alongside this PRD). Build trigger: Byron's explicit go — then P0 begins overnight on Opus 4.8.

AgentLand PRD Gen-1.3 — spec only; nothing is built yet.Back to landing
Mockup — Gen-1 preview