PIBotBeta← protocolized.io

How PIBot Works

Protocol Institute Oracle — technical reference

What is PIBot?

PIBot is the Protocol Institute’s knowledge infrastructure — a retrieval-augmented generation (RAG) system that answers questions grounded in the PI corpus. It is not a fine-tuned model. The underlying language model is Claude Sonnet; what makes it PI-specific is the corpus material injected as retrieval context at query time and a system prompt encoding PI’s intellectual commitments, scope, and vocabulary.

PIBot is the Protocol Institute’s research assistant: it answers from the Institute’s own corpus and cites what it draws on.

Two voices, one corpus. PIBot serves two audiences with distinct response styles selected by a context field in the request body:

The corpus

SourceScaleCoverage
Summer of Protocols PDFs82 docs · 750 vectorsResearch papers, theoretical essays, protocol fiction, game materials (2023–2024). 11 cover letters / title pages deprecated from retrieval.
Protocolized Substack118+ posts · 1,080 vectorsFictions (58), Articles (47), Obliquities (5+); 38 author profiles; 13 collection cards. Synced daily via GitHub Actions.
Protocol Institute YouTube91 talks · 2,940 vectorsGuest Talks (34), Town Halls (20), Protocol School 2025 (13), Researcher Salons (9), Symposium 2024 (7), Bridge Atlas (5). English captions via yt-dlp.
Bibliography270+ refs · 278 vectorsExternal works cited by PI corpus; scored 0–3 for protocol relevance; abstracts + OA PDFs where available.
Discord community5,580 vectors12 general + forum channels including #idle-protocol-musings, #protocol-watch, #field-reports, #reading-room; threaded exchanges; starred highlights weighted 1.0×.
SIG meeting archives5,346 vectorsSix Special Interest Groups: Formal Protocol Theory (SIGFPT), Memory Research Group (MRG), Protocols for Business (SIGPfB), Protocol Fiction (ProtFiSIG), Psychohistory (SIGPSY), Distributed Robotics (DRG). Includes AI-generated meeting summaries, discussion threads, and 91 published meeting pages from protocol-institute.org.
Community-shared links9,674 vectorsExternal URLs linked in Discord/SIG messages; fetched server-side, chunked, scored 0–3 for protocol relevance by Claude Haiku; score-0 entries deleted. source_count tracks how many Discord messages referenced each URL.
PI lexicon560 vectors914 terms extracted from the PI corpus via Claude Haiku; triage a+b ingested. PI-coined vocabulary: “hardness”, “Whitehead advance”, “protocol tai chi”, and 40+ others also injected directly into the system prompt.
Discord channel guide78 vectorsAll active guild channels with Haiku-generated blurbs, SIG meeting cadence, and next scheduled event time. Used only for introductions and navigation queries — not corpus RAG.
Devlog32 vectorsPIBot’s own development log (one vector per session). Queried at top-3 alongside all other namespaces so PIBot can answer questions about its own architecture and history.
Conversation memory23 vectors (growing)Spooled Discord bot conversations and submitted public web chats, re-ingested as discord_conversation and web_conversation chunk types. High-scoring matches surface as “Similar conversation” links rather than regular sources.

Total: ~26,300 vectors across 11 Pinecone namespaces. The corpus is live: Discord and SIG channels sync incrementally every 30 minutes via a local daemon; Substack syncs daily via GitHub Actions; the lexicon and PDFs are updated manually when the source set changes.

Ingest pipeline pattern

Every corpus source follows a three-layer pattern. Deviating requires explicit justification. This pattern is what makes retrieval work well — it is worth reproducing exactly when adding new sources.

Layer 1 — Haiku enrichment

One Claude Haiku call per document. Input: title + authors + type/tags + first ~1,500 chars of text. Output saved to sources/<type>/enriched_meta.json keyed by document ID:

{
  "summary":        "Two concrete sentences about the specific argument or contribution.",
  "categories":     ["protocol-theory", "governance"],
  "primary_author": "Author Name",
  "all_authors":    ["Author Name"]
}

Shared category vocabulary (used across all sources for cross-corpus query consistency): protocol-fiction · protocol-theory · protocol-watching · editorial · research-report · technology-ai · governance · announcement · interview · memory-archival · organizations. Idempotent; skips existing entries unless --force is passed.

Layer 2 — Body chunks

Text chunked at 512-token windows with 64-token overlap. The text sent to Voyage for embedding is prefixed with title and summary information — anchoring the chunk to its topic even when topic keywords don’t appear in the chunk itself. The stored display text is clean (no prefix). Vector IDs are SHA-256 hashes of raw chunk text, so duplicate passages across documents are automatically deduplicated.

# Embedding input (not stored):
"Title: {title}\nAuthors: {authors}\nType: {type}\nSummary: {summary}\n\n{chunk_text}"

# Stored in Pinecone metadata:
{ text: "{chunk_text}", title, authors, date, url, chunk_type, source, ... }

Layer 3 — Document summary vector

One additional vector per document for “what is this document about” queries. Vector ID: {id}__doc_summary (or {slug}__post_summary for Substack). When a summary vector ranks in the top results, the worker fires a follow-up filtered query to fetch real body chunks from the same document — giving Claude actual prose rather than an abstract. Summary hits are replaced in the result set by their body-chunk siblings before the LLM sees them.

Checklist for adding a new source

To reproduce this pattern for a new corpus source: (1) write an enrichment script that calls Claude Haiku and saves to sources/<type>/enriched_meta.json; (2) write an ingest script that chunks, prefixes, embeds via Voyage, and upserts to a named Pinecone namespace; (3) create an incremental state file keyed by document ID or timestamp watermark; (4) add a normalize<Type>() function in api/worker.js that maps raw Pinecone metadata to the standard source object shape; (5) add the namespace to mergeResults() with an appropriate tier weight; (6) register the source in config/source_registry.json and config/corpus_map.json; (7) add a daemon step in bin/daemon.py.

The Pinecone index

Single index c3po — 1,024 dimensions, cosine metric, serverless (AWS us-east-1), Protocol Institute org account. All corpus namespaces queried in parallel on every request; results merged and tier-weighted before passing to Claude.

NamespaceVectorsRetrieval weightNotes
pdfs7501.0×Body chunks + doc_summary; 82 documents; 11 deprecated cover letters excluded from intro recs
substack1,0801.0×Body chunks, post_summary, collection_card, author_profile
definitions5601.0×PI lexicon terms; vector IDs: lexicon__{term_slug}__{source_slug}
videos2,9400.90×Body chunks + video_summary; 91 YouTube talks
sig5,3460.90×chunk_types: sig_meeting_summary, sig_meeting_body, sig_discussion, sig_message, sig_reply, sig_meeting_page; 6 SIG channels; sig_meeting_page chunks fire a parallel filtered query to ensure meeting pages surface even below TOP_K_EACH
bibliography2780.85×ref_summary + body chunks; relevance score (0–3) stored in metadata
discord5,580starred 1.0× · unstarred 0.70×message, thread, forum_post chunk types; star_count in metadata drives weight split
transcripts230.85× base; tiered boostdiscord_conversation + web_conversation; scores ≥0.60 → 1.10× boost and surfaced as “Similar conversation” link rather than source
discord_links9,6740.55–0.75×Haiku relevance score (1–3) × source popularity bonus; score-0 entries deleted from index
discord_guide78nav/intro onlyChannel blurbs, SIG cadence, next event time; not included in corpus RAG; used only for “where should I post about X?” nav queries and #introductions channel recommendations
meta32top-3 alwaysPIBot devlog sessions; always retrieved at top-3 alongside all other namespaces so PIBot can answer questions about its own build history
Secondary retrieval: when a doc_summary, post_summary, or video_summary vector ranks in the top results, the worker fires a follow-up filtered query (chunk_type ≠ *_summary AND title = {matched_title}) to fetch the actual body chunks from that document. Summary hits are replaced in the result set by their body-chunk siblings before Claude sees them — so Claude always reads prose, not abstracts.

Query pipeline (step by step)

Every POST /query, ask_pibot MCP call, and Discord @mention follows the same pipeline:

  1. Embed the query — Voyage AI voyage-3 encodes the question into a 1,024-dim vector (input_type: "query").
  2. Parallel namespace query — all 9 corpus namespaces queried simultaneously at top_k = TOP_K_EACH (typically 5); meta queried at top-3; transcripts queried at top-3; sig_meeting_page sub-query fired in parallel for the sig namespace.
  3. Normalize — each namespace has a normalize*() function mapping raw Pinecone metadata to a standard shape: { source, type, label, title, authors, date, url, summary, excerpt, score, ... }. Namespace-specific fields (e.g. sig_display, channel_name, domain) are preserved alongside.
  4. Merge and tier-weight — mergeResults() applies the tier weights above, deduplicates by URL, extracts transcript cache hits (score ≥ 0.52), and returns the top MAX_SOURCES items.
  5. Secondary retrieval — summary-type vectors in the top results trigger follow-up body-chunk queries for their documents. Results merged back in.
  6. Build context block — buildContextBlock() formats each source as a labeled excerpt block for the Claude prompt: [PDF — "Title" — Author(s) — YYYY]\n{excerpt}, with namespace-appropriate prefixes.
  7. Claude Sonnet — system prompt cached with cache_control: ephemeral; user message is "Question: {q}\n\nRelevant archive excerpts:\n\n{context}". Discord requests use the DISCORD_SYSTEM_PROMPT variant. History array passed for multi-turn sessions.
  8. Return — response includes answer, sources (with metadata for badge/link rendering), and cache_hits (high-scoring transcript matches surfaced separately).

The language model and system prompt

Claude Sonnet is used throughout — query answering, document enrichment, meeting summarization, link relevance scoring (Haiku for the latter two). The research material is dense and cross-disciplinary; strong synthesis matters more than fast extraction.

System prompt structure. The system prompt is an inline document (~1,100 tokens) in api/worker.js, maintained alongside the retrieval code. It is not loaded from a file at runtime. It was originally inspired by SOUL.md (the conceptual identity document in the repo root), but has since evolved independently and is now substantially more detailed. An agent reproducing this pattern should treat the inline prompt as the source of truth. It contains seven sections:

  1. Role and mission — who PIBot is; corpus-grounded research assistant, not a general chatbot.
  2. About the Protocol Institute — SoP provenance, current programs, leadership, independence. Prevents hallucinated org facts.
  3. Scope declaration — 11 topic areas explicitly in scope; 4 categories explicitly out of scope. Prevents false denials (“I don’t know about that”) for topics that are in the corpus.
  4. Intellectual commitments — 9 substantive PI analytical stances (e.g. “hardness is a design variable”, “context tank not think tank”). Shapes how Claude frames answers.
  5. Voice, analytical moves, format — scholarly but accessible; specific about sources; 5 characteristic analytical moves (cross-domain comparison, hardness analysis, historical situating, formalization ladder, stakeholder analysis).
  6. Indexed corpus map — lists the key named papers, magazine, YouTube series, and SIG groups by name so Claude never falsely denies having them.
  7. Protocol lexicon — 40+ PI-coined or PI-specific terms with compact definitions injected inline. Prevents chunk-boundary definition splits and ensures PI vocabulary is used correctly even when a term doesn’t appear verbatim in retrieved chunks.

The Discord variant appends a DISCORD VOICE OVERRIDE block that supersedes the voice and length instructions. Both variants exceed the 1,024-token Anthropic prompt-cache threshold and cache independently (cache_control: ephemeral).

Rate limits: 20 queries per IP per hour via the web UI. After 8 turns, the conversation can be downloaded as Markdown and continued in Claude, or accessed without a turn limit via MCP.

Delivery interfaces

Three interfaces share the same corpus, query engine, and Cloudflare Worker. They differ in voice, turn model, and how they reach users.

Web UI — pibot.protocolized.io

Browser-based chat at the root URL. Full research-librarian voice. 8-turn session limit (download as Markdown to continue). Submitted conversations can be shared publicly or kept private; public submissions are indexed into the transcripts namespace for self-memory. The /chats route is a public transcript browser.

Discord bot — PIBot

Gateway bot (discord.py, WebSocket) running on the host machine under launchd (org.protocol-institute.c3po-bot). All bot requests pass context: "discord" to the Worker, selecting the 2–3 sentence office-manager response style.

Completed bot conversations are spooled to data/spool/bot_conversations/; the listener daemon picks them up each cycle and ingests them into the transcripts namespace.

MCP server — /mcp

PIBot is available as a Model Context Protocol server (JSON-RPC 2.0 + Streamable HTTP) at https://pibot.protocolized.io/mcp. Connect it to Claude Code or Claude Desktop to query the corpus directly inside your AI client — no turn limit, no browser required.

ToolWhat it doesAuthLimit
search_corpusSemantic search — returns ranked excerpts with metadata and URLs; no LLM call. Filter by namespace: pdfs, substack, videos, bibliography, discord, sig, discord_links, definitions, or all. Result limit 1–20 (default 10).None100 calls/IP/day
ask_pibotFull RAG: embed → retrieve → Claude Sonnet synthesis. Accepts history array for multi-turn sessions. Web voice (not Discord office-manager).Bearer tokenCircuit-breaker shared with web UI

search_corpus is open — no key required. Good for agentic workflows that need raw retrieval without LLM cost or turn limits.

ask_pibot requires a Bearer token (each call invokes Claude Sonnet and Voyage AI at real cost). To request access email team@protocol-institute.org.

Claude Code

Search only (no key needed) — run once in your terminal:

claude mcp add pibot --transport http https://pibot.protocolized.io/mcp

Full access with Bearer token:

claude mcp add pibot --transport http https://pibot.protocolized.io/mcp \
  --header "Authorization: Bearer <your-key>"

Claude Desktop

Add to claude_desktop_config.json (on Mac: ~/Library/Application Support/Claude/):

{"mcpServers": {"pibot": {
  "type": "http",
  "url": "https://pibot.protocolized.io/mcp",
  "headers": {"Authorization": "Bearer <your-key>"}
}}}

For search-only without auth, omit the headers key.

Multi-turn conversations via MCP: ask_pibot accepts a history array of {"role": "user"|"assistant", "content": "..."} objects alongside your question. Pass prior turns to maintain context. The same hourly and daily circuit breakers that govern the web UI apply — calls return an error if the budget is exhausted and auto-reset at the next hour or midnight PT.

Infrastructure

ComponentTechnology
Cloudflare WorkerSingle V8 isolate at pibot.protocolized.io serving web UI, RAG API, MCP server, and Discord Interactions endpoint. PI org account (7e8c7969b2464d23795c555bc6a32af8).
PineconeIndex c3po — 1,024d cosine, serverless aws/us-east-1, PI org account. 11 namespaces, ~26,300 vectors.
Voyage AIModel voyage-3 (1,024d). PI org account. Same model for ingest and query.
Cloudflare KVRate limiting (20 web queries/IP/hour; 100 MCP search calls/IP/day); circuit breaker flag; transcript storage (90-day TTL); usage accumulators.
Cloudflare Queuec3po-oracle queue for Discord slash command deferred responses.
Circuit breakerKV flag + hourly cron; sleeps when hourly spend exceeds $4 or daily spend exceeds $30. Auto-resets at next hour / midnight PT.
Ingest daemonbin/daemon.py — 14-step sync cycle every 30 minutes, launchd-managed (org.protocol-institute.c3po.daily). Logs to ~/Library/Logs/c3po/daemon.log.
Discord bot processbin/c3po_bot.py — discord.py WebSocket gateway, launchd-managed with KeepAlive (org.protocol-institute.c3po-bot). Logs to ~/Library/Logs/c3po/c3po_bot.log.
Substack syncGitHub Actions workflow (.github/workflows/sync-substack.yml) — daily at 08:00 UTC; commits state files back to repo.
Source codeProtocol-Institute/c3po (transferred from vgururao/c3po on 2026-05-31).
PIBot is in active development. The corpus, weights, system prompt, and delivery interfaces expand continuously. Development is logged publicly at protocolized.io/resources/c3po-devlog and tracked in ARCHITECTURE.md in the repository.