Protocol Institute Oracle — technical reference
PIBot is the Protocol Institute’s knowledge infrastructure — a retrieval-augmented generation (RAG) system that answers questions grounded in the PI corpus. It is not a fine-tuned model. The underlying language model is Claude Sonnet; what makes it PI-specific is the corpus material injected as retrieval context at query time and a system prompt encoding PI’s intellectual commitments, scope, and vocabulary.
PIBot is the Protocol Institute’s research assistant: it answers from the Institute’s own corpus and cites what it draws on.
Two voices, one corpus. PIBot serves two audiences with distinct response styles selected by a context field in the request body:
context absent): research-librarian voice — dense, source-specific, 3–5 paragraphs. For researchers and extended sessions.context: "discord"): office-manager voice — 2–3 sentences, one named resource, plain prose. A DISCORD VOICE OVERRIDE block appended to the base system prompt supersedes all length and format instructions.| Source | Scale | Coverage |
|---|---|---|
| Summer of Protocols PDFs | 82 docs · 750 vectors | Research papers, theoretical essays, protocol fiction, game materials (2023–2024). 11 cover letters / title pages deprecated from retrieval. |
| Protocolized Substack | 118+ posts · 1,080 vectors | Fictions (58), Articles (47), Obliquities (5+); 38 author profiles; 13 collection cards. Synced daily via GitHub Actions. |
| Protocol Institute YouTube | 91 talks · 2,940 vectors | Guest Talks (34), Town Halls (20), Protocol School 2025 (13), Researcher Salons (9), Symposium 2024 (7), Bridge Atlas (5). English captions via yt-dlp. |
| Bibliography | 270+ refs · 278 vectors | External works cited by PI corpus; scored 0–3 for protocol relevance; abstracts + OA PDFs where available. |
| Discord community | 5,580 vectors | 12 general + forum channels including #idle-protocol-musings, #protocol-watch, #field-reports, #reading-room; threaded exchanges; starred highlights weighted 1.0×. |
| SIG meeting archives | 5,346 vectors | Six Special Interest Groups: Formal Protocol Theory (SIGFPT), Memory Research Group (MRG), Protocols for Business (SIGPfB), Protocol Fiction (ProtFiSIG), Psychohistory (SIGPSY), Distributed Robotics (DRG). Includes AI-generated meeting summaries, discussion threads, and 91 published meeting pages from protocol-institute.org. |
| Community-shared links | 9,674 vectors | External URLs linked in Discord/SIG messages; fetched server-side, chunked, scored 0–3 for protocol relevance by Claude Haiku; score-0 entries deleted. source_count tracks how many Discord messages referenced each URL. |
| PI lexicon | 560 vectors | 914 terms extracted from the PI corpus via Claude Haiku; triage a+b ingested. PI-coined vocabulary: “hardness”, “Whitehead advance”, “protocol tai chi”, and 40+ others also injected directly into the system prompt. |
| Discord channel guide | 78 vectors | All active guild channels with Haiku-generated blurbs, SIG meeting cadence, and next scheduled event time. Used only for introductions and navigation queries — not corpus RAG. |
| Devlog | 32 vectors | PIBot’s own development log (one vector per session). Queried at top-3 alongside all other namespaces so PIBot can answer questions about its own architecture and history. |
| Conversation memory | 23 vectors (growing) | Spooled Discord bot conversations and submitted public web chats, re-ingested as discord_conversation and web_conversation chunk types. High-scoring matches surface as “Similar conversation” links rather than regular sources. |
Total: ~26,300 vectors across 11 Pinecone namespaces. The corpus is live: Discord and SIG channels sync incrementally every 30 minutes via a local daemon; Substack syncs daily via GitHub Actions; the lexicon and PDFs are updated manually when the source set changes.
Every corpus source follows a three-layer pattern. Deviating requires explicit justification. This pattern is what makes retrieval work well — it is worth reproducing exactly when adding new sources.
One Claude Haiku call per document. Input: title + authors + type/tags + first ~1,500 chars of text. Output saved to sources/<type>/enriched_meta.json keyed by document ID:
{
"summary": "Two concrete sentences about the specific argument or contribution.",
"categories": ["protocol-theory", "governance"],
"primary_author": "Author Name",
"all_authors": ["Author Name"]
}
Shared category vocabulary (used across all sources for cross-corpus query consistency): protocol-fiction · protocol-theory · protocol-watching · editorial · research-report · technology-ai · governance · announcement · interview · memory-archival · organizations. Idempotent; skips existing entries unless --force is passed.
Text chunked at 512-token windows with 64-token overlap. The text sent to Voyage for embedding is prefixed with title and summary information — anchoring the chunk to its topic even when topic keywords don’t appear in the chunk itself. The stored display text is clean (no prefix). Vector IDs are SHA-256 hashes of raw chunk text, so duplicate passages across documents are automatically deduplicated.
# Embedding input (not stored):
"Title: {title}\nAuthors: {authors}\nType: {type}\nSummary: {summary}\n\n{chunk_text}"
# Stored in Pinecone metadata:
{ text: "{chunk_text}", title, authors, date, url, chunk_type, source, ... }
One additional vector per document for “what is this document about” queries. Vector ID: {id}__doc_summary (or {slug}__post_summary for Substack). When a summary vector ranks in the top results, the worker fires a follow-up filtered query to fetch real body chunks from the same document — giving Claude actual prose rather than an abstract. Summary hits are replaced in the result set by their body-chunk siblings before the LLM sees them.
To reproduce this pattern for a new corpus source: (1) write an enrichment script that calls Claude Haiku and saves to sources/<type>/enriched_meta.json; (2) write an ingest script that chunks, prefixes, embeds via Voyage, and upserts to a named Pinecone namespace; (3) create an incremental state file keyed by document ID or timestamp watermark; (4) add a normalize<Type>() function in api/worker.js that maps raw Pinecone metadata to the standard source object shape; (5) add the namespace to mergeResults() with an appropriate tier weight; (6) register the source in config/source_registry.json and config/corpus_map.json; (7) add a daemon step in bin/daemon.py.
Single index c3po — 1,024 dimensions, cosine metric, serverless (AWS us-east-1), Protocol Institute org account. All corpus namespaces queried in parallel on every request; results merged and tier-weighted before passing to Claude.
| Namespace | Vectors | Retrieval weight | Notes |
|---|---|---|---|
pdfs | 750 | 1.0× | Body chunks + doc_summary; 82 documents; 11 deprecated cover letters excluded from intro recs |
substack | 1,080 | 1.0× | Body chunks, post_summary, collection_card, author_profile |
definitions | 560 | 1.0× | PI lexicon terms; vector IDs: lexicon__{term_slug}__{source_slug} |
videos | 2,940 | 0.90× | Body chunks + video_summary; 91 YouTube talks |
sig | 5,346 | 0.90× | chunk_types: sig_meeting_summary, sig_meeting_body, sig_discussion, sig_message, sig_reply, sig_meeting_page; 6 SIG channels; sig_meeting_page chunks fire a parallel filtered query to ensure meeting pages surface even below TOP_K_EACH |
bibliography | 278 | 0.85× | ref_summary + body chunks; relevance score (0–3) stored in metadata |
discord | 5,580 | starred 1.0× · unstarred 0.70× | message, thread, forum_post chunk types; star_count in metadata drives weight split |
transcripts | 23 | 0.85× base; tiered boost | discord_conversation + web_conversation; scores ≥0.60 → 1.10× boost and surfaced as “Similar conversation” link rather than source |
discord_links | 9,674 | 0.55–0.75× | Haiku relevance score (1–3) × source popularity bonus; score-0 entries deleted from index |
discord_guide | 78 | nav/intro only | Channel blurbs, SIG cadence, next event time; not included in corpus RAG; used only for “where should I post about X?” nav queries and #introductions channel recommendations |
meta | 32 | top-3 always | PIBot devlog sessions; always retrieved at top-3 alongside all other namespaces so PIBot can answer questions about its own build history |
doc_summary, post_summary, or video_summary vector ranks in the top results, the worker fires a follow-up filtered query (chunk_type ≠ *_summary AND title = {matched_title}) to fetch the actual body chunks from that document. Summary hits are replaced in the result set by their body-chunk siblings before Claude sees them — so Claude always reads prose, not abstracts.Every POST /query, ask_pibot MCP call, and Discord @mention follows the same pipeline:
voyage-3 encodes the question into a 1,024-dim vector (input_type: "query").top_k = TOP_K_EACH (typically 5); meta queried at top-3; transcripts queried at top-3; sig_meeting_page sub-query fired in parallel for the sig namespace.normalize*() function mapping raw Pinecone metadata to a standard shape: { source, type, label, title, authors, date, url, summary, excerpt, score, ... }. Namespace-specific fields (e.g. sig_display, channel_name, domain) are preserved alongside.mergeResults() applies the tier weights above, deduplicates by URL, extracts transcript cache hits (score ≥ 0.52), and returns the top MAX_SOURCES items.buildContextBlock() formats each source as a labeled excerpt block for the Claude prompt: [PDF — "Title" — Author(s) — YYYY]\n{excerpt}, with namespace-appropriate prefixes.cache_control: ephemeral; user message is "Question: {q}\n\nRelevant archive excerpts:\n\n{context}". Discord requests use the DISCORD_SYSTEM_PROMPT variant. History array passed for multi-turn sessions.answer, sources (with metadata for badge/link rendering), and cache_hits (high-scoring transcript matches surfaced separately).Claude Sonnet is used throughout — query answering, document enrichment, meeting summarization, link relevance scoring (Haiku for the latter two). The research material is dense and cross-disciplinary; strong synthesis matters more than fast extraction.
System prompt structure. The system prompt is an inline document (~1,100 tokens) in api/worker.js, maintained alongside the retrieval code. It is not loaded from a file at runtime. It was originally inspired by SOUL.md (the conceptual identity document in the repo root), but has since evolved independently and is now substantially more detailed. An agent reproducing this pattern should treat the inline prompt as the source of truth. It contains seven sections:
The Discord variant appends a DISCORD VOICE OVERRIDE block that supersedes the voice and length instructions. Both variants exceed the 1,024-token Anthropic prompt-cache threshold and cache independently (cache_control: ephemeral).
Rate limits: 20 queries per IP per hour via the web UI. After 8 turns, the conversation can be downloaded as Markdown and continued in Claude, or accessed without a turn limit via MCP.
Three interfaces share the same corpus, query engine, and Cloudflare Worker. They differ in voice, turn model, and how they reach users.
Browser-based chat at the root URL. Full research-librarian voice. 8-turn session limit (download as Markdown to continue). Submitted conversations can be shared publicly or kept private; public submissions are indexed into the transcripts namespace for self-memory. The /chats route is a public transcript browser.
Gateway bot (discord.py, WebSocket) running on the host machine under launchd (org.protocol-institute.c3po-bot). All bot requests pass context: "discord" to the Worker, selecting the 2–3 sentence office-manager response style.
discord_guide query instead of corpus RAG, returning the top 3 relevant channels with blurbs and meeting schedules.POST /interactions (Ed25519 verified); command enqueued to Cloudflare Queue; Worker queue consumer runs the RAG pipeline and posts back via Discord followup webhook.Completed bot conversations are spooled to data/spool/bot_conversations/; the listener daemon picks them up each cycle and ingests them into the transcripts namespace.
PIBot is available as a Model Context Protocol server (JSON-RPC 2.0 + Streamable HTTP) at https://pibot.protocolized.io/mcp. Connect it to Claude Code or Claude Desktop to query the corpus directly inside your AI client — no turn limit, no browser required.
| Tool | What it does | Auth | Limit |
|---|---|---|---|
search_corpus | Semantic search — returns ranked excerpts with metadata and URLs; no LLM call. Filter by namespace: pdfs, substack, videos, bibliography, discord, sig, discord_links, definitions, or all. Result limit 1–20 (default 10). | None | 100 calls/IP/day |
ask_pibot | Full RAG: embed → retrieve → Claude Sonnet synthesis. Accepts history array for multi-turn sessions. Web voice (not Discord office-manager). | Bearer token | Circuit-breaker shared with web UI |
search_corpus is open — no key required. Good for agentic workflows that need raw retrieval without LLM cost or turn limits.
ask_pibot requires a Bearer token (each call invokes Claude Sonnet and Voyage AI at real cost). To request access email team@protocol-institute.org.
Search only (no key needed) — run once in your terminal:
claude mcp add pibot --transport http https://pibot.protocolized.io/mcp
Full access with Bearer token:
claude mcp add pibot --transport http https://pibot.protocolized.io/mcp \
--header "Authorization: Bearer <your-key>"
Add to claude_desktop_config.json (on Mac: ~/Library/Application Support/Claude/):
{"mcpServers": {"pibot": {
"type": "http",
"url": "https://pibot.protocolized.io/mcp",
"headers": {"Authorization": "Bearer <your-key>"}
}}}
For search-only without auth, omit the headers key.
ask_pibot accepts a history array of {"role": "user"|"assistant", "content": "..."} objects alongside your question. Pass prior turns to maintain context. The same hourly and daily circuit breakers that govern the web UI apply — calls return an error if the budget is exhausted and auto-reset at the next hour or midnight PT.| Component | Technology |
|---|---|
| Cloudflare Worker | Single V8 isolate at pibot.protocolized.io serving web UI, RAG API, MCP server, and Discord Interactions endpoint. PI org account (7e8c7969b2464d23795c555bc6a32af8). |
| Pinecone | Index c3po — 1,024d cosine, serverless aws/us-east-1, PI org account. 11 namespaces, ~26,300 vectors. |
| Voyage AI | Model voyage-3 (1,024d). PI org account. Same model for ingest and query. |
| Cloudflare KV | Rate limiting (20 web queries/IP/hour; 100 MCP search calls/IP/day); circuit breaker flag; transcript storage (90-day TTL); usage accumulators. |
| Cloudflare Queue | c3po-oracle queue for Discord slash command deferred responses. |
| Circuit breaker | KV flag + hourly cron; sleeps when hourly spend exceeds $4 or daily spend exceeds $30. Auto-resets at next hour / midnight PT. |
| Ingest daemon | bin/daemon.py — 14-step sync cycle every 30 minutes, launchd-managed (org.protocol-institute.c3po.daily). Logs to ~/Library/Logs/c3po/daemon.log. |
| Discord bot process | bin/c3po_bot.py — discord.py WebSocket gateway, launchd-managed with KeepAlive (org.protocol-institute.c3po-bot). Logs to ~/Library/Logs/c3po/c3po_bot.log. |
| Substack sync | GitHub Actions workflow (.github/workflows/sync-substack.yml) — daily at 08:00 UTC; commits state files back to repo. |
| Source code | Protocol-Institute/c3po (transferred from vgururao/c3po on 2026-05-31). |
ARCHITECTURE.md in the repository.