Praetor Rebuild Roadmap
Status: 2026-05-14. This roadmap replaces the older phase plan. The old plan is
historical context only; active work should follow this file.
Update 2026-05-19: product identity and Agent/Role semantics are refined by
docs/ADR-002-praetor-as-ai-company.md and
docs/PRAETOR_PAPERCLIP_MATURITY_ROADMAP.zh-CN.md. Where older roadmap or
brief language says "role-first" or "users do not manage agents", use the newer
decision: users manage Agent employees; Roles are the responsibility and
permission frame for those Agents.
Praetor is being rebuilt as a new product, not as a Hermes Agent fork. Hermes is
reference material for agent settings, tool/capability routing, gateway adapters,
session/audit patterns, and executor discipline. Praetor's product shape is a
company-governance operating system led by a proactive AI CEO.
North Star
Praetor should let a solo founder run an AI company from a browser-first
chairman's office:
- The owner talks primarily to the CEO, not directly to worker agents.
- The CEO decides whether to answer directly, ask one question, create a
mission, staff a team, request approval, or keep working in the background.
- The CEO is proactive and learns the owner's style. Low-risk internal work
should happen without approval ceremony.
- Missions create real artifacts. Developer and tester roles must be able to
use Codex CLI through a controlled executor bridge to modify files and run
tests.
- Company Memory, CEO Operating Memory, and Company Playbook are separate.
- Web UI is the main office. Telegram is the first private mobile CEO channel.
- Approval exists for real risk: external actions, spending, destructive
changes, credentials, security/legal boundaries, and governance changes.
Product Model
The active mental model is:
Chairman / Owner
-> Praetor CEO
-> Mission / Decision / Approval / Briefing
-> Project Manager
-> Developer / Tester / Reviewer / Marketing / Legal / Security
-> Codex CLI / API model / browser QA / future executors
User-facing objects:
CEO Conversation: default interaction surface.Mission: tracked unit of work that needs execution, artifacts, staffing,
workspace writes, or background progress.
Agent: AI company employee with role, responsibility, manager, permissions,
current work, heartbeat, budget, and output history.
Role: durable job description / authority frame used by Agents.Decision: durable business choice with rationale.Approval: explicit permission boundary crossing.Board Briefing: owner-facing planning or completion artifact.Company Memory: company facts, decisions, docs, open questions.CEO Operating Memory: owner preferences, CEO style, approval habits.Company Playbook: reusable work methods / SOPs / skills.Runtime Run: executor attempt, logs, changed files, test result.
Technical Direction
Use modern but practical technology:
- Frontend: React + Vite + TypeScript, TanStack Query, React Router, Tailwind
tokens. Keep UI browser-first and dashboard-native.
- Backend: Python 3.12 + FastAPI + Pydantic v2.
- Worker: Python background worker using SQLite-backed jobs.
- Storage: filesystem as source of truth for workspace artifacts; SQLite for
index, state, audit, run metadata, and UI queries.
- Executor bridge: local host-side
praetor-execdFastAPI service. - First coding executor: Codex CLI, non-interactive task mode.
- First private messaging channel: Telegram owner-only CEO access.
- Tests: pixi tasks for Python smokes, TypeScript typecheck/build for web,
focused end-to-end smoke tests for flagship flows.
Do not lock Praetor to a default product tech stack. When creating a new
software product, the CEO should infer from context, use learned company
preferences, or ask the owner when the choice materially matters.
Execution Rules
- [ ] Every phase below must end with code wired into the real app, not just docs.
- [ ] Every behavior must have either a smoke test, typecheck/build coverage, or
a documented manual verification path.
- [ ] Do not preserve old UI/routes/data models just because they exist. Remove
stale structures once their replacement is live.
- [ ] Prefer vertical slices over broad scaffolding.
- [ ] Keep approvals sparse by default. Add auditability and rollback before
adding more permission prompts.
- [ ] Keep raw agent discussion as evidence. Convert it into decisions, tasks,
artifacts, findings, playbook entries, or briefings.
Phase 0 - Rebase Product Direction
Goal: align the repo with the new Praetor product architecture.
- [x] Define proactive CEO mode in
docs/PRAETOR_CEO_INTERACTION_POLICY.zh-TW.md. - [x] Add
ceo_memory_updateandplaybook_entryplanner actions. - [x] Separate Company Memory from CEO Operating Memory in the data model.
- [x] Surface Company Playbook and CEO Operating Memory in
/memory. - [x] Update planner smoke tests for asynchronous CEO turns.
- [x] Rename user-facing "skills" surfaces to "Company Playbook" where they are
not developer/admin internals.
- [x] Review all docs for obsolete "agent playground", "skill registry as agent
page", or approval-heavy language.
- [x] Mark old v0 route/page specs as historical if they conflict with this file.
- [x] Create a short architecture decision record:
docs/ADR-001-praetor-not-hermes-runtime.md.
Acceptance:
- [x]
pixi run py-compile - [x]
pixi run planner-smoke - [x]
pixi run safety-policy-smoke - [x]
npm --prefix apps/web run typecheck - [x]
ROADMAP.mdis the active checklist used for future work.
Phase 1 - CEO Intake And Learning Core
Goal: make the CEO the primary intelligent router.
Intake Router
- [x] Define
CEOIntakeRouteenum:
quick_reply, direct_answer, clarifying_question,
lightweight_action, mission, approval, decision_response.
- [x] Replace scattered route heuristics with a single
CEOIntakeRouter. - [x] Preserve fast rule routes for obvious read-only questions.
- [x] Add JSON classifier fallback with strict allowed routes.
- [x] Add unit/smoke tests for:
direct status question does not create mission,
product discussion routes to planner,
"start building" routes to planner,
"remember this" routes to planner for CEO memory,
"create a reusable method" writes playbook.
CEO Operating Memory
- [x] Store CEO memory as first-class records, not only wiki append text.
- [x] Add fields:
category,confidence,source,last_confirmed_at,
active, superseded_by.
- [x] Add API:
GET /api/memory/ceo,
PATCH /api/memory/ceo/{id},
POST /api/memory/ceo/{id}/disable,
POST /api/memory/ceo/{id}/forget.
- [x] Add
/memoryUI controls: edit, disable, forget. - [x] Add
/memoryUI control: promote CEO memory to standing order. - [x] Add audit events for every CEO memory mutation.
Proactive Trust Policy
- [x] Add
proactive_trustto onboarding as the recommended default. - [x] Add learned approval policy records:
approve_once, approve_for_mission, always_allow_type,
always_ask_type, never_allow_type.
- [x] Add hard boundaries:
external send, public publish, spending, destructive delete,
credential/security change, legal/financial commitment, new external
integration.
- [x] Update approval UI to teach the CEO, not just approve/reject.
- [x] Add Telegram approval actions that map to the same learned policy model.
Acceptance:
- [x] CEO answers simple project status without mission creation.
- [x] CEO can remember "be more proactive" without owner approval.
- [x] Owner can remove a bad CEO memory from
/memory. - [x] Approval "always allow this type" changes future behavior.
Phase 2 - Company Memory And Company Playbook
Goal: make memory useful, separated, and correct.
Memory Structure
- [x] Split
/memoryAPI response into:
company_wiki, decisions, open_questions, documents,
company_playbook, ceo_operating_memory, knowledge_updates.
- [x] Stop relying on page names such as
CEO Memory.mdfor semantic state. - [x] Add memory scopes everywhere:
company, ceo, mission.
- [x] Add migration for old
Wiki/CEO Memory.md:
owner preferences -> CEO Operating Memory,
company facts -> Company Wiki candidates,
ambiguous entries -> needs review.
- [x] Make memory search scope-aware.
- [x] Ensure worker prompts never receive full CEO Operating Memory.
Company Playbook
- [x] Rename
AgentSkillSpecinternally or wrap it asPlaybookEntry. - [x] Add playbook fields:
trigger_examples, procedure, required_capabilities, risk_level,
approval_policy, learned_from, usage_count, last_used_at.
- [x] Add API:
GET /api/memory/playbook,
POST /api/memory/playbook,
PATCH /api/memory/playbook/{id},
POST /api/memory/playbook/{id}/disable.
- [x] Add playbook matching in CEO planning context.
- [x] Add "used playbook" to mission work trace.
- [x] Add UI for active/draft/deprecated playbook entries.
Memory Promotion
- [x] Update
MemoryPromotionReviewto produce separate candidates:
company fact, CEO preference, playbook entry, open question, decision,
do-not-promote.
- [x] Low-risk CEO preferences can auto-apply.
- [x] High-risk governance/strategy changes require decision or standing order.
- [x] Playbook candidates from successful reviewed missions can auto-save as
draft or active depending on risk.
Acceptance:
- [x]
/memoryclearly separates Company Memory, CEO Memory, and Playbook. - [x] A mission closeout can create a playbook candidate.
- [x] A CEO preference never appears as normal company wiki knowledge.
Phase 3 - Mission And Team Operating System
Goal: make missions behave like managed company work, not chat threads.
Mission Creation
- [x] Add
MissionType:
new_product, existing_repo, research, operations,
planning_only, general.
- [x] Add
ProjectContext:
mode, workspace_root, repo_url, allowed_paths,
protected_paths, tech_stack_decision.
- [x] CEO creates mission only when work requires execution, tracking, file
artifacts, delegation, background run, or owner explicitly says to start.
- [x] Add mission proposal UI showing:
title, why mission is needed, team, workspace, artifacts, budget,
approval boundaries.
Team Model
- [x] Define default durable roles:
CEO, Project Manager, Product Manager, Developer, Tester, Reviewer,
Marketing Lead, Design Lead, Legal Counsel, Security Officer,
Memory Steward.
- [x] Make PM responsible to CEO for mission delivery.
- [x] Add role disagreement records:
participant, position, rationale, PM decision, escalation if needed.
- [x] Add PM authority:
low-risk sequencing, local implementation tradeoffs, role coordination.
- [x] Add escalation:
strategy, budget, external, legal/security/privacy, unresolved conflict.
- [x] Allow PM to request new mission-scoped agents.
- [x] CEO auto-approves low-risk role expansion under
proactive_trust.
Work Trace
- [x] Convert agent-to-agent discussion into structured
WorkTraceEvent. - [x] Keep raw turns for audit but show concise trace in mission UI.
- [x] Add event types:
staffing, delegation, disagreement, PM decision, executor run,
test result, review finding, role request, briefing ready.
Acceptance:
- [x] Owner can discuss idea with CEO, then say "start building".
- [x] CEO creates mission and proposes team.
- [x] PM owns final team-level decision inside mission scope.
- [x] Team can request Tester and continue after CEO approval/auto-approval.
Phase 4 - Codex CLI Executor V1
Goal: real product development through non-interactive Codex tasks.
Bridge Contract
- [x] Finish
praetor-execdinstall flow for host-side Codex CLI. - [x] Verify bridge binds only to loopback by default.
- [x] Require auth token on every bridge request.
- [x] Validate executor type:
codexfirst,claude_codelater. - [x] Enforce workspace allowlist and denylist at bridge level.
- [x] Capture stdout, stderr, exit code, timing, run log path.
- [x] Support
start,status,cancel, and event/log polling.
Task Spec
- [x] Define
ExecutorTaskSpec:
role, mission_id, task_id, workdir, goal, context_files,
allowed_paths, protected_paths, constraints, expected_outputs,
test_commands, timeout, approval_policy.
- [x] Generate Codex prompt from structured task spec.
- [x] Require final response format:
summary, changed files, tests run, known limitations,
follow-up recommendations.
- [x] Persist task spec and result.
Run Normalization
- [x] Extend
RunAttempt:
changed_files, tests_run, summary, known_limitations,
raw_log_path, exit_code, requires_owner_action.
- [x] Parse Codex final output into normalized result.
- [x] Detect common pause states:
auth required, approval needed, test failure, timeout, blocked.
- [x] Attach changed files and test output to mission timeline.
Developer Role
- [x] Developer uses Codex CLI for scoped implementation.
- [x] Tester uses Codex CLI for smoke/test tasks.
- [x] Reviewer can use API model first, Codex if code inspection is needed.
- [x] Developer output always goes to review before CEO briefing.
Acceptance:
- [x] Praetor can run a Codex task in a scoped workspace.
- [x] Changed files and tests are visible in mission UI.
- [x] Failed tests produce a follow-up Developer task.
- [x] Codex auth failure becomes a clear runtime health issue.
Phase 5 - Flagship Development Workflow
Goal: prove Praetor can actually build or modify software.
Scenario A: Existing Repo
- [x] Owner selects or enters existing repo/workspace.
- [x] CEO asks only for missing high-impact info.
- [x] PM creates task plan.
- [x] Developer/Codex implements a small change.
- [x] Tester/Codex runs relevant tests.
- [x] Reviewer checks acceptance criteria.
- [x] CEO provides briefing with changed files, tests, screenshots if relevant.
Scenario B: New Product
- [x] Owner discusses idea with CEO.
- [x] CEO decides whether tech stack is obvious, learned, or needs owner choice.
- [x] PM writes product spec and technical spec.
- [x] Developer/Codex scaffolds a simple app/tool.
- [x] Tester runs smoke tests.
- [x] CEO returns local run instructions and briefing.
Briefing And Owner Review
- [x] Board Briefing includes:
summary, completed work, files, test commands/results, screenshots,
risks, decisions needed, next mission recommendations.
- [x] Owner can continue, stop, request changes, ask tester to rerun, or create
next mission.
- [x] Completed mission can promote decisions, facts, and playbook entries.
Acceptance:
- [x] One existing repo change can be completed end to end.
- [x] One new simple product can be created end to end.
- [x] Owner can test locally from CEO briefing.
Phase 6 - Web UI: Chairman's Office
Goal: make the browser UI the real product surface.
Information Architecture
- [x]
/office: CEO chat, daily briefing, pending decisions, active missions,
risk/blocked rail, recent results, runtime health.
- [x]
/missions: mission portfolio and filters. - [x]
/missions/:id: mission detail, team, work trace, files, runs, tests,
decisions, approvals, briefing.
- [x]
/memory: Company Memory, CEO Operating Memory, Company Playbook,
Decisions, Open Questions.
- [x]
/runtime: models, Codex bridge, executor health, costs, run history. - [x]
/settings: owner profile, workspace, autonomy, Telegram, hard
boundaries.
UI Cleanup
- [x] Remove legacy Jinja pages once React equivalents are live:
memory, decisions, meetings, inbox, agents, models, mission pages.
- [x] Keep Jinja only for onboarding/login/mobile fallback if still useful.
- [x] Keep React
main.tsxsmall; move UI into components/hooks/pages. - [x] Add skeleton, empty, error, degraded states for every route.
- [x] Add persistent Needs Decision rail.
- [x] Add command palette jump to approvals, mission, runtime, memory.
Visual Direction
- [x] Follow Praetor brand: calm, decisive, structured, authoritative.
- [x] Avoid chat-app bubbles as the whole product.
- [x] Avoid agent playground visuals.
- [x] Prioritize dense but readable operational information.
Acceptance:
- [x] Owner can operate the flagship workflow from Web UI.
- [x] No primary workflow depends on old Jinja pages.
- [x]
npm --prefix apps/web run typecheck - [x]
npm --prefix apps/web run build
Phase 7 - Private CEO Channel: Telegram First
Goal: let the owner manage CEO decisions privately from mobile.
- [x] Owner-only pairing flow remains mandatory.
- [x] Add Telegram commands:
/brief, /status, /missions, /approvals, /pause, /resume,
/ask.
- [x] Approval notifications include teaching options:
approve once, approve for mission, always allow type, always ask,
reject, never allow.
- [x] High-risk approvals can deep-link to Web UI.
- [x] Telegram can answer status and briefing without creating mission.
- [x] Telegram cannot bypass hard boundaries.
- [x] Add gateway abstraction for future channels:
Signal, WhatsApp, Email, Slack.
Acceptance:
- [x] Telegram
/briefreturns current CEO briefing. - [x] Telegram approval changes learned approval policy.
- [x] Telegram unauthorized user is rejected and audited.
Phase 8 - Governance, Audit, And Safety
Goal: make high autonomy trustworthy without making it bureaucratic.
- [x] Create hard-boundary policy module independent of UI.
- [x] Add policy evaluation before executor run, external action, spending,
file deletion, secret access, deployment/publish.
- [x] Add checkpoint snapshots before risky file modifications.
- [x] Add audit event browser with filters.
- [x] Add run replay context: task spec, prompt summary, logs, changed files,
test result, approvals.
- [x] Add "why did CEO do this?" explanation panel from audit + memory.
- [x] Add recovery path for interrupted mission jobs and bridge crashes.
Acceptance:
- [x] Proactive actions are auditable.
- [x] Risky actions pause with clear reason and options.
- [x] Restart during mission marks jobs interrupted and visible.
Phase 9 - Runtime, Providers, And Cost
Goal: support serious local-first operation.
- [x] Runtime settings API supports API model and subscription executor mode.
- [x] Codex CLI is the first-class subscription executor.
- [x] Claude Code adapter is behind the same interface, after Codex is stable.
- [x] Provider profiles support OpenAI-compatible, Anthropic, local endpoints.
- [x] Add fallback behavior for CEO planner and reviewer model.
- [x] Add runtime health:
model configured, bridge reachable, Codex logged in, workspace available,
queue health, recent failure.
- [x] Add cost/token/run duration summary.
- [x] Add budget policy per mission and learned owner preference.
Acceptance:
- [x] Runtime page explains exactly what is healthy/degraded.
- [x] Mission stops or pauses cleanly on budget/time limit.
- [x] Codex login expiration is surfaced as an actionable issue.
Phase 10 - Packaging And Deployment
Goal: make Praetor installable and recoverable.
Paperclip-inspired install/operator maturity work is tracked in
docs/PRAETOR_PAPERCLIP_MATURITY_ROADMAP.zh-CN.md Phase 1, and should be kept
in sync with this phase.
- [ ] One-line local install tested on fresh macOS.
- [ ] One-line local install tested on Linux.
- [x] WSL2 path tested or documented honestly.
- [ ] Docker compose app stack starts Web/API/Worker.
- [x]
praetor-execdsetup script installs bridge and config. - [x] Backup and restore includes:
state DB, workspace, bridge logs, config, CEO memory, playbook.
- [x] Update script preserves owner data and workspace.
- [x] Uninstall script can keep or purge data.
- [x] Security docs explain loopback bridge, tokens, workspace allowlists,
Telegram owner-only model.
Acceptance:
- [ ] A new user can install, onboard, connect Codex, and run a simple mission.
- [x] Backup restore is verified locally.
Phase 11 - Post-V1 Evolution
These are important but not required before the first real v1.
Company template/import-export work is tracked in
docs/PRAETOR_PAPERCLIP_MATURITY_ROADMAP.zh-CN.md Phase 11, and should be kept
in sync with this phase.
- [ ] Long-running Codex sessions for Developer agents.
- [ ] More executor adapters: Claude Code, OpenCode, custom local agent.
- [x] Browser QA with screenshots and visual checks.
- [x] GitHub integration: issues, PRs, CI results.
- [ ] Email and Signal CEO channels.
- [x] Richer Company Playbook editor and versioning.
- [ ] Multi-workspace / multi-company support.
- [ ] Advanced role tuning and permission profiles.
- [ ] Cloud/private remote deployment mode.
Verification Matrix
Run as relevant after each phase:
- [x]
pixi run app-import-check - [x]
pixi run packaging-smoke - [x]
pixi run py-compile - [x]
pixi run planner-smoke - [x]
pixi run high-risk-memory-smoke - [x]
pixi run safety-policy-smoke - [x]
pixi run organization-smoke - [x]
pixi run office-smoke - [x]
pixi run telegram-smoke - [x]
pixi run bridge-smoke - [x]
pixi run bridge-e2e - [x]
pixi run browser-qa-smoke - [x]
pixi run playbook-version-smoke - [x]
npm --prefix apps/web run typecheck - [x]
npm --prefix apps/web run build
Definition Of Done
An item is done only when:
- [ ] The code path is wired into the real app.
- [ ] The UI, API, worker, or executor behavior is reachable by a user flow.
- [ ] The behavior is covered by a smoke, typecheck/build, or documented manual
verification.
- [ ] The relevant docs are updated.
- [ ] Old conflicting structures are removed or explicitly marked historical.
Active Reference Docs
- PRAETOR_PRODUCT_BRIEF.zh-TW.md
- docs/PRAETOR_CEO_INTERACTION_POLICY.zh-TW.md
- docs/PRAETOR_SYSTEM_SPEC.zh-TW.md
- docs/PRAETOR_UI_SPEC.zh-TW.md
- docs/PRAETOR_BRAND_SPEC.zh-TW.md
- docs/PRAETOR_MEMORY_PROMOTION.md
- docs/PRAETOR_EXECUTOR_BRIDGE_SPEC.zh-TW.md
- docs/PRAETOR_SURFACES_SPEC.zh-TW.md
- docs/PRAETOR_REPO_ARCHITECTURE.zh-TW.md
Historical material may remain in the repo, but it should not override this
roadmap.