FounderCLI

Playbook12 min read

How to run an AI agent in shadow mode before granting production authority

A six-stage, 16-week-or-shorter playbook for collecting production evidence while humans retain consequential decisions and rollback remains available.

By FounderCLI Editorial

Run a no-authority, versioned decision trial—not a passive demo—and retain evidence to compare the agent with the authorized path. Treat each result as local evidence about one frozen model–pipeline–policy cohort, not proof that the agent is safe, accurate, autonomous, compliant, or ready for broad authority.

Use this method for one bounded workflow with a domain governor, independent inspection where feasible, durable records, and a human or deterministic fallback. It is not a substitute for legal, security, privacy, or domain review.

1. Stage 0 — Decide whether the workflow is admissible (day 0)

  • Prerequisites: Bounded workflow, governor, record retention, fallback, and completed required reviews.
  • Inputs: Workflow map and consequence analysis.
  • Outputs: Signed trial charter or no-run decision.
  • Owners: Executive risk owner and domain governor.
  • Decision: Proceed, narrow the scope, or stop.
  • Controls: Keep final authority with the governor; retain the human or deterministic path.
  • Stop condition: Stop if a critical role is vacant, exact cases cannot be retained, fallback fails, or required review remains incomplete.

The NIST AI RMF Core calls for documented risk responsibilities and leadership accountability. It also calls for differentiated human–AI oversight roles. Name an executive risk owner, workflow owner, domain governor with final decision authority, evaluation lead independent from the builder where feasible, security or platform owner, incident and rollback owner, and record custodian.

Bound the trial to one decision and an explicit exclusion list, deny the agent production writes and credentials, and preserve the authorized human or deterministic path. Make the human’s approval technically and procedurally binding. Reserve the consequential decision for the authorized human rather than treating an approval prompt as sufficient. Treat human approval as an authority boundary, not a correctness score. Specify and test the monitoring, intervention, independent-review, and recovery controls required by the workflow.

2. Stage 1 — Freeze the protocol, cohort, and gates (weeks 1–2)

  • Prerequisites: Signed charter, taxonomy, baseline records, risk analysis, and current policy.
  • Inputs: Those charter, taxonomy, baseline, risk, and policy records.
  • Outputs: Cohort manifest, authority matrix, label dictionary, metrics, threshold card, stop card, and rollback plan.
  • Owners: The evaluation lead, domain governor, platform owner, and executive risk owner.
  • Decision: Admit the cohort, narrow it, or do not start shadow operation.
  • Controls: Freeze the model, pipeline, policy, tools, inputs, labels, and measurement definitions listed in the cohort manifest.
  • Stop condition: Close the cohort for any configuration change, lost record integrity, bypassed human authority, severe erroneous proposal, control failure, or provider or security incident.

Use a 16-week maximum: protocol weeks 1–2, instrumentation weeks 3–4, shadow operation weeks 5–12, adjudication weeks 13–14, and authority decision weeks 15–16. Define the cohort by exact model and version, inference parameters, system and task prompts, pipeline or code digest, retrieval sources, tool allowlist, input schema, and policy or rubric version.

Every change to the frozen model, pipeline, or policy configuration closes that cohort, quarantines the affected result, and starts a separately measured cohort. The reason is methodological: an ACL study found that evaluation depends on the target use-case distribution. It also found that correlations among test prompts can change model rankings. Do not pool changed cohorts as though they were identical.

Predeclare workflow-specific thresholds for coverage, independently supported correctness, omissions, severity, abstention, missed slots, provider failures, latency, overrides, and incidents. Set those thresholds for this workflow; do not present a number here as a universal default or as endorsed by a source.

Make the threshold card usable: metric | threshold | evidence source | owner | consequence | unresolved-case treatment. Define coverage against eligible slots, use independent outcome evidence for correctness, require an omission search, and specify how unresolved cases affect promotion or incident handling. The domain governor signs correctness and severity; the evaluation lead signs denominator and evidence integrity; the platform or security owner signs reliability metrics.

Also predeclare stop and rollback triggers for lost record integrity, bypassed human authority, configuration drift, severe erroneous proposals, control failures, and provider or security incidents. Thresholds that are not met should produce a named consequence rather than an informal explanation.

3. Stage 2 — Build the immutable case envelope (weeks 3–4)

  • Prerequisites: Cohort manifest, record schema, access controls, and an available fallback.
  • Inputs: Cohort manifest and record schema.
  • Outputs: Tested case store, separate machine and human forms, usage ledger, comparison job, access controls, and successful fallback and restore tests.
  • Owners: The record custodian, evaluation lead, and platform owner.
  • Decision: Admit live inputs only after the end-to-end control test passes.
  • Controls: Use append-only records, separate machine and human forms, restricted access, and a restore rehearsal.
  • Stop condition: Stop admission if recording, access, comparison, fallback, or restore cannot be demonstrated.

Configure an append-only record to capture the case ID and time, immutable input snapshot or digest, cohort ID, agent proposal, evidence and tool trace, abstention reason, sealed timestamp, human decision and rationale, comparison labels, independent outcome evidence, revisions, usage, latency, errors, and approvals. Atlan describes durable traces containing every step, conclusion, human correction, and mistake selected for non-repetition.

Seal the machine proposal before the human decision. If the reviewer sees it, label the case reviewer-exposed and do not treat it as an independent comparison. Preserve every correction as a new exact revision with provenance, and require separate approval before a case-specific correction becomes reusable policy or knowledge. Admit live inputs only after a synthetic end-to-end test confirms recording, access, comparison, fallback, and restore.

Recording example — FounderCLI: Run 7 saved ten operations (376,566 input tokens, 48,946 output tokens, and 1,227,839 active model milliseconds). A completed response was omitted from the older Payload ledger after deterministic assembly failed. Its conservative emitted total was 390,152 input tokens, 50,626 output tokens, 11 calls, and 1,292,146 active model milliseconds. Run 8 saved nine operations (158,891 input tokens, 37,415 output tokens, and 858,213 active model milliseconds). These figures are model usage, not human labor or money saved. Run 7 exposed source-ID mismatches, passage-alias problems, a draft-slug collision, and an omitted usage receipt. Completed outputs were recovered without another provider call for source and passage fixes. Later implementations checkpointed usage before deterministic artifact handling. They also deduplicated replay by operation identity. The failures remain in the record.

4. Stage 3 — Run each production slot through the shadow loop (weeks 5–12)

  • Prerequisites: Admitted cohort, eligible production events, staffed human path, and functioning fallback.
  • Inputs: Each eligible event and frozen cohort.
  • Outputs: Proposals, human dispositions, terminal paths, usage, and incidents.
  • Owners: Case decision maker, platform owner, and record custodian.
  • Decision: The human decides whether to publish, revise, or reject; the system records missed-slot, provider-failure, and configuration-change paths when they occur.
  • Controls: Reconcile eligible production slots against stored cases every day, enforce tool boundaries, use deadlines or turn caps, monitor in real time, and keep the fallback available.
  • Stop condition: Stop or revert on a bypass of human authority, a provider or security incident, a severe erroneous proposal, or a failed control.

Route every eligible slot to exactly one terminal path: publish, where the human authorizes; revise, where the human changes the proposal and the exact difference is retained; reject, with a human reason; missed-slot, where the agent is absent or late and the human proceeds; provider-failure, where fallback runs and the failure is logged; or configuration-change, where the result is quarantined, the cohort closes, and a new cohort begins. Only the human-authorized action may reach production during shadow mode.

Make abstention and policy conflict defined outcomes that return evidence to a human instead of forcing a prediction. Enforce tool boundaries, deadlines or turn caps, deterministic fallbacks, real-time monitoring, and immediate human intervention. Record failed attempts and fallback use rather than hiding them through retries.

In LogicWeave’s operational case, the review interface combined incoming data, deterministic results, history, candidates, AI research, conflicts, and abstention reasons while leaving authority with the analyst. For this playbook, count missed slots and provider failures in the cohort denominator. This is a reporting convention for the workflow experienced, not a claim that every other denominator is misleading.

5. Stage 4 — Adjudicate disagreements, omissions, and outcomes (weeks 13–14)

  • Prerequisites: Complete cohort ledger, independent outcome evidence, and separate inspection from the builder where feasible.
  • Inputs: The complete cohort ledger and independent outcome evidence.
  • Outputs: Locked evaluation table, severity review, unresolved-case queue, and proposed control changes.
  • Owners: The domain governor and independent evaluation lead.
  • Decision: Invalidate incomplete evidence, continue the unchanged cohort, or send proposed modifications to a new cohort.
  • Controls: Keep machine and human dispositions separate, search for missed findings, reconcile the cohort denominator, and report by configuration.
  • Stop condition: Stop promotion if evidence is incomplete, a severe omission is found, or a stop-condition crossing is unresolved.

Label machine and human dispositions separately. Record agreement, exact revision, rejection, abstention, missed slot, provider failure, policy conflict, agent-only finding, human-only finding, and unresolved outcome. Do not score agreement as correctness until the domain governor checks independent evidence and searches explicitly for findings missed by both the agent and the original human decision.

Atlan reports that model confidence measured how convincing the evidence appeared rather than whether a conclusion was correct. A familiar incident could score highly while identifying the wrong cause. Report by cohort: eligible-case count, completion and missed-slot rates, disposition matrix, exact-revision rate, abstentions, independently supported errors, omissions by severity, overrides, appeals, provider failures, latency, usage, incidents, and stop-condition crossings.

Use disagreements to identify policy or data boundaries. Keep exact revisions and omission-search results as records for adjudication and later control decisions. Invalidate incomplete evidence, continue the unchanged cohort, or send proposed modifications into a new cohort rather than repairing the result retrospectively.

Adjudication example — FounderCLI: Machine review left three advisory sourcing findings. Human review dismissed them as editorial recommendations. It approved the exact Post revision. Machine findings and human dispositions remained separate. An August 24 audit found that Evaluation 1 lacked exact-final-candidate machine review. It found that Evaluations 2 and 3 contained reviews only for earlier hashes. Those evaluations’ 57 seeded claims are draft, audit-only records. No complete frozen human-label or missed-finding corpus establishes reliability for these audited cases.

6. Stage 5 — Rehearse rollback, then make the authority decision (weeks 15–16)

  • Prerequisites: Locked evaluation, threshold card, incident and appeal logs, independent review, rollback plan, and tested fallback.
  • Inputs: Records plus restore-test result.
  • Outputs: Rollback-rehearsal record and signed authority decision with effective date and monitoring owner.
  • Owners: Executive risk owner and domain governor; independent evaluator attests evidence integrity.
  • Decision: After rehearsal, reject, continue unchanged, start a new cohort, or consider narrow authority for the tested cohort.
  • Controls: Rehearse reversion before signing; preserve override and appeal, credential isolation, and monitoring.
  • Stop condition: On an incident or stop-condition crossing, remove permissions, invoke fallback, preserve records, open incident review, notify named owners, and require a new signed cohort after any configuration or policy change.

Before the authority decision, use the rollback plan and tested fallback to rehearse reversion to manual or deterministic operation. Record the result. The executive risk owner and domain governor then sign one decision with an effective date and monitoring owner. Require every threshold to pass, no stop condition to be crossed, independent review to be complete, rollback to succeed, and domain, security, operational, and executive owners to sign before authority expands.

Authority-gate example — FounderCLI: The editorial-operations record reports that Runs 7 and 8 reached final human review. Run 7 used foundercli-editorial-v1.5 through v1.10. Run 8 used v1.10. Human reviewers approved exact Post hashes. As of September 14, the Playbook had not completed current production review. Historical publication does not establish current quality, broad workflow generality, or readiness for autonomous publication. New pipeline, model, or policy configurations form separate cohorts. Historical records do not validate v1.14.

If authority expands, grant only the minimum action, tools, data, duration, and volume required. Preserve human override and appeal, continuous monitoring, credential isolation, and automatic reversion to the manual or deterministic path. OWASP guidance supports least-privilege tools and explicit authorization (tool security). It recommends independent validation of sensitive actions. It also recommends logging decisions, tool calls, and outcomes. A decision not to expand authority is a successful trial outcome when the evidence is weak, the operational trade-off is poor, or the controls are inadequate.

Evidence boundaries and trade-offs

In one LogicWeave blind test of 37 lower-confidence records, the reported outcomes were nine exact automatic decisions, 26 safe-review decisions, two unresolved policy cases, and zero wrong automatic decisions. LogicWeave also reported 180 automatic decisions and 51 safe reviews in a controlled 231-case set. It expressly limited that result to the controlled evaluation rather than production accuracy or future files.

Atlan reports that an earlier sequential investigator usually took 10+ minutes per investigation and occasionally up to 30, slower than the humans it was meant to help. For your cohort, test whether bounded execution meets the workflow’s latency requirement.

Treat these case studies as observed, configuration-specific evidence—not proof of generalized performance or endorsement. The NIST AI RMF describes real-time monitoring and human intervention as practical approaches for safety risks. It says independent review can improve testing. It says recovery controls may be options.

Plan four workflow-specific trade-offs. Delay improvements until a new cohort. Keep tasks requiring restricted tools on the human path. Budget storage and data handling for complete records. Count stop-rule cases as incomplete coverage; measure each effect against workflow requirements.

The FounderCLI editorial-operations record states that FounderCLI operates a shadow-mode editorial system. It also states that FounderCLI has an interest in presenting this governance design as useful. The first-party FounderCLI case was designed and documented by the same publication producing this playbook. FounderCLI uses Codex and CometAPI in the system being evaluated. Neither provider sponsors this article. These are publisher disclosures, not independent proof of performance. Where feasible, an evaluator independent of the publication or system builder should review the promotion packet.

7. Reusable operator checklist

The evaluation lead maintains this checklist; the domain governor signs each gate.

Turn real production inputs into an auditable, configuration-specific comparison while the human-controlled path retains every consequential decision. Name the governor and evaluator, then draft the cohort manifest, threshold card, stop card, and rollback test before sending real input.

Shadow mode as a versioned decision trial

Prepare records before live input, append and adjudicate each slot, then rehearse reversion before the reversible authority decision.

Sources

  1. AI RMF Core - AIRC

    First-party record · airc.nist.gov · Accessed September 14, 2026

    Primary NIST AI RMF Core guidance on continuous lifecycle risk management, documented human-AI roles, leadership accountability, and human oversight. Safely fetched from https://airc.nist.gov/airmf-resources/airmf/5-sec-core/; extracted text SHA-256 b3d70b2d8e6ff5d25e75087ed06525edd6170103f2fddc6e111a2b66260d660d.

  2. AI Operations Review System Case Study | LogicWeave

    First-party record · www.logicweave.ai · Accessed September 14, 2026

    Public operational case record with blind evaluation, machine-versus-human decision separation, abstention and policy-conflict handling, controlled release evidence, and explicit limits on accuracy and generalization. Safely fetched from https://www.logicweave.ai/case-studies-ai-operations-review-system/; extracted text SHA-256 6ad40ba0e84686be16354fcdfe69646df6ef9cb2971bdb20edd802b64d7d4a1c.

  3. Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks

    Primary research · aclanthology.org · Accessed September 14, 2026

    Peer-reviewed evidence that evaluation conclusions depend on the target-use-case distribution and that correlated prompts can change model rankings. Safely fetched from https://aclanthology.org/2024.acl-long.560/; extracted text SHA-256 8b5d3cd660de6b2b4ddc7dc2e250c39fa8ceba798d4a01e3fcf4a11219ed0628.

  4. NIST.AI.100-1.pdf

    First-party record · nvlpubs.nist.gov · Accessed September 14, 2026

    Primary NIST guidance supporting separation of development and validation roles, real-time monitoring, human intervention, shutdown, independent review, removal, and recovery controls. Safely fetched from https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf; extracted text SHA-256 15880ddc108f40bafb461ec6075b5fc1a3c8cb5b752a60022cd2e914c9f388fa.

  5. Loop Engineering in Production: Putting AI Agents on Call

    First-party record · blog.atlan.com · Accessed September 14, 2026

    First-party operational evidence supporting bounded and read-only tools, deterministic stop checks, trace review, human validation before action, credential isolation, and human-controlled expansion of permissions. Safely fetched from https://blog.atlan.com/engineering/loop-engineering-in-production-putting-ai-agents-on-call/; extracted text SHA-256 9a70f7c73ee20dbf298ec8bb6254aebfabc412d8f85b1f67e3236f623c863bed.

  6. editorial-operations.txt

    First-party record · 134.199.215.220 · Accessed September 14, 2026

  7. AI Agent Security - OWASP Cheat Sheet Series

    Primary research · cheatsheetseries.owasp.org · Accessed September 14, 2026

    Primary OWASP guidance supporting least privilege, bounded tool permissions, explicit authorization, independent validation of sensitive actions, and monitoring. Safely fetched from https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html; extracted text SHA-256 87a6dd8d5a66a5829f5eefa532561e61c8dc6c2fd69fd003ec2c43aaf5f8f267.

Related