Skip to content

Safety Model

Agent teams build on the same safety infrastructure that governs every bnerd AI session. No new permission system is introduced; the team layer adds connectivity between sessions, not new safety semantics.

SafetyMode floor

Every team recipe declares a safety floor (defaulting to read-only). When you run a team, the --safety flag may only tighten this floor — it can never loosen it:

# The cloud-app recipe declares safety: non-destructive.
# This flag tightens it further to read-only for this run.
bnerd team run cloud-app "explore the codebase" --safety=read-only

# This would be rejected: full is looser than non-destructive.
bnerd team run cloud-app "…" --safety=full  # error: flag loosens the recipe floor

The effective mode is passed to every teammate's tool registry at spawn. A teammate cannot exceed the team's effective mode regardless of what its role file declares.

Per-role scoped tool registries

Each teammate's tool registry is built in this order (most restrictive wins):

  1. SafetyMode floor — the team's effective mode, established first
  2. Role base toolset — the appropriate category of tools for the role (read, write, etc.) filtered by the safety floor
  3. disallowedTools — tools explicitly blocked by the role file, removed unconditionally
  4. tools allowlist — if present, only the listed tool names survive; everything else is dropped (Claude Code tool names are translated to bnerd names at this step)
  5. Team coordination tools — always registered last, not removable by role policy: team_send_message, team_task_claim, team_task_complete, team_task_list, team_task_get

Unknown Claude Code tool names (Agent, Task, Skill) log a warning and are dropped in phase 1.

The coordinator driver's tool surface

Every team is driven by a coordinator holding a separate, always-SafetyRead toolset: team_spawn_teammate, team_task_create/team_task_update, task list/get, team_pause/team_resume, team_stop_teammate, team_replace_teammate, team_send_message, team_status, team_finish, and team_delegate. Three things can hold this toolset — bnerd team run's scripted driver session, your own conversation (once it starts a team — see Agent Teams → Two execution flows), or a teammate the conversation has handed coordination to (a delegated coordinator — see Delegated coordinator) — but never more than one at a time, and team_delegate itself only exists on the conversation's/driver's own binding: a delegate cannot chain-delegate to a second teammate.

None of this toolset is reachable by an ordinary teammate — every tool in it refuses a caller whose session is a roster member (unless that caller is the currently-published delegate calling on its own scoped binding), as defense in depth on top of never being registered on a plain teammate's registry.

Critically, this toolset excludes any approval-resolution tool — no team_task_approve, no way to resolve a pending tool-use confirmation. This is an explicit anti-laundering invariant: an agent that could approve its own team's pending confirms or completion gates would let it launder around the human-in-the-loop boundary. See Human-gate tasks below for what actually resolves those.

A delegate can still dismiss a queued approval — never grant one

team_stop_teammate and team_finish are in this toolset. Stopping a teammate denies every tool-use confirm and ask_question it has pending (they resolve to "no"/"skip" so its session goroutine can unwind rather than block forever), and team_finish stops the whole roster the same way on its way out — see Approvals stay yours in the AI Assistant docs. A delegated coordinator holding those two tools is therefore the first agent-side principal that can make a pending human gate go away — but only by dismissing it, never by granting it. The anti-laundering boundary is about resolving an approval as "yes"; a drain-to-deny on shutdown is not that, and it already happens the same way when the conversation itself (never delegated) calls team_finish or when the TUI's own exit grace runs. Nothing here lets an agent approve its own team's pending work. (A requires_approval task's own completion approval is a separate mechanism — see Human-gate tasks below — and is not touched by either tool; it is simply left awaiting_approval on the persisted board if the team ends before you resolve it.)

Confirmations

Tools classified as requiring confirmation surface a [CONFIRM] prompt in the headless output (or a modal in the TUI). The --auto-approve flag controls what happens:

Policy Tool confirmations Task-completion approvals
never (default) Prompt for every one (y/N) Prompt for every one
safe Still prompt for every one Auto-approved
all Auto-approved Auto-approved

Task-completion approvals (requires_approval: true tasks) are a separate approval class from the tool-use confirmations above — see Human-gate tasks below.

safe does not auto-approve read-only tools

A StreamConfirmRequested event does not yet carry the tool's safety level, so safe has no way to tell a read from a write at the moment it must decide — and it prompts rather than guessing. safe therefore differs from never only in that it clears task-completion approvals without a prompt. Use --auto-approve=all (with --safety=read-only, which makes every available tool read-only anyway) for a genuinely unattended run.

Pending confirmations are queued under a ULID. An MCP client can resolve one with bnerd_team_approve(team_id, confirm_id, decision) — but only when it shares a process with the team, which the stock bnerd mcp-server does not (see MCP reality). Today, a chat-driven team launched from the TUI surfaces these as inline prompts in the conversation instead.

Gate machinery

The task list encodes four levels of gates that prevent premature completion:

Dependency gates

A task with blocked_by: [3, 5] cannot be claimed until tasks 3 and 5 are completed. Eligibility is re-evaluated automatically when any task transitions to completed — no explicit unblock call is needed.

Multi-gate completion (review pairs)

A task can declare named gates that must each be satisfied before it can complete:

{
  "subject": "add POST /networks to hq",
  "gates": [
    { "name": "code-review",    "by_role": "code-reviewer" },
    { "name": "security-review","by_role": "security-reviewer" }
  ]
}

Each gate is satisfied by a teammate calling team_task_satisfy_gate(task_id, gate_name, decision, at_ref?). The task transitions to completed only when all declared gates are satisfied. This is the canonical pattern for critic–actor pairs.

Tip-alignment

Every gate satisfaction carries an optional at_ref (commit SHA or artifact content-hash). When a task tries to complete, the runtime verifies that all gates were satisfied against the same at_ref. If they diverge, the task is held in awaiting_approval and a StreamTipMismatch event is emitted. The operator is notified (bnerd team status / bnerd_team_status) and must reconcile before the task can proceed — see Human-gate tasks below for who resolves an awaiting_approval task.

This prevents a reviewer from approving a stale commit while the author has already pushed new changes.

Human-gate tasks (RequiresApproval)

A task with requires_approval: true cannot transition to completed without an explicit approval decision — and neither the coordinator driver nor any teammate can supply one (anti-laundering invariant, see The coordinator driver's tool surface above). Only the operator can resolve it, one of two ways:

  • Operator policy at launch — bnerd team run --auto-approve=safe or --auto-approve=all auto-approves every completion approval as the task reaches it (no per-task prompt); --auto-approve=never (the default) prompts on stderr/stdin.
  • The TUI's inline prompts — a chat-driven team launched from the TUI surfaces every pending task-completion approval inline in the conversation. This is how it is resolved today for a TUI-launched team.
  • bnerd_team_task_approve(team_id, task_id, decision, feedback?) — an MCP tool, but it only resolves a team running in the same process as the MCP server. The stock bnerd mcp-server is its own separate stdio process, so it cannot reach a team launched from the TUI, bnerd web, or a separate bnerd team run — see MCP reality.

This is the standard pattern for phase-boundary gates:

Phase-2 gate task
  requires_approval: true
  blocked_by: [all phase-1 task IDs]

Phase-2 tasks are blocked until the gate task is approved, so no teammate starts phase-2 work until you (or your --auto-approve policy) have reviewed phase-1 outcomes and given the go-ahead.

Plan-approval gates — not built yet

A team recipe can mark a role plan_approval: true in its frontmatter today, and that flag parses into RoleDef.PlanGated — but nothing currently reads it. There is no spawn-time plan-mode gate, no team_request_plan_approval (or team_plan_decide) tool, and no adjudication path for a teammate's plan before it starts executing. Declaring plan_approval: true on a role has no runtime effect yet. See Roadmap → Agent teams for the planned plan_first + team_plan_decide design.

Secrets masking

Teammates inherit the same secrets-scrubbing pipeline as the main chat session. API keys, tokens, and other sensitive values detected in tool call arguments or outputs are masked before they reach the model context. Masking is applied on every emit path — it cannot be disabled by a role or task instruction.

Role-declared skills

A role's skills: list (in its role file frontmatter) is force-activated on spawn — the Coordinator loads every skill source reachable from the run's workdir/mission-root and activates each named skill by name, unconditionally, regardless of whether the skill's own applies_to would otherwise match a team context. This is a request, not a filter: unlike the auto-suggested skills a plain chat session picks up by keyword match, a role that lists skills: [verify-before-done, destructive-action-guard] gets exactly those, because it asked for them by name.

A role that lists no skills gets none injected — skills are opt-in per role, not applied unconditionally by tool capability. If you want every write-capable role in a recipe to carry verify-before-done and destructive-action-guard, list them explicitly on each such role.

A skill name that doesn't resolve to a loaded skill is reported back (surfaced on StreamTeammateSpawned.SkillWarnings and as a spawn-time warning) rather than silently dropped.

File-write locking

Two write-capable teammates editing the same file simultaneously is prevented by a coordinator-owned file lock map. If a teammate's fs_write_file or fs_patch_file call targets a file currently held by another teammate, the tool returns an error suggesting coordination via team_send_message. Read-only teams do not incur this overhead.

Pause as a safety valve

team_pause(reason) — a coordinator/driver-only tool, unreachable by any teammate — transitions the Coordinator to paused. All teammates finish their current tool call and idle; new task claims are refused. The driver's own team_resume() tool call lifts the pause while the run is still live.

No CLI resume yet

There is no bnerd team resume subcommand today. A team that ends without calling team_finish — the TUI's exit grace on quit, bnerd web's or bnerd team run's on SIGINT/SIGTERM, or bnerd team run's own end-of-run cleanup — is recorded as interrupted on disk (see Exit behavior). A TUI-launched team can be resumed by reopening its conversation (bnerd x --ai-continue) — see AI Assistant → Resuming an interrupted team. A bnerd web- or bnerd team run-launched team has no resume path: the only way forward for those is to start the run again.

Pausing is the intended response when something unexpected surfaces mid-run:

> pause the team, something looks wrong with task #7

The coordinator records the reason and the time; the bnerd team status output shows the pause state so you can review and decide.