§ 04 · DOCS

How the agents work

Memory, discipline, and economics — the operating traits behind every awk agent.

The difference between a chatbot and a colleague is discipline: memory that carries between conversations, a process the agent follows the same way every time, and judgment about when to act versus when to ask. This page explains the operating traits every awk agent brings to the work — and, just as importantly, where each one stops.

Memory

Auto memory

Agents recall the context that matters at the start of every turn, and save what's worth keeping so it persists between conversations and sessions. You tell the team something once — a convention, a preference, a constraint — and you don't have to repeat it. Recall is automatic; saving is deliberate, so the memory stays a curated set of durable facts rather than a transcript of everything ever said.

Decision records

When an agent makes a consequential, non-obvious choice, it records the decision, the reasoning behind it, and the alternatives it weighed — linked to the evidence (tickets, PRs, prior findings). These records are pinned durably, so when a long thread is compacted the why behind a choice is recalled exactly later, by the same agent or by a teammate resuming the thread, instead of being vaguely reconstructed.

Decision records capture the choices an agent judged consequential — not a verbatim transcript of the whole conversation. They're a precise, structured record of what was decided and why.

Team memory

A learning one colleague captures is available to the whole team. When one agent works out how your repo is laid out, which reviewer owns a service, or how you like PRs structured, that knowledge is there for every agent on your team — so patterns stay consistent across everyone's work. Team memory is scoped to your team; it is never shared across tenants.

Discipline

Set workflows

Agents work role-specific checklists end to end — the same way every time, the way an experienced engineer follows a runbook. The engineer gathers requirements, plans, gets the plan reviewed, implements, reviews the code, and only then opens a PR. QA runs its own checklist: it turns the requirements into test cases, tests the feature, captures the findings and artefacts — screenshots and screen recordings — and hands the work back with those findings when something needs fixing. The steps don't get skipped because the agent is in a hurry.

OG
✓understand✓plan✓build✓review✓PR
QA
✓requirements✓test✓evidence✓report✓hand back

Gated actions

Actions that are irreversible or reach outside your team pause for your explicit approval before they run: pushing code, merging a pull request, sending an email or message to another person, or changing what a file or document is shared with. The agent proposes the action; you approve or hold it. The gate is fail-closed — an action it can't classify is held, never run silently. This is containment by design: routine, reversible work flows without friction, and the consequential steps wait for a human — you approve or hold each one right in the Slack thread the agent is working in.

OG requests approval
git push → main
Approve
Hold

Autonomy dial

You set how much the agents lean on you versus their owner-agent for the judgment calls — the "should I do it this way or that way?" moments — across a five-position dial, from always ask you through balanced to the team decides. The default leans toward the team, so you're not consulted on every small fork, but you can turn it all the way toward yourself.

The dial is advisory: it changes who is asked about a decision, not what an agent is allowed to do. It never widens an agent's authority — irreversible and outward-facing actions always pause for approval regardless of where the dial is set (see Gated actions above).

You decideTeam decides
Mostly the team

The owner-agent fields most calls; you see the ones that matter. (Default.)

Advisory — sets who fields judgment calls. It never widens what an agent may do; irreversible actions always pause for approval.

Scale & economics

Model router

Not every task needs the most powerful model. A built-in router matches each task to the most economical model that can do it well — a cheap classifier sizes the work up front, and heavier models are reserved for the work that actually earns them. You get frontier-grade results on the hard parts without paying frontier prices for the routine ones.

taskclassifylighterheavier

Subagent fan-out

For a large job, an agent decomposes the work into subtasks and runs them in parallel rather than grinding through serially — reading several parts of a codebase at once, or implementing independent pieces concurrently. Parallelism is bounded by your plan's workstream limit; see Billing & credits for the per-plan figures.

job
parallel

Context management

Long-running work would balloon an agent's context if left unchecked. The harness keeps only the context that's still relevant and auto-compacts the rest, so the working window stays small. That keeps every turn fast and economical — and is part of why economical models can deliver senior-level results.

Re-expandable history

Compaction here isn't a one-way squash into a lossy summary. When the harness compacts older work, it leaves a pointer to every step it set aside — so the agent can re-expand the exact original detail on demand when a later decision needs it, instead of working from a blurred paraphrase. The window stays small day-to-day, but the full detail is a step away when it matters.

Cost per workstream

Every workstream freezes the credits it consumed when the work completes, and posts the total back to you, so you can see what a given piece of work actually cost. Credits are billed at a fixed rate — see Billing & credits for the current price per credit and how consumption works.

Workstream · PROJ-142Metering
💳 Credits used
0$0.00
Illustrative · one run · $0.01 / credit

A per-step, per-ticket cost breakdown (a full cost explorer) is on the roadmap; today, cost is attributed and reported at the workstream level.

Rapid startup

To verify a change, agents don't just reason about it — they spin your service up and check it actually runs. That's the difference between "this should compile" and "I ran it and here's what happened." See Preview links & localhost access for how you can watch a running service yourself.

Multimodal input

Agents read more than text. Attach an image, a PDF, or an audio clip in a thread and the agent works from its contents directly; video is supported on request as an opt-in per team. See Attachments / rich media for the supported types, size limits, and per-team controls.