Run configuration
Every option on the New Run Configuration tab, with defaults and limits.
Every option on New Run → Configuration, in the order it appears. For how to use them, see Configure a run.
Autorun#
- Default
- Off
- Values
- On / Off
- Effect
- Off: check in after every iteration. On: iterate without checking in while the Score is below the threshold. Also switchable on the live dashboard.
Starting a run with Autorun on asks for confirmation (Start with Autorun?), unless you've ticked Don't show again.
Agent Roles#
One card per role: Analyst, PM, Developer, Integration, QA, Visual QA, Test Author, Feedback and Summary.
- Default
- Assigned from your keys
- Values
- Any of the 8 providers
- Effect
- Where this role's calls go.
- Default
- The provider's starter model
- Values
- Any model in that provider's catalog
- Effect
- The model the role uses. Shows its cost tier.
- Default
- The model's maximum output
- Values
- 256 or more, in steps of 1,000; at most the model's cap
- Effect
- Caps each response. Raised automatically for reasoning models that are set too low.
- Default
- —
- Values
- Button
- Effect
- Copies this card's provider and model to all nine roles.
Starter models, used when roles are assigned to a newly keyed provider:
claude-haiku-4-5
gpt-5-mini
gemini-2.5-flash-lite
deepseek-flash
grok-4.3
muse-spark-1.3
kimi-k2.7-code
qwen3.8-flash
Each is the lowest-cost model judged able to carry all nine roles. Model ids change as providers rename models; if a starter model leaves the catalog, the role card flags it for you to repick.
Scorecard#
- Default
- 85
- Values
- 0–100, in steps of 5
- Effect
- The Score needed for Successful (with zero criticals). The target for Autorun. Never stops a run by itself.
- Default
- 1.0
- Values
- 0–5, in steps of 0.25
- Effect
- Pull in the Score's weighted geometric mean. 0 ignores the dimension.
- Default
- 1.0
- Values
- 0–5, in steps of 0.25
- Effect
- As above.
- Default
- 1.0
- Values
- 0–5, in steps of 0.25
- Effect
- As above.
- Default
- 1.0
- Values
- 0–5, in steps of 0.25
- Effect
- As above.
- Default
- 1.0
- Values
- 0–5, in steps of 0.25
- Effect
- As above.
↺ Reset restores all of these.
Orchestration#
- Default
- 10
- Values
- 1–50
- Effect
- Hard stop on the number of iterations.
- Default
- Off
- Values
- Off, or $0.25 upward in $0.25 steps
- Effect
- Predictive spend check before each iteration after the first.
- Default
- 3
- Values
- 1–20
- Effect
- Parallel QA reviewers per iteration.
↺ Reset restores all of these.
Dev Mode#
- Default
- Full-Stack
- Values
- Full-Stack / Role-Based
- Effect
- Whether developers can touch any file, or specialize.
- Default
- 5
- Values
- 1–20
- Effect
- Developers writing code in parallel.
- Default
- 1
- Values
- 1 or more
- Effect
- HTML and JavaScript specialists.
- Default
- 1
- Values
- 1 or more
- Effect
- Server-code specialists.
- Default
- 1
- Values
- 1 or more
- Effect
- CSS, image and visual specialists.
In Role-Based mode, the total number of developers is Frontend + Backend + Design.
QA Mode#
- Default
- Code
- Values
- Code / Visual / Interactive
- Effect
- Code: source review only. Visual: screenshots reviewed by Visual QA. Interactive: screenshots plus up to six automated clicks.
Fixed quality rules#
These aren't editable in the app:
- 1 / 4 / 20 issues halve a dimension
- 0 criticals and ≥ 95% of functional criteria passing
- ≥ 90% of coverage criteria passing and no missing planned files
- 0 criticals and ≤ 2 majors
- 0 criticals and ≤ 3 majors
- A change of 5 points or more counts as improving or regressing