Run configuration

Every option on the New Run Configuration tab, with defaults and limits.


Every option on New Run → Configuration, in the order it appears. For how to use them, see Configure a run.

Autorun#

Autorun
Default
Off
Values
On / Off
Effect
Off: check in after every iteration. On: iterate without checking in while the Score is below the threshold. Also switchable on the live dashboard.

Starting a run with Autorun on asks for confirmation (Start with Autorun?), unless you've ticked Don't show again.

Agent Roles#

One card per role: Analyst, PM, Developer, Integration, QA, Visual QA, Test Author, Feedback and Summary.

Provider
Default
Assigned from your keys
Values
Any of the 8 providers
Effect
Where this role's calls go.
Model
Default
The provider's starter model
Values
Any model in that provider's catalog
Effect
The model the role uses. Shows its cost tier.
Max output tokens
Default
The model's maximum output
Values
256 or more, in steps of 1,000; at most the model's cap
Effect
Caps each response. Raised automatically for reasoning models that are set too low.
Apply provider + model to all roles
Default
—
Values
Button
Effect
Copies this card's provider and model to all nine roles.

Starter models, used when roles are assigned to a newly keyed provider:

Anthropic
claude-haiku-4-5
OpenAI
gpt-5-mini
Gemini
gemini-2.5-flash-lite
DeepSeek
deepseek-flash
xAI
grok-4.3
Meta
muse-spark-1.3
Kimi
kimi-k2.7-code
Qwen
qwen3.8-flash

Each is the lowest-cost model judged able to carry all nine roles. Model ids change as providers rename models; if a starter model leaves the catalog, the role card flags it for you to repick.

Scorecard#

Score threshold
Default
85
Values
0–100, in steps of 5
Effect
The Score needed for Successful (with zero criticals). The target for Autorun. Never stops a run by itself.
Correctness weight
Default
1.0
Values
0–5, in steps of 0.25
Effect
Pull in the Score's weighted geometric mean. 0 ignores the dimension.
Completeness weight
Default
1.0
Values
0–5, in steps of 0.25
Effect
As above.
Integrity weight
Default
1.0
Values
0–5, in steps of 0.25
Effect
As above.
Quality weight
Default
1.0
Values
0–5, in steps of 0.25
Effect
As above.
Accessibility weight
Default
1.0
Values
0–5, in steps of 0.25
Effect
As above.

↺ Reset restores all of these.

Orchestration#

Max iterations
Default
10
Values
1–50
Effect
Hard stop on the number of iterations.
Max cost
Default
Off
Values
Off, or $0.25 upward in $0.25 steps
Effect
Predictive spend check before each iteration after the first.
QA agents
Default
3
Values
1–20
Effect
Parallel QA reviewers per iteration.

↺ Reset restores all of these.

Dev Mode#

Mode
Default
Full-Stack
Values
Full-Stack / Role-Based
Effect
Whether developers can touch any file, or specialize.
Developer agents (Full-Stack)
Default
5
Values
1–20
Effect
Developers writing code in parallel.
Frontend (Role-Based)
Default
1
Values
1 or more
Effect
HTML and JavaScript specialists.
Backend (Role-Based)
Default
1
Values
1 or more
Effect
Server-code specialists.
Design (Role-Based)
Default
1
Values
1 or more
Effect
CSS, image and visual specialists.

In Role-Based mode, the total number of developers is Frontend + Backend + Design.

QA Mode#

QA Mode
Default
Code
Values
Code / Visual / Interactive
Effect
Code: source review only. Visual: screenshots reviewed by Visual QA. Interactive: screenshots plus up to six automated clicks.

Fixed quality rules#

These aren't editable in the app:

Half-lives (critical / major / minor)
1 / 4 / 20 issues halve a dimension
Correctness passes at
0 criticals and ≥ 95% of functional criteria passing
Completeness passes at
≥ 90% of coverage criteria passing and no missing planned files
Quality passes at
0 criticals and ≤ 2 majors
Accessibility passes at
0 criticals and ≤ 3 majors
Trend step
A change of 5 points or more counts as improving or regressing