Skip to content

Configuration Reference

SAM uses environment variables for platform configuration. User-specific settings (cloud provider tokens, agent API keys) are stored encrypted in the database, not as environment variables.

These are Cloudflare Worker secrets, set during deployment. Pulumi auto-generates security keys on first deploy.

SecretDescription
ENCRYPTION_KEYAES-256-GCM master key. Used for BetterAuth session cookies and user credential encryption unless a purpose-specific override below is set (auto-generated)
BETTER_AUTH_SECRETOptional purpose-specific override for BetterAuth session cookie signing/encryption. Falls back to ENCRYPTION_KEY when unset (apps/api/src/lib/secrets.ts)
CREDENTIAL_ENCRYPTION_KEYOptional purpose-specific override for AES-GCM encryption of user cloud/agent credentials. Falls back to ENCRYPTION_KEY when unset (apps/api/src/lib/secrets.ts)
JWT_PRIVATE_KEYRSA-2048 private key for signing tokens (auto-generated)
JWT_PUBLIC_KEYRSA-2048 public key for token verification (exposed via JWKS)
DEPLOY_SIGNING_PRIVATE_KEYEd25519 private key for signing deployment apply payloads (auto-generated)
DEPLOY_SIGNING_PUBLIC_KEYEd25519 public key derived during deployment for deployment node verification (auto-generated)
VAPID_PRIVATE_KEYBase64url P-256 private scalar used to authenticate Web Push delivery (auto-generated)
VAPID_PUBLIC_KEYUncompressed base64url P-256 public key returned to browsers at runtime (derived during deployment)
VAPID_SUBJECTRFC 8292 contact URI for Web Push, defaulting to the deployment app origin (generated during deployment)
CF_API_TOKENCloudflare API token for infrastructure, DNS, Origin CA certificate issuance, observability, AI Gateway, Containers, and admin logs. Requires Account → Containers → Edit and Account → SSL and Certificates → Edit.
CF_AIG_TOKENOptional narrower Cloudflare AI Gateway Unified Billing token
CF_ZONE_IDCloudflare zone ID for DNS record management
CF_ACCOUNT_IDCloudflare account ID
DEVCONTAINER_CACHE_CLOUDFLARE_API_TOKENOptional narrower Cloudflare token for managed devcontainer registry credentials
DEVCONTAINER_CACHE_CLOUDFLARE_ACCOUNT_IDOptional Cloudflare account override for managed devcontainer registry credentials
GITHUB_CLIENT_IDOptional fallback GitHub App client ID for OAuth; runtime admin config takes precedence
GITHUB_CLIENT_SECRETOptional fallback GitHub App client secret for OAuth; runtime admin config takes precedence
GITHUB_APP_IDOptional fallback GitHub App ID for installation tokens; runtime admin config takes precedence
GITHUB_APP_PRIVATE_KEYOptional fallback GitHub App private key (PEM or base64); runtime admin config takes precedence
GITHUB_APP_SLUGOptional fallback GitHub App URL slug; runtime admin config takes precedence
GITHUB_WEBHOOK_SECRETOptional fallback GitHub App webhook HMAC secret; runtime admin config takes precedence
GITLAB_HOSTOptional fallback GitLab OAuth host, such as https://gitlab.com; runtime admin config takes precedence
GITLAB_CLIENT_IDOptional fallback GitLab OAuth application ID; runtime admin config takes precedence
GITLAB_CLIENT_SECRETOptional fallback GitLab OAuth secret; runtime admin config takes precedence
TRIAL_CLAIM_TOKEN_SECRETTrial onboarding HMAC secret (auto-generated)

Unless the deploy sets them itself (as it does BASE_DOMAIN), Worker variables come from [vars] in apps/api/wrangler.toml, falling back to the default in the code. On a self-hosted instance, a GitHub Environment variable of the same name replaces that value at deploy time, but only if the deploy forwards it: scripts/deploy/sync-wrangler-config.ts must read it, and the Sync Wrangler Config steps in .github/workflows/deploy-reusable.yml must pass it from vars. After a deploy, the Worker’s Settings → Variables and Secrets page in the Cloudflare dashboard shows the value it got. To change any other variable, edit wrangler.toml in your fork; updates then need a manual merge (see Updating an Existing Self-Hosted Instance).

VariableDefaultDescription
BASE_DOMAIN—Root domain for the deployment (e.g., example.com)
PREVIEW_BASE_DOMAINpreview.BASE_DOMAINFull isolated hostname used for interactive HTML previews
PREVIEW_URL_TTL_SECONDS300Lifetime of project/file/version-scoped interactive preview URLs in seconds
PREVIEW_SIGNING_KEYgeneratedDeployment-owned HMAC key generated and persisted by Pulumi; not a manual prerequisite
VERSION—Deployment version string
SETUP_TOKEN—Plaintext first-run setup token generated during deploy and readable in the Cloudflare dashboard while setup is incomplete
SETUP_FORCE(unset)Set to true to reopen /setup for lockout recovery
SETUP_RATE_LIMIT_MAX_ATTEMPTS10Max setup-token attempts per identifier/window
SETUP_RATE_LIMIT_WINDOW_SECONDS900Setup-token attempt window in seconds
D1_SESSION_MODEfirst-primaryD1 Sessions API anchor for the Worker fetch handler. first-primary runs each request against one D1 session whose first query goes to the primary and whose later queries may be served by a caught-up read replica — same freshness, one wide-area round trip per request instead of one per query. disabled sends every query straight at the primary. An unrecognised value falls back to first-primary and is logged once per isolate. scheduled() and Durable Objects always use the unsessioned binding.
PLATFORM_CONFIG_CACHE_MS60000Per-isolate cache TTL for the resolved platform integration config (the GITHUB_*/GITLAB_*/GOOGLE_LOGIN_* fallbacks above and their runtime admin overrides). Resolving costs 14 D1 queries and runs on the auth preamble of every authenticated request. After a config change, isolates that already hold a cached copy converge within this window. Set to 0 to disable caching and always re-read D1.
GITHUB_INSTALLATION_TOKEN_CACHE_TTL_SECONDS3000KV cache TTL for GitHub App installation tokens. The default is shorter than GitHub’s one-hour token lifetime. Set to 0 to disable writes for new cache entries.
GITHUB_INSTALLATION_TOKEN_REFRESH_MARGIN_SECONDS300Cached GitHub App installation tokens that expire within this many seconds are minted again instead of reused, so a long cache TTL can never hand out an expiring token. Capped at 1800, half the one-hour token lifetime.
GITHUB_REPO_ACCESS_CACHE_TTL_SECONDS300KV cache TTL for per-user, per-installation, per-repository GitHub access checks used by the Files page. Set to 0 to disable writes for new cache entries.
GITHUB_TREE_CACHE_TTL_SECONDS86400KV cache TTL for immutable Git tree responses keyed by commit SHA. Branch refs are resolved to a commit SHA before lookup. Set to 0 to disable writes for new cache entries.
PROJECT_MULTIPLAYER_CACHE_TTL_MS10000Per-isolate cache TTL for project multiplayer state counts used by trigger-bearing pages. Set to 0 to disable the cache.
CREDENTIAL_ATTRIBUTION_CACHE_TTL_MS10000Per-isolate cache TTL for project credential attribution health used by trigger-bearing pages. Set to 0 to disable the cache.

Set in GitHub Settings → Environments → production:

VariableDescriptionExample
BASE_DOMAINDeployment domainexample.com
RESOURCE_PREFIXDomain-derived Cloudflare resource name prefixsa379a6
PULUMI_STATE_BUCKETR2 bucket for Pulumi statesa379a6-pulumi-state
CF_CONTAINER_ENABLEDOptional instant-session runtime toggle. Generated deploys default to true; set false to force VM runtime.false
WORKER_SECRET_BULK_MAX_OPSOptional deploy-script limit for queued Worker secret create/update/delete operations in one wrangler secret bulk payload. Defaults to 100; range 1–100.75
D1_READ_REPLICATION_MODEOptional D1 read-replication mode applied on every deploy. Defaults to auto; set disabled to remove replicas.disabled
D1_SESSION_MODEOptional Worker D1 Sessions anchor. Defaults to first-primary; set disabled to send every query to the primary. An unrecognised value falls back to the default and is logged once per isolate.disabled
D1_RESTORE_RECOVERY_WINDOW_DAYSOptional D1 restore window for accounts with narrower retention. Defaults to 30; range 1–30.7
D1_MIGRATION_CHURNING_TABLESOptional comma-separated <binding>.<table> subset of the reviewed retention/expiry table list. May narrow the built-in list but cannot expand it.OBSERVABILITY_DATABASE.platform_errors
D1_MIGRATION_CHURNING_TABLE_MAX_DECREASE_PERCENTMaximum allowed decrease for reviewed churning tables. Defaults to 50; range 0–100. A decrease exactly at the limit is accepted.25

The reviewed default churning selectors are DATABASE.deployment_releases, DATABASE.github_webhook_deliveries, DATABASE.project_files, DATABASE.registry_credential_rate_limits, DATABASE.session_snapshots, DATABASE.sessions, DATABASE.trial_waitlist, DATABASE.trigger_executions, DATABASE.verifications, DATABASE.webhook_deliveries, and OBSERVABILITY_DATABASE.platform_errors. All other application tables retain zero row-decrease tolerance. Leave D1_MIGRATION_CHURNING_TABLES unset to use the complete reviewed default list.

RESOURCE_PREFIX is generated from BASE_DOMAIN as s plus the first six hex characters of the domain’s SHA-256 hash. The self-host onboarding flow fills it in for you.

These optional Worker variables bound the server-side OCI registry lookups used when a deployment release submits tag-based images. Digest-pinned images are stored without registry network resolution.

VariableDefaultDescription
DEPLOYMENT_IMAGE_RESOLVE_REQUEST_TIMEOUT_MS10000Per-registry request timeout
DEPLOYMENT_IMAGE_RESOLVE_TOTAL_TIMEOUT_MS60000Total tag-resolution wall-clock budget per release submission
DEPLOYMENT_IMAGE_RESOLVE_MAX_FETCH_ATTEMPTS200Maximum outbound registry/token fetches per resolver instance
DEPLOYMENT_IMAGE_RESOLVE_MAX_REDIRECTS2Maximum manually validated HTTPS redirects per outbound request
DEPLOYMENT_IMAGE_RESOLVE_TOKEN_RESPONSE_MAX_BYTES65536Maximum bearer-token JSON response size
DEPLOYMENT_IMAGE_RESOLVE_MAX_CONCURRENT_FETCHES4Maximum simultaneous outbound resolver fetches
DEPLOYMENT_IMAGE_RESOLVE_MAX_SERVICES50Maximum tag-based image references resolved per release submission

Required GitHub Actions secrets include CF_API_TOKEN, CF_ACCOUNT_ID, CF_ZONE_ID, R2_ACCESS_KEY_ID, R2_SECRET_ACCESS_KEY, and PULUMI_CONFIG_PASSPHRASE. GitHub App/OAuth secrets (GH_CLIENT_ID, GH_CLIENT_SECRET, GH_APP_ID, GH_APP_PRIVATE_KEY, GH_APP_SLUG, GH_WEBHOOK_SECRET) and Google login OAuth secrets (GOOGLE_LOGIN_CLIENT_ID, GOOGLE_LOGIN_CLIENT_SECRET) are optional environment fallbacks; fresh deployments can set them through /setup instead. The separate Google infra/GCP OAuth pair (GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET) is used only for WIF and can be configured by a superadmin at /admin/integrations; runtime values override the environment fallback. Service-account JSON users need no infrastructure OAuth client. Deploy signing keys are generated and persisted by Pulumi during deployment; GitHub Environment values are only needed for explicit key overrides.

These variables affect the local sam CLI process only. They are not Worker runtime variables or GitHub Actions secrets.

VariableDefaultDescription
SAM_CLI_MAX_API_RESPONSE_BYTES1048576Maximum API response body bytes the CLI reads before truncating/aborting.

Codex and Claude Code guided subscription login have no feature-on environment variable. They are available by default when the deployment includes the SANDBOX, CREDENTIAL_SETUP_SESSION, and SETUP_SESSION_POOL Worker bindings generated by SAM’s deployment configuration. Omitting one of those bindings disables the guided flow. SANDBOX_ENABLED continues to control separate administrative Sandbox runtime surfaces and is not required for guided login.

VariableDefaultDescription
MAX_CONCURRENT_SETUP_SESSIONS2Maximum concurrent guided credential-setup sessions.
SETUP_SESSION_TTL_MS900000Guided session lifetime before automatic teardown.
SETUP_SESSION_CAPTURE_POLL_MS3000Interval for checking device-login and credential-capture state.
CODEX_DEVICE_AUTH_REQUEST_TIMEOUT_MS30000Timeout for each Codex app-server JSON-RPC request.
CLAUDE_SETUP_ENTER_DELAY_MS1000Delay before sending Enter as a separate stdin write after pasting Claude’s browser-displayed code.
CLAUDE_SETUP_EXCHANGE_TIMEOUT_MS120000Maximum wait for Claude’s CLI code exchange before a visible timeout.
CLAUDE_SETUP_REJECTION_SETTLE_MS400Wait for Claude CLI Ink redraws to settle before classifying an OAuth error.
CLAUDE_SETUP_VERIFICATION_POLL_MS500Interval for checking the sandbox handoff file for Claude’s browser-displayed code.
CLAUDE_SETUP_TTY_COLUMNS512PTY width for claude setup-token, reducing opaque-token wrapping.
CLAUDE_SETUP_OUTPUT_BUFFER_BYTES32768Maximum in-memory Claude PTY output retained for parsing.
CLAUDE_VERIFICATION_CODE_MAX_LENGTH1024Maximum accepted length of Claude’s browser-displayed code#state value.
CLAUDE_SETUP_ERROR_DETAIL_MAX_LENGTH160Maximum sanitized Claude CLI diagnostic length shown to the user.
CLAUDE_OAUTH_TOKEN_MAX_LENGTH8192Maximum captured Claude OAuth token length.
SETUP_SESSION_SWEEP_MAX_CANDIDATES50Maximum expired sessions cleaned up by one scheduled sweep.
POOL_LEASE_BUFFER_MS300000Grace period after session TTL before a leaked capacity lease self-prunes.

The variables below tune the Instant (Cloudflare Container) runtime — how long a session stays awake, how long a wake may take, and how many snapshot restores are attempted before a session is failed. See Instant Sessions for what each of these means to a user.

VariableDefaultDescription
CF_CONTAINER_ENABLEDtrueEnables Cloudflare Container instant sessions for matching profiles and zero-config runtime selection. Set false to force cloud VM runtime.
CF_CONTAINER_SLEEP_AFTER1hNormal inactivity window before an Instant container sleeps. Sleep remains recoverable through the runtime-neutral session snapshot.
CF_CONTAINER_ACTIVE_WORK_MAX_MS7200000Defensive maximum lifetime for an active-work keepalive lease.
CF_CONTAINER_KEEPALIVE_RENEW_INTERVAL_MS300000Interval used to renew the container activity timeout while prompt work is active.
CF_CONTAINER_WAKE_TIMEOUT_MS120000Maximum time for a sleeping container to launch, restore its snapshot, and accept the triggering request.
CF_CONTAINER_RECOVERY_MAX_ATTEMPTS2Maximum snapshot restore attempts before SAM reconciles the runtime, workspace, agent session, and active task to a visible terminal recovery failure.
INSTANT_STALE_CALLBACK_MARGIN_MS60000 (60 sec)Freshness margin used to reject destructive (error/failed) callbacks arriving from a superseded Instant container generation after the runtime row was reconciled by a completed recovery.
CF_CONTAINER_CREATE_WORKSPACE_TIMEOUT_MS120000Budget for the synchronous instant-session create-workspace request, which includes the repository clone inside the container.
CF_CONTAINER_CLONE_FILTERblob:noneGit partial-clone filter forwarded to instant containers as STANDALONE_CLONE_FILTER. Set off to force full clones.

CF_CONTAINER_RECOVERY_MAX_ATTEMPTS has a deployment-safety minimum of 2; smaller positive values resolve to 2.

Sleeping and reclaimed Instant and VM sessions are restored from a snapshot of the agent’s home directory and the repository work in progress. An ordinary sleep requires a complete snapshot before SAM tears down VM compute. When snapshots keep failing, a bounded fallback can instead sleep an idle VM session on the Git recovery point an earlier snapshot saved, keeping the transcript but not every file (SAM could not save a complete snapshot). None of these limits are surfaced in the UI, so operators should set expectations deliberately — see What gets restored.

VariableDefaultDescription
SESSION_SNAPSHOT_TTL_DAYS7Snapshot retention. A session sleeping longer than this cannot be fully restored.
SESSION_SNAPSHOT_TOTAL_BUDGET_BYTES268435456 (256 MiB)Max combined size of the home + work-in-progress snapshot. The higher default favors bounded retained R2 state over keeping a VM alive when a typical agent harness has accumulated substantial durable state.
SESSION_SNAPSHOT_ENTRY_THRESHOLD_BYTES268435456 (256 MiB)Largest single file the snapshot scanner will include. This matches the total budget so durable agent state databases are not skipped solely because they are larger than the former 50 MiB cap.
SESSION_SNAPSHOT_TRANSFER_IDLE_TIMEOUT_MS30000 (30 sec)No-progress timeout for each snapshot upload or download.
SESSION_SNAPSHOT_UPLOAD_URL_TTL_SECONDS900 (15 min)Lifetime of direct R2 upload URLs used so large snapshots do not traverse the Worker request-body boundary. Current agents bind exact length and SHA-256; busy legacy VM agents stream through a current same-user VM relay that independently authenticates both nodes and removes callback credentials before R2. When R2 S3 credentials are unavailable, SAM retains the Worker upload path.
SESSION_SNAPSHOT_REQUEST_TIMEOUT_MS300000 (5 min)Budget for the vm-agent to accept the final checkpoint request. Durable completion is governed by progress reporting rather than this fixed wall clock.
SESSION_SNAPSHOT_PROGRESS_IDLE_TIMEOUT_MS120000 (2 min)No-progress watchdog for an accepted final checkpoint. Current vm-agents periodically advance D1 progress while walking HOME or uploading artifacts; if progress stops, SAM records a degraded snapshot and that sleep attempt fails, counting against the sleep failure budget below.
SESSION_SNAPSHOT_POLL_INTERVAL_MS1000 (1 sec)Interval used while the Worker waits for a VM agent’s asynchronous final checkpoint to commit in D1.
SESSION_SNAPSHOT_OPERATION_TIMEOUT15mVM-agent checkpoint/restore deadline, using Go duration syntax. Before the first restore RPC, TaskRunner pins this duration plus SESSION_SNAPSHOT_REQUEST_TIMEOUT_MS as its retry window; retries and restarts cannot renew it.
SESSION_SNAPSHOT_PROGRESS_REPORT_INTERVAL15sVM-agent throttle for best-effort progress callbacks during data-scaled snapshot work. This uses Go duration syntax and is passed to newly provisioned VMs and Instant containers.
SESSION_SNAPSHOT_PROGRESS_REPORT_TIMEOUT5sVM-agent timeout for each best-effort snapshot progress callback. This uses Go duration syntax and is passed to newly provisioned VMs and Instant containers.
SESSION_SNAPSHOT_JSON_BODY_MAX_BYTES262144 (256 KB)Maximum snapshot coordination request size accepted by the Worker.
SESSION_SNAPSHOT_R2_PREFIXsession-snapshotsPrivate object prefix. Session objects are deleted by the Worker from D1 lifecycle state, not by object age.
SESSION_SNAPSHOT_RECOVERY_MAX_ATTEMPTS3Replacement-VM wake attempts allowed in a burst. Once spent, further wakes are refused (recovery_attempts_exhausted) until SESSION_SNAPSHOT_RECOVERY_ATTEMPT_DECAY_MS has passed since the last failed attempt; the snapshot’s own expiry remains the hard limit.
SESSION_SNAPSHOT_RECOVERY_ATTEMPT_DECAY_MS900000 (15 min)How long a spent wake-attempt burst stays spent before a new wake may be tried
SESSION_RECOVERY_LINEAGE_MAX_DEPTH256How many wake-to-wake links a wake follows back to the conversation’s first run to decide whether that run explicitly pinned its region. A wake without such a request only prefers the region it slept in.
SESSION_SLEEP_AFTER_MS900000 (15 min)ProjectData-recorded idle interval before SAM automatically sleeps a VM session. Runtime heartbeats do not extend this clock. Completed and failed tasks queue sleep immediately. Their still-active final prompt becomes eligible after this interval from its later activity or the task’s end, so a working agent is not slept mid-turn. Ledger cleanup also uses it to protect the final response.
SESSION_SLEEP_SWEEP_BATCH_SIZE10Maximum due session sleep candidates selected and individually claimed by one scheduled sweep.
SESSION_SLEEP_SWEEP_WALL_BUDGET_MS20000 (20 sec)Soft wall-clock budget for bounded D1/ProjectData eligibility and claim work. After a durable claim, final snapshot and teardown run through the scheduled event’s out-of-band lifetime. Remaining unclaimed rows stay due for the next sweep.
SESSION_SLEEP_RETRY_DELAY_MS300000 (5 min)Retry delay after a fail-closed automatic sleep attempt.
SESSION_SLEEP_FAILURE_MAX_ATTEMPTS3Failed full-snapshot sleep attempts in one episode before SAM tries the fallback: release an idle VM session’s compute, keeping its transcript and the exact Git commit, branch and uncommitted changes an earlier snapshot saved. Instant sessions, and VM sessions without such a recovery point, end the episode blocked instead. An episode ends when the session sleeps or wakes, or a person sends a message.
SESSION_SLEEP_FAILURE_MAX_ELAPSED_MS900000 (15 min)Time since an episode’s first sleep attempt after which the fallback is tried whatever the attempt count, once at least one attempt has failed. With the default sweep and retry delay this is about three attempts.
SESSION_SLEEP_MAX_ATTEMPTS9Ceiling on failed attempts in one sleep episode, counting full snapshots and fallback attempts. It is raised to at least SESSION_SLEEP_FAILURE_MAX_ATTEMPTS + 1 so the fallback always gets a try. At the ceiling the episode ends blocked: automatic sleep stops and the chat says so. A task failure starts a fresh episode; a failed task whose episode ends blocked has its runtime torn down, with a chat notice. Raising this value re-arms rows exhausted before bounded episodes existed that are still below the new limit.
FAILED_TASK_PRESERVATION_MAX_WAIT_MS28800000 (8 hours)Longest a failed task’s runtime stays awake waiting for its preservation sleep, measured from the latest of the failure, an in-place wake and the start of the agent’s current turn. A turn that never ends defers that sleep without spending an attempt; past this wait the sweep tears the runtime down and says so in the chat (releaseStalledFailedTaskPreservation()).
SESSION_SLEEP_CLAIM_LEASE_MS600000 (10 min)Time after which an interrupted automatic-sleep claim can be safely reclaimed.
HARNESS_BACKGROUND_WORK_LEASE_MS300000 (5 min)Finite sleep-protection lease renewed by normalized harness background-work lifecycle signals. Expiry fails open to ordinary idle-sleep eligibility so a missing terminal signal cannot pin compute forever.
HARNESS_BACKGROUND_WORK_MAX_DURATION_MS1800000 (30 min)Absolute ceiling, measured from the last harness lifecycle progress edge rather than the last heartbeat, on how long background work may defer sleep. The sliding lease above is refreshed by periodic re-reports, so an adapter faithfully re-reporting a stale task set (for example an abandoned run_in_background dev server) would otherwise pin compute awake indefinitely.
ACP_ACTIVITY_ADMISSION_ENABLEDtrueEnables Worker-side admission control for ACP activity callbacks. Redundant intermediate state is coalesced to protect ProjectData load; terminal/error and final idle transitions still bypass coalescing.
ACP_ACTIVITY_COALESCE_WINDOW_MS2000 (2 sec)Minimum interval between redundant intermediate ProjectData activity writes, and the first retry delay for a coalesced report whose flush hit a retryable ProjectData failure (including a CPU-limit reset or lost connection); each further retry doubles the delay, capped at ACP_ACTIVITY_COALESCE_TTL_MS. Activity transitions, new prompt epochs, and terminal/error reports bypass this window.
ACP_ACTIVITY_COALESCE_TTL_MS60000 (1 min)Maximum lifetime for a pending coalesced activity report before it is evicted and left to probe-backed session-activity reconciliation.
ACP_ACTIVITY_COALESCE_MAX_PENDING512Maximum pending coalesced activity reports retained by one Worker isolate. Capacity evictions are logged as activity telemetry rather than silently dropped.
ACP_ACTIVITY_BINDING_CACHE_TTL_MS30000 (30 sec)Short-lived cache for already-authorized ACP session bindings used to avoid ProjectData reads during callback storms. Callback JWT authorization and D1 node/workspace liveness checks still run per request.
ACP_ACTIVITY_BINDING_CACHE_MAX_ENTRIES2048Maximum cached ACP activity bindings retained by one Worker isolate.
SESSION_SNAPSHOT_RECOVERY_CLAIM_LEASE_MS600000 (10 min)Time after which an interrupted replacement-runtime wake claim can be reconciled or reclaimed.
SESSION_LIFECYCLE_ERROR_MAX_LENGTH2048Maximum session lifecycle and agent activity failure diagnostic detail stored in lifecycle records.
SESSION_SNAPSHOT_PURGE_ENABLEDtrueEnables bounded expiry cleanup: terminalizes the sleeping chat, deletes its R2 objects, then removes D1 metadata. Chats slept by the fallback expire on the same seven-day schedule.
SESSION_SNAPSHOT_PURGE_BATCH_SIZE250Maximum expired snapshot rows deleted per daily purge.
SESSION_SLEEP_IN_FLIGHT_MAX_AGE_MS1800000 (30 minutes)Absolute ceiling for preserving in-flight sleep lifecycle rows (scheduled, preparing, stopping, retry-eligible failed) from terminal session destroyers. Rows older than this are treated as wedged and must escape through bounded repair by runSessionSleepLifecycleRepair() in apps/api/src/scheduled/session-sleep-lifecycle-repair.ts instead of deferring forever.
SESSION_SLEEP_IN_FLIGHT_REPAIR_BATCH_SIZE25Maximum stale post-capture in-flight sleep rows repaired per scheduled sweep by runSessionSleepLifecycleRepair() in apps/api/src/scheduled/session-sleep-lifecycle-repair.ts. The repair only completes restorable, unexpired preparing/stopping rows as sleeping; it does not wake or replay work. Values above 100 are capped.
TERMINAL_SESSION_RECONCILE_PROJECT_BATCH_SIZE25Maximum projects inspected for stale active ProjectData session ledgers per scheduled sweep. Values above 200 are capped.
TERMINAL_SESSION_RECONCILE_BATCH_SIZE25Maximum active ProjectData chat_sessions candidates reconciled per project per scheduled sweep. Values above 200 are capped.
TERMINAL_SESSION_SUMMARY_RECONCILE_BATCH_SIZE25Maximum active D1 session_summaries candidates reconciled globally per scheduled sweep. Values above 200 are capped.
TERMINAL_SESSION_RECONCILE_DEFER_MS3600000 (1 hour)Retry delay for live-head, snapshot-protected, or temporarily ineligible terminal-session ledger candidates. Values above 86400000 (24 hours) are capped.
TERMINAL_NODE_LIFECYCLE_REPAIR_BATCH_SIZE25Maximum active-looking workspace rows on terminal/deleted nodes repaired per scheduled sweep by runTerminalNodeLifecycleRepair() in apps/api/src/scheduled/terminal-node-lifecycle-repair.ts. The repair marks non-sleeping workspaces stopped, closes non-terminal agent sessions and open compute usage, and routes ProjectData cleanup through the sleeping-snapshot guard. Values above 100 are capped.
TERMINAL_NODE_LIFECYCLE_REPAIR_WALL_BUDGET_MS10000 (10 seconds)Wall-clock budget for runTerminalNodeLifecycleRepair() in apps/api/src/scheduled/terminal-node-lifecycle-repair.ts inside the scheduled sweep. Values above 30000 are capped so this repair cannot monopolize the cron event.
REQUIRE_APPROVAL(unset)Default signup approval gate. Superadmins can override it at runtime in Admin → Users without redeploying; when no runtime override exists, this value is used. The first genuine human becomes superadmin regardless of this flag — see First Login & Admin Access.
TRIAL_ANONYMOUS_USER_IDsystem_anonymous_trialsId of the internal anonymous-trial sentinel user, excluded from first-user superadmin checks. Override only if your deployment uses a different sentinel id.
CAPACITY_POOL_BACKFILL_SCOPE_BATCH_SIZE25Maximum user scopes and maximum project scopes reconciled by one unscoped capacity-pool backfill pass. Values above 200 are capped; rerun the backfill to continue.
CAPACITY_POOL_SCHEDULED_RECONCILIATION_INTERVAL_MS86400000 (24 hours)Minimum interval between scheduled capacity-pool reconciliation runs. Set to 0 to let every operational cron sweep reconcile; explicit UI/API reconcile requests are not throttled by this setting.
CAPACITY_POOL_LEGACY_WORKLOAD_MAPPING_JSONbuilt-in slicesEnvironment fallback for the versioned legacy small/medium/large to workload requirements adapter. Persisted platform_settings.capacityPools.legacyWorkloadMapping.v1 wins when present. Values are workload slices, not old whole-VM shapes.
CAPACITY_POOL_PLATFORM_DEFAULTS_JSONbuilt-in defaultsEnvironment fallback for platform resource requirement defaults used when no task/trigger/skill/profile/project/user layer sets a field. Persisted platform_settings.capacityPools.platformDefaults.v1 wins when present. Values are validated by the shared ResourceRequirements validator and must provide every field.
CAPACITY_POOL_SELECTION_SETTINGS_JSONbuilt-in scoring weightsEnvironment fallback for default capacity-pool selection weights and ranking rollout. Persisted platform_settings.capacityPools.selectionSettings.v1 wins when present. Candidate priority remains explicit pool policy; price comparisons are normalized by unit and currency, with unknown price sorted after known comparable prices.
ORIGIN_CA_CERT_VALIDITY_DAYS7Validity for per-node Cloudflare Origin CA certificates issued from node-generated CSRs. Must be one of Cloudflare’s supported values: 7, 30, 90, 365, 730, 1095, or 5475.

The rolloutCohortPercent field in capacity-pool selection settings accepts 0–100 (default 100). resolvePlacementRollout in services/placement-rollout.ts assigns stable user/pool cohorts. Enabled cohorts use the pool’s configured ranking; other cohorts use native balanced ranking while the reuse selector records which host the configured strategy would select and why the selections differ. The same eligible host set feeds both comparisons. Pool precedence, membership, credential generation, workload role, aggregate reservations, and paid allocation fences remain enforced at every percentage. Reducing rollout changes ranking; it never restores legacy size labels as allocation authority. Settings and plan columns remain additive, and readers accept plans without rollout diagnostics.

For the user-facing explanation of what these settings control — pool scopes and precedence, allowed offerings, strategies, exhaustion policies, and resource requirements — see the Compute Pools guide.

Deploy the normal additive migrations before starting the updated Worker. Existing tasks, workspaces, credentials, and recorded hardware remain in place. Background reconciliation creates missing default pools from existing credentials and resumes in bounded batches; opening the settings page is not required. Larger installations may need several scheduled passes before every scope is ready.

In project Infrastructure settings, inspect the effective default pool before starting new work. A project default takes precedence over a personal default, which takes precedence over installation capacity. These states need different responses:

Pool stateWhat to do
Migration pendingAllow reconciliation to finish; if it persists, check scheduled reconciliation errors and the affected credential.
Configured emptySelect a supported offering in that pool. An empty configured pool intentionally blocks new allocation.
Source disabledRe-enable or replace the pool’s credential source.
Catalog unavailableCheck provider access and retry after inventory refresh. A failed refresh preserves the last valid inventory.
Configured readyStart a small test workload and check its requested resources and provider-native hardware in the node details.

An administrator can verify completion in D1 by inspecting capacity_pools: migration_state must be complete for the affected pool. The durable user and project backfill cursors are stored in platform_settings under capacityPools.backfill.userCursor.v1 and capacityPools.backfill.projectCursor.v1. Their presence indicates resumable progress, not an error. Do not delete pools or reset cursors to resolve an unavailable credential.

Existing nodes without verified pool and provider identity may finish their current work but are not automatically treated as eligible pool capacity. New work must pass the current pool, credential, and resource checks. Previously recorded hardware remains visible even if an offering is later removed. Old browser, API, CLI, and MCP size fields remain accepted as compatibility inputs; new resource fields take precedence at the same configuration layer. Saved reservations survive retry rather than adopting changed defaults.

For a ranking rollback, reduce rolloutCohortPercent in the effective selection settings. This uses balanced ranking for the excluded cohort while retaining pool authorization and capacity checks. It does not roll back migrations, revive removed offerings, or permit reuse of unverified nodes. Keep the additive schema and saved plans; do not drop columns or recreate tables as a rollback step.

Activity coalescing and binding caches are per Worker isolate, so burst reduction scales with the number of active isolates for the same session. Delayed flushes carry their original observed event time, and ProjectData rejects stale writes so a delayed intermediate report cannot overwrite a newer idle/error state from another isolate.

VariableDefaultDescription
LIBRARY_PROJECT_DELETE_CLEANUP_BATCH_SIZE1000Maximum project-owned library objects listed and deleted per R2 page after project deletion. Values above R2’s 1,000-object page maximum are capped.

Deployment release and compose artifact retention

Section titled “Deployment release and compose artifact retention”

The scheduled Worker first reconciles provably stale non-terminal compose releases, then prunes terminal deployment releases outside the protected window (apps/api/src/scheduled/d1-retention.ts:runDeploymentReleaseRetention()). Terminal retention always retains the newest releases per environment and the version reported in deployment_environments.observed_applied_seq. The stale reconciler only marks a created/applying compose-artifact release failed when D1 shows old release status activity, stable authenticated deployment-node observed state, no recent release fetch/apply events, a valid manifest, and a release version that is not the observed applied version. Unknown statuses, malformed manifests, missing observed state, active applying observations, and recent release events fail closed. Compose artifact cleanup then re-derives references from the remaining manifests (apps/api/src/scheduled/compose-image-artifact-cleanup.ts:runComposeImageArtifactCleanup()).

VariableDefaultDescription
DEPLOYMENT_RELEASE_RETENTION_ENABLEDtrueEnables bounded terminal release pruning.
DEPLOYMENT_RELEASE_RETENTION_COUNT3Newest releases protected per environment, in addition to observed-applied and non-terminal releases.
DEPLOYMENT_RELEASE_RETENTION_BATCH_SIZE250Maximum release rows deleted per run.
DEPLOYMENT_RELEASE_RETENTION_INTERVAL_HOURS24Minimum interval between release retention runs.
DEPLOYMENT_RELEASE_RETENTION_LAST_RUN_KV_KEYcleanup:deployment-releases:last-runKV interval marker.
DEPLOYMENT_RELEASE_RECONCILIATION_ENABLEDtrueEnables stale non-terminal compose release reconciliation before terminal retention.
DEPLOYMENT_RELEASE_RECONCILIATION_BATCH_SIZE50Maximum stale non-terminal releases marked failed per retention run.
DEPLOYMENT_RELEASE_RECONCILIATION_STALE_HOURS168Minimum release status age before reconciliation can terminalize a stale non-terminal release.
DEPLOYMENT_RELEASE_RECONCILIATION_ACTIVITY_GRACE_HOURS6Recent release-event window that protects active fetch/apply work from reconciliation.
COMPOSE_IMAGE_ARTIFACT_CLEANUP_BATCH_SIZE250Maximum abandoned compose archives deleted per daily run.

Pulumi updates the existing assets bucket lifecycle resource on upgrades and creates the same rules on clean installs (infra/resources/storage.ts:r2BucketLifecycle). temp-uploads/ is transient browser-upload staging; tts/ is a regenerable audio cache. Durable library/ content is deleted only with its project, and reachable compose-image-artifacts/ are governed by deployment release retention. Archived ProjectData tool payloads stay private and retrievable through Worker/MCP access paths while message rows retain their text in the Durable Object, so these durable prefixes do not have age-only lifecycle rules.

Pulumi optionDefaultObject prefixDescription
sessionSnapshotTtlDays7session-snapshots/Worker-owned retention from actual sleep; no age-only R2 lifecycle
diagnosticIncidentTtlDays7configured privatePrivate diagnostic artifact retention
tempUploadTtlDays1temp-uploads/Abandoned presigned browser upload retention
ttsTtlDays30tts/Regenerable TTS audio-cache retention
n/an/aproject-data/tool-payloads/Private ProjectData archive; Worker-owned retention only
n/an/aresource-history/Private workspace resource chunks; Worker-owned retention only

All TTL options must be positive integers. Set overrides with pulumi config set against the target stack before running its deployment workflow.

Google login and Google infrastructure authorization are independent credential families:

VariablesPurposeRuntime precedenceRedirect URIs
GOOGLE_LOGIN_CLIENT_ID, GOOGLE_LOGIN_CLIENT_SECRETBetterAuth user login/setup or superadmin runtime D1 → Worker env → unset/api/auth/callback/google
GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRETKeyless GCP/WIF setup onlySuperadmin runtime D1 → Worker env → unset/auth/google/callback and /api/deployment/gcp/callback

Configuring one family never enables or modifies the other. Users who choose service-account JSON do not need either infrastructure OAuth variable.

VariableDefaultDescription
GCP_SERVICE_ACCOUNT_JSON_MAX_BYTES65536Maximum UTF-8 byte size accepted by PUT /api/gcp/service-account
GCP_DEFAULT_ZONEus-central1-aDefault Compute zone
GCP_IMAGE_FAMILYubuntu-2404-lts-amd64Compute image family. Native image overrides may be a family name or a Compute Engine image/family reference.
GCP_IMAGE_PROJECTubuntu-os-cloudCompute image project
GCP_DISK_SIZE_GB50Default boot disk size for GCP legacy callers and native requests without bootDiskSizeGb. A native VM request with bootDiskSizeGb overrides this value before the Compute Engine insert call.
GCP_TOKEN_CACHE_TTL_SECONDS3300Maximum derivative access-token cache TTL; actual TTL is capped by Google’s returned expiry
GCP_IDENTITY_TOKEN_EXPIRY_SECONDS600SAM identity-token lifetime for WIF
GCP_OPERATION_POLL_TIMEOUT_MS300000Maximum wait for GCP asynchronous operations
GCP_API_TIMEOUT_MS30000GCP OAuth, IAM, and Compute request timeout
GCP_STS_SCOPEhttps://www.googleapis.com/auth/cloud-platformWIF STS exchange scope
GCP_SA_IMPERSONATION_SCOPEShttps://www.googleapis.com/auth/computeComma-separated scopes for WIF service-account impersonation
GCP_SA_TOKEN_LIFETIME_SECONDS3600WIF impersonated access-token lifetime
GCP_STS_TOKEN_URLhttps://sts.googleapis.com/v1/tokenWIF STS endpoint override for controlled environments
GCP_IAM_CREDENTIALS_BASE_URLGoogle IAM Credentials APIWIF impersonation base URL override

The service-account JWT bearer flow always uses https://oauth2.googleapis.com/token; it has no endpoint override, and uploaded token_uri values are ignored. Source credentials are encrypted in D1. Only derivative short-lived tokens are cached.

VariableDefaultDescription
TASK_TITLE_MODEL@cf/google/gemma-4-26b-a4b-itWorkers AI model for title generation
TASK_TITLE_MAX_LENGTH100Max characters in generated title
TASK_TITLE_TIMEOUT_MS5000Timeout before falling back to truncation
TASK_TITLE_GENERATION_ENABLEDtrueSet false to disable AI generation
TASK_TITLE_SHORT_MESSAGE_THRESHOLD100Messages at or below this length bypass AI
TASK_TITLE_MAX_RETRIES2Max retry attempts on failure
TASK_TITLE_RETRY_DELAY_MS1000Base delay between retries (exponential backoff)
TASK_TITLE_RETRY_MAX_DELAY_MS4000Max delay cap for backoff
TASK_TITLE_ERROR_DIAGNOSTIC_MAX_LENGTH512Max sanitized provider-error diagnostic length
VariableDefaultDescription
BRANCH_NAME_PREFIXsam/Prefix for generated task output branches. Include the trailing separator (for example agent/).

Task workspaces are checked out on the generated output branch, and SAM refuses to auto-push a completed task while the workspace is still on the project’s default branch. See Where the work lands.

VariableDefaultDescription
DEBUG_AGENT_MODEL@cf/zai-org/glm-5.2Workers AI model for superadmin deployment diagnosis
DEBUG_AGENT_MAX_TURNS6Maximum model/tool turns per diagnosis
DEBUG_AGENT_RUN_TOKEN_LIMIT96000Combined token ceiling per diagnosis
DEBUG_AGENT_MODEL_OUTPUT_TOKENS4096Maximum output tokens requested per model turn
DEBUG_AGENT_DAILY_TOKEN_LIMIT480000Daily diagnosis token budget, counted per feature
DEBUG_AGENT_TOOL_RESULT_LIMIT50Maximum rows returned by a diagnosis tool
DEBUG_AGENT_TOOL_RESULT_BYTES32768Maximum serialized bytes per model-visible tool result
DEBUG_AGENT_MAX_WINDOW_HOURS24Maximum selectable diagnosis window
DEBUG_AGENT_TIMEOUT_MS120000Timeout for each diagnosis model request
DEBUG_AGENT_HARD_DEADLINE_MS900000Hard deadline for an active diagnosis
DEBUG_AGENT_STALE_HEARTBEAT_MS120000Orphan reconciler heartbeat threshold
DEBUG_AGENT_RETRY_BASE_DELAY_MS2000Initial transient step retry delay
DEBUG_AGENT_RETRY_MAX_DELAY_MS60000Maximum transient step retry delay
DEBUG_AGENT_STEP_MAX_RETRIES3Maximum classified transient retries per step

The /admin/errors view remains superadmin-only and may show local user IDs, IP addresses, and user-agent strings. Before any tool result enters model context, SAM recursively removes those fields plus credential-shaped values such as API tokens, JWTs, authorization headers, private keys, and long secret-like strings. Cloudflare credentials stay server-side and are never included in model messages or saved diagnosis text.

VM failures use a durable local SQLite outbox and a private R2 artifact. Generated deployments set the R2 prefix and object lifecycle from Pulumi; the remaining Worker bounds can be overridden through deployment environment variables.

Worker variableDefaultDescription
MAX_VM_AGENT_ERROR_BODY_BYTES32768Maximum VM error batch body
MAX_VM_AGENT_ERROR_BATCH_SIZE10Maximum errors per VM batch
MAX_VM_AGENT_ERROR_SOURCE_LENGTH256Maximum redacted VM error source length
OBSERVABILITY_ERROR_MESSAGE_MAX_LENGTH2048Maximum persisted observability error message length
OBSERVABILITY_ERROR_STACK_MAX_LENGTH4096Maximum persisted observability stack length
OBSERVABILITY_ERROR_USER_AGENT_MAX_LENGTH512Maximum persisted observability user-agent length
VM_INCIDENT_R2_PREFIXdiagnostic-incidentsPrivate object prefix; generated from the Pulumi output
VM_INCIDENT_ARTIFACT_MAX_BYTES2097152Maximum compressed artifact size
VM_INCIDENT_REGISTRATION_MAX_BYTES262144Maximum registration JSON body
VM_INCIDENT_MANIFEST_MAX_BYTES131072Maximum redacted manifest
VM_INCIDENT_PREVIEW_MAX_BYTES131072Maximum redacted model/UI preview
VM_INCIDENT_MAX_ARTIFACTS_PER_NODE50Active artifact quota per node
VM_INCIDENT_MAX_BYTES_PER_NODE104857600Active expected-byte quota per node
VM_INCIDENT_RETENTION_DAYS7Private object and active metadata retention
VM_INCIDENT_METADATA_RETENTION_DAYS30Expired metadata retention after object deletion
VM_INCIDENT_PENDING_TIMEOUT_MINUTES30Incomplete-upload timeout and upload-lease duration
VM_INCIDENT_RECONCILE_BATCH_SIZE50Maximum artifacts/incidents repaired per scheduled pass (minimum: 6)

The VM Agent process accepts the corresponding ERROR_REPORT_* overrides for flush interval, batch size/bytes, outbox size and path, SQLite busy timeout, HTTP timeout, retry bounds, attempts, spool path/bytes, artifact bytes, retention, collector timeout/count/concurrency, document bytes, recursive value depth/items, string bytes, structured event limit, response-read bytes, and persisted-error bytes. Generated deployments pass these validated values through cloud-init into the VM Agent systemd service, so overrides apply to newly provisioned nodes. Defaults are listed in apps/api/.env.example; the common defaults are a 32 KiB error batch, 1,000-row outbox, 2 MiB artifact, 20 MiB spool, and 24-hour local retention.

Pulumi options diagnosticIncidentPrefix (default diagnostic-incidents) and diagnosticIncidentTtlDays (default 7, any positive integer) configure the private prefix and an independent R2 lifecycle rule. They do not require a separate bucket or manually managed Worker variable. The prefix cannot begin with the application-owned namespaces agents, cli, compose-image-artifacts, library, resource-history, session-snapshots, temp-uploads, or tts, because the lifecycle would otherwise expire unrelated objects.

VariableDefaultDescription
PLATFORM_FEEDBACK_PROJECT_IDunsetBootstrap/environment fallback for the project that receives user issue reports and automated triage draft Ideas. The Admin → Integrations runtime setting is preferred and overrides it.
PLATFORM_FEEDBACK_TRIAGE_WINDOW_MINUTES60Lookback window for grouping recent platform errors
PLATFORM_FEEDBACK_TRIAGE_ERROR_LIMIT100Maximum platform error rows scanned per triage sweep
PLATFORM_FEEDBACK_TRIAGE_GROUP_LIMIT5Maximum grouped feedback candidates processed per triage sweep
PLATFORM_FEEDBACK_TRIAGE_EVIDENCE_LIMIT10Maximum bounded error references retained per grouped feedback record
PLATFORM_FEEDBACK_TRIAGE_CLAIM_TTL_MS600000Claim lease duration before a later sweep can reclaim the group
PLATFORM_FEEDBACK_TRIAGE_MAX_FAILURES3Maximum failed attempts before a group is rejected from auto-triage
PLATFORM_FEEDBACK_TRIAGE_FAILURE_REASON_MAX_LENGTH240Maximum characters stored or returned for sanitized failure reasons
PLATFORM_FEEDBACK_TRIAGE_BUDGET_DEFER_MS86400000Retry delay for per-run budget deferrals
PLATFORM_FEEDBACK_INCIDENT_DISPATCH_LEASE_TTL_MS7200000Dispatch lease before a failed incident trigger handoff can be reclaimed
PLATFORM_FEEDBACK_INCIDENT_AGENT_LEASE_TTL_MS3600000Agent claim lease before another task can reclaim a private incident
PLATFORM_FEEDBACK_INCIDENT_MAX_DISPATCH_ATTEMPTS3Agent-reported failed dispatch attempts before an incident is rejected
PLATFORM_FEEDBACK_INCIDENT_REOPEN_COOLDOWN_MS1800000Minimum elapsed time after terminal resolution/expiry before a newer occurrence can reopen the same signature; set 0 to disable cooldown-only suppression
PLATFORM_FEEDBACK_INCIDENT_RECLAIM_LIMIT25Maximum expired dispatch leases reclaimed by one incident sweep
PLATFORM_FEEDBACK_INCIDENT_MAX_AGE_MS2592000000Maximum active incident age before expiry
PLATFORM_FEEDBACK_INCIDENT_STALE_SINGLETON_MAX_AGE_MS259200000Maximum age for one-off pending incidents with no recurrence
PLATFORM_FEEDBACK_INCIDENT_STALE_SINGLETON_EXPIRY_BATCH_SIZE25Maximum stale singleton incidents expired per sweep
PLATFORM_FEEDBACK_INCIDENT_MIN_DISPATCH_SEVERITYerrorMinimum severity admitted to automatic VM incident dispatch
PLATFORM_FEEDBACK_INCIDENT_MIN_DISPATCH_BATCH_SIZE2Dispatch immediately once this many eligible incidents are ready
PLATFORM_FEEDBACK_INCIDENT_MIN_PENDING_AGE_MS1800000Dispatch a smaller eligible batch after this pending age
PLATFORM_FEEDBACK_INCIDENT_DISPATCH_RATE_WINDOW_MS3600000Rate-cap window for each incident trigger
PLATFORM_FEEDBACK_INCIDENT_MAX_DISPATCHES_PER_TRIGGER_WINDOW1Maximum dispatches one incident trigger may submit per rate window
PLATFORM_FEEDBACK_INCIDENT_AUTO_TRIGGER_ENABLEDtrueAuto-create one private incident trigger when pending incidents exist and no incident trigger exists
PLATFORM_FEEDBACK_INCIDENT_TRIGGER_LIMIT5Maximum active incident triggers inspected per sweep
PLATFORM_FEEDBACK_INCIDENT_TRIGGER_NAMEbuilt-inName for the auto-created private incident trigger
PLATFORM_FEEDBACK_INCIDENT_TRIGGER_TEMPLATEbuilt-inPrompt template for the auto-created private incident trigger
PLATFORM_FEEDBACK_INCIDENT_SUMMARY_LIMIT10Maximum grouped incidents included in one incident-trigger backlog summary
PLATFORM_FEEDBACK_INCIDENT_EVIDENCE_REF_LIMIT10Maximum bounded evidence references retained per incident
PLATFORM_FEEDBACK_INCIDENT_EVIDENCE_MAX_BYTES32768Maximum serialized evidence bytes retained per incident
PLATFORM_FEEDBACK_INCIDENT_RESOLUTION_NOTE_MAX_LENGTH2000Maximum private incident resolution-note length

Resolved and expired incident signatures reopen only when a newer occurrence arrives after PLATFORM_FEEDBACK_INCIDENT_REOPEN_COOLDOWN_MS; older lookback-window occurrences remain closed. VM-agent incidents resolved with fix evidence also wait for occurrences from nodes reporting the current VM_AGENT_REQUIRED_VERSION, because Worker deploys do not update already-running VM binaries. Dispatch attempts are consumed only when the incident task reports its own failure; platform-side handoff/session failures release the dispatch without incrementing dispatch_attempts.

Automated triage and superadmin-initiated diagnosis read the same DEBUG_AGENT_DAILY_TOKEN_LIMIT value but count against independent per-feature counters, so worst-case daily spend across both is twice this value. Automated triage treats budget exhaustion as a retryable deferral: daily exhaustion retries after the next UTC day starts, and per-run exhaustion uses PLATFORM_FEEDBACK_TRIAGE_BUDGET_DEFER_MS. Incident trigger agents run from the private grouped backlog and dispatch one agent for a bounded backlog summary, not one agent per occurrence. Automatic incident dispatch ignores pending signatures already linked to open tracked work and warning-only signatures below the configured severity floor, then applies the batch/age gate and per-trigger rate cap before reserving incidents.

The in-app Report an Issue flow files user-submitted reports as draft Ideas in the effective private feedback project. Configure it from Admin → Integrations when possible; PLATFORM_FEEDBACK_PROJECT_ID remains the environment fallback when no runtime setting is saved. The feature is hidden entirely — both UI entry points disappear and GET /api/report-issue/config returns enabled: false — when no effective project exists or the effective project does not exist in this deployment’s database.

VariableDefaultDescription
REPORT_ISSUE_TITLE_MAX_LENGTH200Truncation ceiling for the stored title (lowers only — see below)
REPORT_ISSUE_DESCRIPTION_MAX_LENGTH5000Truncation ceiling for the stored description (lowers only)
REPORT_ISSUE_CONTENT_MAX_LENGTH65536Maximum stored Idea body, including attached technical references
RATE_LIMIT_REPORT_ISSUE_POST20Report submissions allowed per clock hour, per authenticated user

The two length variables apply after request validation, so they can only lower the stored length — the request schema and the report dialog both enforce the built-in 200 / 5,000 caps regardless of what you set here.

See Reporting Issues for the user-facing flow and the untrusted-evidence Idea format.

SAM loads OpenCode Zen and OpenCode Go model choices through the authenticated model-catalog API, backed by Models.dev and cached in KV. If the upstream catalog or cache is unavailable, SAM falls back to the static catalog shipped with the app.

VariableDefaultDescription
MODEL_CATALOG_SOURCE_URLhttps://models.dev/api.jsonSource URL for the dynamic model catalog
MODEL_CATALOG_CACHE_TTL_SECONDS3600KV cache TTL for normalized dynamic model catalog payloads
MODEL_CATALOG_FETCH_TIMEOUT_MS5000Timeout for the upstream catalog fetch before static fallback

The dashboard’s Active Tasks list (GET /api/dashboard/active-tasks, apps/api/src/routes/dashboard.ts) reads up to DASHBOARD_ACTIVE_TASK_CANDIDATE_LIMIT of the user’s most recently started active tasks, ranks them by newest message (or by when they started, if they have none yet), and only then applies the display limit. The open dashboard refreshes the list every 15 seconds while its tab is visible.

VariableDefaultDescription
DASHBOARD_ACTIVE_TASK_LIMIT6Most recently active tasks the list shows
DASHBOARD_ACTIVE_TASK_CANDIDATE_LIMIT100Active tasks read and ranked before the display limit applies (also the maximum; each project’s candidates share one SQL statement’s binds)
DASHBOARD_INACTIVE_THRESHOLD_MS900000 (15 min)A working task whose last message is newer than this shows Active; older shows Working

Conservative Cache-Control budgets for stable and semi-stable API GETs, letting the browser serve a cached body instantly while it revalidates in the background. All values are seconds and are clamped to [0, 86400]; an unparseable or negative value falls back to the default rather than caching for longer.

Authenticated responses are always emitted as private with Vary: Cookie, so neither a shared cache nor a second account in the same browser can be served another user’s body. Only the unauthenticated /api/config/* endpoints are marked public. Endpoints returning real-time data (chat messages, task status, session and workspace state) are deliberately excluded.

VariableDefaultDescription
PUBLIC_CONFIG_CACHE_MAX_AGE_SECONDS60max-age for the unauthenticated /api/config/* endpoints
PUBLIC_CONFIG_CACHE_SWR_SECONDS300stale-while-revalidate for /api/config/*
MODEL_CATALOG_CACHE_MAX_AGE_SECONDS60max-age for GET /api/model-catalog/:agentType
MODEL_CATALOG_CACHE_SWR_SECONDS300stale-while-revalidate for the model catalog response
PROJECT_REFERENCE_CACHE_MAX_AGE_SECONDS0max-age for project agent-profile and skill lists (0 = always revalidate)
PROJECT_REFERENCE_CACHE_SWR_SECONDS30stale-while-revalidate for project agent-profile and skill lists

Direct node allocation and workspace creation use isolated NodeLifecycle Durable Object instances. These optional Worker runtime overrides accept positive integers; invalid or unset values use the defaults.

VariableDefaultDescription
NODE_PROVISIONING_REQUEST_TIMEOUT_MS5000 (5 s)Allocation/reconciliation request budget, also used for background readiness and workspace dispatch.
NODE_PROVISIONING_RETRY_INTERVAL_MS30000 (30 s)Delay between durable provisioning attempts.
NODE_PROVISIONING_MAX_AGE_MS900000 (15 min)Maximum provisioning intent age before retries stop.
NODE_PROVISIONING_MAX_ATTEMPTS30Maximum provisioning attempts before retries stop.

Reaching either the age or attempt limit stops allocation retries. Diagnostic publication then uses a separate retry budget with the same interval and attempt limit; the age limit applies only to allocation and reconciliation. The unresolved intent is retained for inspection. Empty provider inventory does not authorize another provider create request or establish cleanup proof.

VariableDefaultDescription
NODE_WARM_TIMEOUT_MS1800000 (30 min)Time a managed auto-provisioned node stays warm after its last active workspace leaves
MAX_AUTO_NODE_LIFETIME_MS14400000 (4 hr)Max lifetime for an auto-provisioned node holding no active workspaces
NODE_WARM_GRACE_PERIOD_MS2100000 (35 min)Cron sweep grace period (must be > warm timeout)
NODE_LIFECYCLE_ALARM_RETRY_MS60000 (1 min)Retry delay for DO alarm failures
NODE_LIFECYCLE_MAX_DESTROYING_AGE_MS86400000 (24 hr)Backstop after which a destroying-state alarm self-cleans; infrastructure teardown remains owned by cron/provider reconciliation
DEFAULT_TASK_AGENT_TYPEopencodeDefault agent for autonomous idea execution

The cleanup sweep measures idleness from a node’s last workspace activity (COALESCE(MAX(workspaces.updated_at), nodes.created_at)), never from nodes.updated_at — heartbeats rewrite updated_at on every beat, so it tracks liveness rather than idleness. The eligibility check is implemented by claimNodeForCleanup() in apps/api/src/scheduled/node-cleanup/shared.ts.

Reaping only ever applies to nodes with node_role = 'workspace' and node_class != 'user-owned'. Deployment nodes host long-running user applications and legitimately hold zero workspaces forever, so they are never reaped by these timers; they are released when their last deployment environment is deleted.

Stopped managed VM nodes created directly through a canonical pool can also be reaped using their server-recorded pool, credential, and native-offering identity. They retain the same workspace-activity and active-claim guards. A short project warm timeout does not bypass the workspace idle window. Runtime teardown preserves the saved snapshot and conversation for recovery on a fresh node.

VariableDefaultDescription
NODE_WORKSPACE_IDLE_TIMEOUT_MS1800000 (30 min)Last-workspace-activity window before an auto-provisioned node_role = 'workspace' node with no active workspaces can be destroyed. Uses COALESCE(MAX(workspaces.updated_at), nodes.created_at), never heartbeat-updated nodes.updated_at.
NODE_ORPHAN_IDLE_TIMEOUT_MSlegacy aliasBackward-compatible alias used only when NODE_WORKSPACE_IDLE_TIMEOUT_MS is unset.
NODE_ABSOLUTE_MAX_LIFETIME_MS86400000 (24 hr)Hard ceiling on auto-provisioned workspace node age. Applies even when a workspace row still reports running, provided no workspace has reported activity within the idle window — this is what stops a stuck workspace row from making a node immortal.
NODE_CLEANUP_SWEEP_LIMIT25Max node candidates processed per cleanup phase per cron run.
NODE_CLEANUP_FAILURE_BACKOFF_MS3600000 (1 hr)Expiring exclusion applied to failed cleanup candidates so a permanent provider error cannot monopolize the bounded page.
NODE_UNHEALTHY_DRAIN_AFTER_MS600000 (10 min)Heartbeat-loss window before SAM posts a session notice and requests sleep on a managed workspace VM.
NODE_UNHEALTHY_RELEASE_AFTER_MS1800000 (30 min)Heartbeat-loss window before SAM attempts strict provider deletion, even if the node still has active workspace rows. The value is clamped above the drain threshold.
NODE_UNHEALTHY_FLEET_MAX_FRACTION0.5Holds destructive cleanup when this fraction of at least three managed workspace VMs lose heartbeat together, indicating a possible heartbeat-intake incident. The hold escalates after the configured release and drain windows combined; it does not delete busy nodes on a timer.
NODE_UNHEALTHY_FLEET_MIN_NODES3Minimum number of managed workspace VMs required before the fleet-wide heartbeat-loss guard applies.
NODE_UNHEALTHY_PRESERVATION_TIMEOUT_MS5000 (5 sec)Per-node budget for attempting chat notices and sleep requests before cleanup proceeds.
NODE_UNHEALTHY_RETRY_MS60000 (1 min)Retry delay after provider deletion of an unhealthy node fails.
NODE_STOPPED_HANDOFF_SWEEP_BUDGET_MS20000 (20 sec)Wall-clock budget for stopped-node handoff. Candidates not started within the budget remain eligible for the next sweep.
NODE_STOPPED_HANDOFF_REQUEST_TIMEOUT_MS5000 (5 sec)Per-candidate provider/DNS deadline during stopped-node handoff, capped by remaining sweep time. Provider failures enter cleanup backoff.
WORKSPACE_CLEANUP_SWEEP_LIMIT50Max workspace candidates processed per cleanup phase per cron run.
NODE_AGENT_BACKGROUND_REQUEST_TIMEOUT_MS5000 (5 s)VM-agent request timeout for background sweeps. Deliberately far below the interactive NODE_AGENT_REQUEST_TIMEOUT_MS (30 s) so a sweep over unreachable nodes cannot exhaust the Worker’s wall-clock budget.
WORKSPACE_DELETION_RETRY_BASE_MS60000 (1 min)Initial retry delay after a VM workspace deletion remains unconfirmed.
WORKSPACE_DELETION_RETRY_MAX_MS3600000 (1 hr)Maximum exponential backoff for an unconfirmed workspace deletion.
WORKSPACE_DELETION_MAX_RESIDENCE_MS86400000 (24 hr)Maximum hot-retry residence before an unconfirmed deletion enters durable operator quarantine; its workspace remains stopping and replacement-fenced.
WORKSPACE_DELETION_ALARM_BATCH_SIZE3Maximum due workspace-deletion entries processed by one NodeLifecycle alarm, sized to stay within the Cloudflare Free-plan D1 query budget on the worst successful linked-workspace path.
WORKSPACE_DELETION_CALLBACK_SIGNAL_TTL_SECONDS300 (5 min)Per-workspace and callback-kind throttle for payload-free workspace.deletion_unconfirmed_callback activity evidence.
WORKSPACE_DELETION_CALLBACK_SIGNAL_CLEANUP_LIMIT25Maximum expired callback telemetry throttle claims removed by one signal attempt.
WORKSPACE_DELETION_DIAGNOSTIC_MAX_LENGTH500 charactersMaximum sanitized deletion-attempt diagnostic stored in workspaces.error_message.

The cron and Durable Object switches are availability brakes: an absent key or KV read error means enabled (fail-open). This differs deliberately from the fail-closed trials entitlement switch. Superadmins can inspect and update both brakes through /api/admin/runtime-controls; emergency operators can use the KV procedure in .claude/rules/55-runaway-cost-emergency-ops.md.

VariableDefaultDescription
CRON_SWEEPS_ENABLED_KV_KEYcontrol-loops:cron-enabledKV key gating the five-minute operational sweep block
DO_ALARMS_ENABLED_KV_KEYcontrol-loops:alarms-enabledShared KV key gating alarm-bearing Durable Objects
CONTROL_LOOP_KILL_SWITCH_CACHE_MS30000In-memory switch cache; runtime clamps it to at most 30 seconds
CONTROL_LOOP_DISABLED_ALARM_RETRY_MS300000 (5 min)Safe alarm recheck interval while DO work is disabled; values below 60 seconds are clamped
CRON_FAILURE_NOTIFICATION_THROTTLE_MS3600000 (1 hr)Per-sweep throttle enforced by a KV cache plus an atomic per-user Notification DO claim; also how long an archive-breaker alert claim is held
CRON_FAILURE_NOTIFICATION_KV_PREFIXcron-failure-notificationKV prefix for notification throttle markers
DIAGNOSIS_COMPLETED_STEP_MIN_DELAY_MS1000Minimum delayed re-arm for an already-completed diagnosis step
ORCHESTRATOR_ZERO_TASK_GRACE_MS600000 (10 min)Grace period before an active mission with no tasks terminalizes
ORCHESTRATOR_MAX_MISSION_LIFETIME_MS86400000 (24 hr)Backstop that force-completes active/completing missions

The scheduled Durable Object billing monitor reads these non-secret variables from the selected GitHub Environment, not from the API Worker runtime:

VariableDefault/fallbackDescription
DO_WALL_TIME_SCRIPT_NAMESnoneOptional comma-separated API Worker filter for wall-time and invocation-rate analysis
DO_INVOCATION_RATE_REGRESSION_RATIO2Recent-versus-seven-day-baseline request-rate failure ratio
DO_CRON_LIVENESS_MAX_AGE_HOURS3Maximum age of the most recent targeted cron.completed event
DO_CRON_LIVENESS_SCRIPT_NAMESDO_WALL_TIME_SCRIPT_NAMESExplicit API Worker service target for cron liveness; the GitHub workflow derives both from RESOURCE_PREFIX and the selected stack when unset
DO_CRON_LIVENESS_ENDPOINTCloudflare Workers Observability query endpointOptional endpoint override for compatible/private telemetry gateways

The selected GitHub Environment’s CF_API_TOKEN secret must include the Cloudflare Workers Observability Write permission. Cloudflare requires that permission for the telemetry query endpoint even though this monitor only reads aggregated liveness telemetry.

Reclaims cloud servers that exist at the provider but which no live database row claims — for example when a server was created but the control plane failed before recording its instance ID.

Because this is the only path that destroys infrastructure on the basis of absent evidence, it fails closed at every step. A server must carry both the current control-plane env value and the exact Pulumi-generated installation marker before SAM consults D1. SAM then re-reads and revalidates the same provider resource immediately before it calls the provider delete API. Provider-account membership, server names, resource prefixes, and absence from this installation’s D1 are not ownership proof.

Pulumi generates the non-secret installation identity automatically on first deploy, persists it in the stack state, and injects it into the Worker as SAM_INSTALLATION_ID; there is no manual GitHub Environment setting. An upgrade does not relabel existing servers. Legacy servers without the marker remain usable and are preserved indefinitely, while servers provisioned after the upgrade participate in normal orphan cleanup. If the Pulumi state is lost or recreated, the new identity safely leaves the old fleet unattributable instead of adopting it destructively. Any missing/malformed identity, ambiguous provider metadata, or failed/malformed D1 lookup skips deletion. Resources surfaced to reconciliation with non-owning metadata emit aggregate operator-visible counters.

VariableDefaultDescription
PROVIDER_ORPHAN_RECONCILIATION_ENABLEDtrueSet to false to disable provider-side reconciliation entirely.
PROVIDER_ORPHAN_MIN_AGE_MS3600000 (1 hr)Minimum server age before it can be treated as an orphan. Must comfortably exceed provisioning time, since a server’s instance ID is recorded only after the provider returns it.
PROVIDER_ORPHAN_DESTROY_LIMIT5Max servers destroyed per reconciliation run.
PROVIDER_ORPHAN_RECONCILE_INTERVAL_MS3600000 (1 hr)Minimum interval between runs. Invoked by the 5-minute cron but self-throttled to this interval via KV.
VariableDefaultDescription
PROJECT_INVITE_TOKEN_BYTES32Random bytes used for generated project invite link tokens
PROJECT_INVITE_DEFAULT_EXPIRY_DAYS7Default lifetime for invite links created without an explicit expiry
PROJECT_INVITE_MAX_EXPIRY_DAYS30Maximum allowed invite link lifetime, including explicit expiry-date input
PROJECT_OFFBOARDING_PLAN_TTL_SECONDS900Lifetime for project member offboarding preview plans before recomputation
VariableDefaultDescription
NOTIFICATION_PROGRESS_BATCH_WINDOW_MS300000 (5 min)Min interval between progress notifications per idea
NOTIFICATION_DEDUP_WINDOW_MS60000 (60s)Dedup window for task_complete notifications
NOTIFICATION_AUTO_DELETE_AGE_MS7776000000 (90 days)Auto-delete old notifications
MAX_NOTIFICATIONS_PER_USER500Max stored notifications per user
NOTIFICATION_PAGE_SIZE50Default page size for notification list
MAX_NOTIFICATION_PAGE_SIZE100Max allowed page size
HUMAN_INPUT_TIMEOUT_MS7200000 (2 hr)Initial needs-input response window
HUMAN_INPUT_ESCALATION_FRACTIONS0.25,0.75Reminder points within the initial response window
HUMAN_INPUT_UNDELIVERED_GRACE_MS7200000 (2 hr)Extension without confirmed push delivery
HUMAN_INPUT_MAX_WAIT_MS86400000 (24 hr)Hard maximum needs-input marker lifetime
WEB_PUSH_TTL_SECONDS86400Push-service message TTL
WEB_PUSH_VAPID_TTL_SECONDS43200VAPID authorization-token lifetime
WEB_PUSH_DELIVERY_TIMEOUT_MS10000Per-attempt push-service timeout
WEB_PUSH_DELIVERY_BUDGET_MS25000Total fan-out budget, hard-capped at 25s below Worker background limit
WEB_PUSH_FANOUT_CONCURRENCY8Maximum concurrent endpoint deliveries
WEB_PUSH_MAX_ATTEMPTS3Bounded transient delivery attempts
WEB_PUSH_MAX_RETRY_AFTER_SECONDS30Maximum honored Retry-After delay
WEB_PUSH_MAX_PAYLOAD_BYTES3500Maximum unencrypted payload size
WEB_PUSH_FAILURE_THRESHOLD5Consecutive failures before disabling a subscription
WEB_PUSH_MAX_SUBSCRIPTIONS_PER_USER8Maximum retained browser endpoints per user
WEB_PUSH_USER_AGENT_MAX_LENGTH512Maximum stored browser description length
RATE_LIMIT_PUSH_SUBSCRIPTION30Subscription mutations per user per hour
VariableDefaultDescription
TRIGGER_STALE_EXECUTION_TIMEOUT_MS1800000 (30 min)Age before running executions are checked against linked task liveness
TRIGGER_STALE_QUEUED_TIMEOUT_MS300000 (5 min)Age before queued executions are checked against linked task liveness
TRIGGER_EXECUTION_HARD_MAX_RESIDENCE_HOURS48Hard maximum execution residence backstop; live linked tasks still control concurrency and incident dispatch use
TRIGGER_EXECUTION_LOG_RETENTION_DAYS90Completed/failed/skipped execution log retention
TRIGGER_EXECUTION_CLEANUP_ENABLEDenabledSet to false to disable the cleanup sweep
TRIGGER_STALE_RECOVERY_BATCH_SIZE100Maximum stale execution candidates processed per sweep
MAX_TRIGGERS_PER_PROJECT20Platform default trigger cap per project. A project owner can raise or lower it per project via UI (Project → Settings → Scaling & Scheduling → Task Limits).
VariableDefaultDescription
WEBHOOK_TRIGGERS_ENABLEDtruePublic generic webhook ingress kill switch
WEBHOOK_CREDENTIAL_CLAIM_TTL_SECONDS600Authenticated one-time MCP webhook credential claim lifetime
WEBHOOK_TRIGGER_MAX_BODY_BYTES65536Maximum JSON request body size
WEBHOOK_TRIGGER_MAX_FILTERS10Maximum deterministic filters per trigger
WEBHOOK_TRIGGER_MAX_FILTER_PATH_LENGTH200Maximum configured filter dot-path length
WEBHOOK_TRIGGER_MAX_FILTER_PATH_DEPTH8Maximum filter nesting depth at evaluation time
WEBHOOK_TRIGGER_MAX_INCLUDED_HEADERS10Maximum safe request headers copied into template context
WEBHOOK_TRIGGER_MAX_HEADER_NAME_LENGTH100Maximum configured included-header name length
WEBHOOK_TRIGGER_MAX_SOURCE_LABEL_LENGTH100Maximum optional source label length
WEBHOOK_TRIGGER_MAX_IDEMPOTENCY_KEY_LENGTH200Maximum accepted Idempotency-Key length
WEBHOOK_INGRESS_RATE_LIMIT_PER_MINUTE120Best-effort pre-auth request damping per client IP/window
WEBHOOK_TRIGGER_RATE_LIMIT_PER_MINUTE60Best-effort request damping per trigger/window
WEBHOOK_INVALID_TOKEN_RATE_LIMIT_PER_MINUTE30Best-effort invalid-token damping per client IP/window
WEBHOOK_RATE_LIMIT_WINDOW_SECONDS60Fixed rate-limit window length
WEBHOOK_DELIVERY_RETENTION_DAYS7Retention for redacted delivery audit metadata
WEBHOOK_DELIVERY_CLEANUP_BATCH_SIZE500Maximum expired audit rows deleted per cleanup pass
WEBHOOK_DELIVERY_DEFAULT_PAGE_SIZE25Default delivery-history page size
WEBHOOK_DELIVERY_MAX_PAGE_SIZE100Maximum delivery-history page size
WEBHOOK_DELIVERY_PROCESSING_LEASE_SECONDS300Lease before an unsubmitted processing delivery can recover

Webhook tokens use the existing ENCRYPTION_KEY as keyed-hash material and do not require a separate deployment secret. See Webhook Triggers for request, credential, filtering, and audit behavior.

Webhook damping uses Cloudflare KV’s eventually consistent read-update-write behavior. It reduces accidental bursts and abuse but is not a strict distributed quota.

These control whether agents can stop and ask the person who started a chat for permission, ask them a question, or send a link to open — see When the Agent Needs You. All three switches are false in the checked-in configuration; set them as GitHub Environment variables to turn them on (see Let agents ask in chat). Turning a switch on applies to agent sessions started afterwards; sleep and wake preserve their recorded interaction settings; turning one off refuses new requests at once, even in running sessions.

A permission request in a Chat session, and every question, waits up to ACP_INTERACTION_PERMISSION_CONVERSATION_DEADLINE_MS; a permission request in a Task waits up to ACP_INTERACTION_PERMISSION_TASK_DEADLINE_MS. Both are cut short a minute (ACP_INTERACTION_DEADLINE_MARGIN_MS) before the agent’s own turn would time out.

VariableDefaultDescription
ACP_INTERACTIONS_ENABLEDfalseTurns on permission requests: an agent can ask the person who started a chat before it acts, and waits for the answer. While false, SAM refuses every request at once. The other two switches need this one too. Turning it off stops new requests; ones already waiting can still be answered until they expire.
ACP_INTERACTION_FORMS_ENABLEDfalseTurns on questions: a short form (choices, short text, numbers, yes/no) the agent asks in a Chat (conversation-mode) session. Agents decide when to ask; this switch only lets them. A form SAM can’t display is cancelled. Questions and answers are stored encrypted, and only the person who started the chat can see them.
ACP_INTERACTION_URLS_ENABLEDfalseTurns on links to open: a tool, usually an MCP server, asks the person who started a Chat session to open an https:// page to sign in or approve something. SAM never opens the link itself and refuses local addresses and localhost callbacks.
ACP_INTERACTION_URL_DEADLINE_MS600000 (10 min)Maximum URL request window, further limited by the live prompt deadline. Existing requests remain answerable when the URL switch is turned off.
ACP_INTERACTION_URL_MAX_CHARS8192Maximum HTTPS URL length. Overrides can lower this bound.
ACP_INTERACTION_URL_ELICITATION_ID_MAX_CHARS256Maximum wrapper URL request ID length. Overrides can lower this bound.
ACP_INTERACTION_URL_REDIRECT_DEPTH2Maximum nesting of explicit redirect/callback query URLs. Overrides can lower this bound; SAM does not fetch redirects.
ACP_INTERACTION_FORM_SCHEMA_MAX_BYTES16384 (16 KiB)Maximum supported form schema size, including descriptions and previews.
ACP_INTERACTION_FORM_SCHEMA_MAX_PROPERTIES20Maximum fields in one supported form.
ACP_INTERACTION_FORM_SCHEMA_MAX_ENUM50Maximum choices per select or multi-select field.
ACP_INTERACTION_ANSWER_MAX_BYTES16384 (16 KiB)Maximum accepted form answer payload size.
ACP_INTERACTION_ANSWER_STRING_MAX_BYTES4096 (4 KiB)Maximum size of a string answer or selected choice.
ACP_INTERACTION_PERMISSION_TASK_DEADLINE_MS1800000 (30 min)Runtime permission deadline for task sessions.
ACP_INTERACTION_PERMISSION_CONVERSATION_DEADLINE_MS7200000 (2 hours)Runtime permission deadline for conversation sessions.
ACP_INTERACTION_DEADLINE_MARGIN_MS60000 (1 min)Safety margin before the enclosing prompt deadline when computing a permission deadline.
ACP_INTERACTION_MAX_DEADLINE_MS14400000 (4 hours)Longest any request may wait. Keep every deadline setting at or below it: while requests are on, a deadline above it makes the VM agent reject the session start, so new agent sessions fail to start until you fix it.
ACP_INTERACTION_MAX_PENDING_PER_SESSION8Most requests that can wait in one chat at once; a further request is refused.
ACP_INTERACTION_OPTIONS_MAX_COUNT16Most buttons a permission request can offer.
ACP_INTERACTION_REQUEST_MAX_BYTES32768 (32 KiB)Largest request detail SAM accepts from an agent.
ACP_INTERACTION_DELIVERY_WINDOW_MS900000 (15 min)How long SAM keeps trying to hand an answer to the agent (never past the request’s deadline) before the card says delivery is unconfirmed.
ACP_INTERACTION_SENSITIVE_PURGE_MS3600000 (1 hour)How long a request’s question and answer are kept after it is settled; after that only its outcome remains.
ACP_INTERACTION_SUMMARY_RETENTION_MS2592000000 (30 days)How long settled requests’ outcomes are kept. The newest 100 per chat are kept regardless.
ACP_INTERACTION_OPTION_ID_MAX_CHARS128Maximum permission option or form field identifier length accepted by the runtime bridge.
ACP_INTERACTION_OPTION_NAME_MAX_CHARS200Maximum permission option or form field label length accepted by the runtime bridge.
ACP_INTERACTION_RUNTIME_RECEIPT_LIMIT256Maximum in-memory idempotency receipts and distinct URL request IDs retained per live SessionHost generation. Exhausting the URL ID limit fails closed until a new runtime generation starts; retained IDs prevent late duplicate completion from binding to a new request.
ACP_INTERACTION_RUNTIME_RESPONSE_MAX_BYTES65536 (64 KiB)Maximum response body read by the runtime for interaction creation and settlement.

The remaining ACP_INTERACTION_* settings (*_BATCH_SIZE, RETRY_*, ALARM_*, *_LAST_SETTLED) tune how SAM stores and delivers requests internally and rarely need changing; their defaults are in packages/shared/src/acp-interactions.ts.

VariableDefaultDescription
ACP_SESSION_DETECTION_WINDOW_MS300000 (5 min)Stale ProjectData ACP heartbeat detection window. VM sessions are not interrupted solely from stale/missing ProjectData heartbeat rows; the timeout must be paired with conclusive runtime/workspace evidence.
ACP_SESSION_HEARTBEAT_INTERVAL_MS60000 (60s)How often VM agent sends heartbeats
ACP_SESSION_RECONCILIATION_TIMEOUT_MS30000 (30s)VM agent startup reconciliation timeout
ACP_SESSION_MAX_FORK_DEPTH10Maximum session fork chain depth
ACP_SESSION_FORK_CONTEXT_MESSAGES20Context messages included when forking

Task-managed ACP prompts use control-plane inactivity classification and the task absolute ceiling. ACP_TASK_PROMPT_TIMEOUT has been removed; legacy values no longer impose a duration-only failure. ACP_PROMPT_TIMEOUT applies only to unmanaged workspace sessions.

VariableDefaultDescription
ACP_MESSAGE_BUFFER_SIZE5000Buffer size for ACP messages
ACP_STDERR_BUFFER_BYTES4096Agent stderr bytes retained for crash reports
ACP_PING_INTERVAL30sWebSocket keepalive ping interval
ACP_PONG_TIMEOUT10sPong response timeout
ACP_PROMPT_RETRY_MAX_RETRIES2Max transient provider prompt retries after the initial attempt
ACP_PROMPT_RETRY_INITIAL_BACKOFF15sInitial backoff before retrying transient provider prompt errors
ACP_PROMPT_RETRY_MAX_BACKOFF2mMax exponential backoff for transient provider prompt retries
ACTIVITY_REREPORT_INTERVAL60sRe-send prompting activity while a prompt is active
ACP_HARNESS_ACTIVITY_REPORT_DEBOUNCE750msDebounce ACP harness/tool-call activity reports before callbacks
ACP_CHECKPOINT_PREEMPT_GRACE30sGraceful ACP cancel/close wait before harness force-stop
ACP_CHECKPOINT_PREEMPT_MAX_GRACE2mMaximum caller-selected checkpoint rollover grace
ACP_CHECKPOINT_ROLLOVER_TIMEOUT2mFull checkpoint restart and strict LoadSession deadline
ACTIVITY_TERMINAL_REPORT_ATTEMPTS5Retry attempts for terminal activity reports
ACTIVITY_TERMINAL_REPORT_BACKOFF1sBackoff between terminal activity report retries
ACP_IDLE_SUSPEND_TIMEOUT30mIdle session auto-suspend timeout
ACP_NOTIF_SERIALIZE_TIMEOUT5sNotification serialization timeout
VariableDefaultDescription
MCP_TOKEN_TTL_SECONDS28800 (8 hours)Sliding inactivity timeout for agent MCP access
MCP_RATE_LIMIT120Max MCP requests per window
MCP_RATE_LIMIT_WINDOW_SECONDS60Rate limit window
MCP_DISPATCH_MAX_DEPTH3Max recursion depth for dispatch_task
MCP_DISPATCH_MAX_PER_TASK5Max dispatched tasks per parent task
MCP_DISPATCH_MAX_ACTIVE_PER_PROJECT10Max active dispatched tasks per project
ORCHESTRATOR_STOP_CAS_MAX_ATTEMPTS2Task-status CAS attempts after a hard stop
VariableDefaultDescription
WHISPER_MODEL_ID@cf/openai/whisper-large-v3-turboTranscription model
MAX_AUDIO_SIZE_BYTES10485760 (10 MB)Max upload audio size
MAX_AUDIO_DURATION_SECONDS60Max recording duration
RATE_LIMIT_TRANSCRIBE30Max transcriptions per user per window
RATE_LIMIT_TRANSCRIBE_WINDOW_SECONDS60Transcription rate-limit window (seconds)
TTS_ENABLEDtrueEnable/disable text-to-speech
TTS_MODEL@cf/deepgram/aura-2-enTTS model
TTS_SPEAKERlunaTTS voice selection
TTS_ENCODINGmp3Audio output format
TTS_MAX_TEXT_LENGTH100000Max characters per TTS synthesis
TTS_TIMEOUT_MS60000TTS synthesis timeout
VariableDefaultDescription
TASK_RUN_MAX_EXECUTION_MS14400000 (4 hr)Age from which each stuck-task sweep checks an in_progress task’s task-scoped runtime liveness. A conclusively dead runtime is failed; a live one is kept, bounded only by TASK_RUN_ABSOLUTE_CEILING_MS
TASK_STUCK_QUEUED_TIMEOUT_MS1200000 (20 min)Timeout for tasks stuck in queued state
TASK_STUCK_DELEGATED_TIMEOUT_MS1860000 (31 min)Timeout for tasks stuck in delegated state
TASK_DO_MISMATCH_GRACE_MS300000 (5 min)Minimum age before reconciling completed TaskRunner state with task-scoped liveness
STUCK_TASK_MAX_CANDIDATES_PER_SWEEP100Maximum active tasks inspected by each recovery sweep
STUCK_TASK_SCAN_CURSOR_KV_KEYscheduled:stuck-tasks:scan-cursor:v1KV key used to resume bounded recovery scans fairly across active tasks
TASK_LIVENESS_MAX_ACP_SESSIONS5Maximum task-scoped ACP sessions inspected per liveness probe
TASK_LIVENESS_PROBE_TIMEOUT_MS5000 (5 sec)Per-candidate timeout for ACP and Instant lifecycle probes used by ProjectData heartbeat deferral, idle cleanup, and stuck-task reconciliation; a timeout is inconclusive and preserves the task and workspace
TASK_LIVENESS_NODE_HEALTH_PROBE_TIMEOUT_MS5000 (5 sec)Per-candidate timeout for stale-VM-node health probes used by ProjectData idle cleanup and stuck-task reconciliation; a timeout is inconclusive and preserves the task and workspace
IDLE_CLEANUP_MAX_CANDIDATES_PER_SWEEP5Maximum exact-session task candidates inspected by a ProjectData idle-cleanup pass; workspace deletion is deferred when this bound cannot prove every reporter-scoped runtime conclusively dead
IDLE_CLEANUP_MAX_RESIDENCE_MS7200000 (2 hr)Maximum residence for a ProjectData idle-cleanup schedule before repeated preserved/error outcomes stop re-arming, preserve the workspace, and surface an attention marker
WORKSPACE_IDLE_TIMEOUT_MS7200000 (2 hr)Installation default for how long an active chat session’s workspace can go without messages or terminal activity before ProjectData retires it, once its runtime is conclusively dead; a project’s Workspace Idle Timeout (30 min to 24 hr) overrides it
WORKSPACE_IDLE_BACKOFF_BASE_MS600000 (10 min)First retry delay after a ProjectData workspace-idle check finds an idle workspace it cannot retire yet: inconclusive task candidates, a live or unprovable runtime, a missing project identity, or a failed check
WORKSPACE_IDLE_BACKOFF_MAX_MS21600000 (6 hr)Maximum retry delay for repeated workspace-idle checks that cannot retire an idle workspace; the delay doubles from the base and resets on new activity or when the session wakes
TASK_RUN_ABSOLUTE_CEILING_MS86400000 (24 hr)Absolute runaway-cost ceiling; fails even a task with a demonstrably live runtime. Measured from the age of the currently allocated runtime generation (workspaces.created_at), not from tasks.started_at, and applied only while a runtime generation exists — a conversation whose workspace has been released holds no compute for it to bound
TASK_RUN_ABSOLUTE_CEILING_SLEEP_GRACE_MS3600000 (1 hr)Longest the absolute ceiling waits for a sleep that is still only in flight (scheduled, capturing, stopping or retrying) once the ceiling has passed. Measured on runtime-generation age, so sleep retries cannot renew it; a restorable sleep record always defers the ceiling. Keep it above SESSION_SLEEP_IN_FLIGHT_MAX_AGE_MS
STALLED_TASK_CLASSIFIER_ENABLEDtrueOnce a task is past TASK_RUN_MAX_EXECUTION_MS and kept alive only by an open agent turn, ask a Workers AI classifier whether the turn is stalled; a confident stalled verdict fails the task with “SAM detected a stalled agent turn after N minutes”. Set false to keep such turns running until the absolute ceiling
STALLED_TASK_CLASSIFIER_MIN_ACTIVITY_AGE_MS3600000 (1 hr)How long the agent’s current turn (or its running work) must have gone on, and the transcript stayed quiet, before the classifier is asked
STALLED_TASK_CLASSIFIER_CONFIDENCE_THRESHOLD0.8Stalled probability required to fail the task; uncertain or failed classifications keep the task running
STALLED_TASK_CLASSIFIER_MODEL@cf/cloudflare/clefWorkers AI model used for the stall classification
STALLED_TASK_CLASSIFIER_SELECTORclefClef selector passed with each classification
STALLED_TASK_CLASSIFIER_TIMEOUT_MS10000 (10 sec)Per-classification timeout; a timeout counts as uncertain
STALLED_TASK_CLASSIFIER_MESSAGE_LIMIT200Most recent transcript rows read for the classification
STALLED_TASK_CLASSIFIER_TRANSCRIPT_MAX_CHARS24000Maximum transcript characters, lightly redacted, sent to the classifier
CLAUDE_CODE_COMPACTION_LOOP_DETECTOR_ENABLEDtrueEnable Claude Code compaction-loop shutdown from recent message evidence
CLAUDE_CODE_COMPACTION_LOOP_RECENT_MESSAGE_LIMIT40Recent task-session messages to inspect for compaction-loop evidence
CLAUDE_CODE_COMPACTION_LOOP_WINDOW_MESSAGES20Rolling recent-message window used for compaction-loop detection
CLAUDE_CODE_COMPACTION_LOOP_MIN_PAIRS3Minimum Compacting... / Compacting completed marker pairs before failing a task
TASK_CALLBACK_TIMEOUT_MS10000Callback response timeout
TASK_CALLBACK_RETRY_MAX_ATTEMPTS3Max callback retry attempts
TASK_RUN_CLEANUP_DELAY_MS5000Delay before task cleanup
TASK_RECONCILIATION_IDLE_MS300000 (5 min)Idle threshold before SAM sends a visible task check-in
TASK_RECONCILIATION_RESPONSE_DEADLINE_MS60000 (1 min)Response deadline after a visible task check-in
TASK_RECONCILIATION_PROMPT_SOFT_STALL_MS1800000 (30 min)In-flight prompt observation threshold before a non-interrupting reconciliation event
TASK_RECONCILIATION_PROMPT_HARD_STALL_MS7200000 (2 hr)In-flight prompt hard-stall threshold before SAM requests prompt cancellation
TASK_RECONCILIATION_ACTIVE_WORK_HARD_STALL_MS7200000 (2 hr)Hard ceiling for deferring an expired task check-in because prompt/tool work is still active
TASK_RECONCILIATION_MIN_ALARM_DELAY_MS10000 (10 sec)Minimum delay before the next reconciliation alarm can fire
TASK_RECONCILIATION_MAX_CANDIDATES_PER_SWEEP5Maximum reconciliation candidates assessed per ProjectData alarm pass
TASK_RECONCILIATION_NODE_CALL_TIMEOUT_MS5000 (5 sec)Bounded timeout for reconciliation check-in delivery and prompt cancellation
TASK_RECONCILIATION_CANDIDATE_LEASE_MS30000 (30 sec)Durable claim floor preventing overlapping alarms from repeating reconciliation; the effective lease is clamped to cover the configured liveness-probe, node-call, and minimum-alarm-delay budgets
TASK_RECONCILIATION_MAX_CHECKINS3Maximum automatic check-ins without confirmed tool progress. At the limit, pause nudges and ask Clef once; human input or new completed tool work resets the budget. Classifier failure never grants more retries.
TASK_RECONCILIATION_PROBE_MAX_ATTEMPTS3Consecutive inconclusive task reconciliation attempts before quarantine
TASK_RECONCILIATION_QUARANTINE_MS300000 (5 min)Cooldown after task reconciliation exhausts its inconclusive-attempt budget
INSTANT_START_STALE_TIMEOUT_MS600000 (10 min)How long an Instant session may sit mid-launch (execution step instant_persistence) before the recovery sweep treats its start as stuck and fails it. Instant starts are accepted and then finished in the background, so this bounds a launch that never completes.

Durable prompt delivery and checkpoint storage

Section titled “Durable prompt delivery and checkpoint storage”

Durable prompt delivery is enabled by default so a follow-up can remain queued while a sleeping VM is replaced and restored. Legacy VM compatibility remains disabled: targets must advertise stable delivery receipts, and receipt ambiguity fails visibly rather than being guessed or replayed.

VariableDefaultDescription
DURABLE_PROMPT_DELIVERY_ENABLEDtruePersist prompts and deliver them from ProjectData alarms, including sleeping-session wake.
PROMPT_DELIVERY_LEGACY_VM_COMPAT_ENABLEDfalseExplicit old-VM compatibility switch; receipt ambiguity still fails visibly and is never guessed or replayed.
PROMPT_DELIVERY_MAX_CANDIDATES_PER_ALARM5Maximum delivery claims started by one alarm pass.
PROMPT_DELIVERY_MAX_ATTEMPTS5Counted delivery attempts before retryable busy/not-ready waits use capped backoff; TTL remains the hard bound.
PROMPT_DELIVERY_RETRY_BASE_MS5000Initial retry delay.
PROMPT_DELIVERY_RETRY_MAX_MS300000Maximum exponential retry delay.
PROMPT_DELIVERY_TTL_MS3600000Maximum unresolved delivery lifetime.
PROMPT_DELIVERY_RECEIPT_TIMEOUT_MS30000Age at which an unconfirmed claim enters receipt reconciliation.
PROMPT_DELIVERY_BACKGROUND_TIMEOUT_MS5000Deadline for pre-send target/recovery preparation; also bounds background VM submit and receipt calls. Preparation timeouts remain retryable without sending a prompt.
PROMPT_DELIVERY_MIN_ALARM_DELAY_MS1000Minimum delay before the next delivery alarm.
ACP_LONG_TURN_SUPERVISOR_ENABLEDfalseReserved long-turn candidate/preemption engine switch; this release leaves it inert.
ACP_LONG_TURN_CHECKPOINT_MS18000000 (5 hr)Reserved checkpoint eligibility threshold.
ACP_CHECKPOINT_PREEMPT_GRACE_MS30000Reserved graceful preemption window.
ORCHESTRATOR_WAIT_RECONCILE_INTERVAL_MS30000D1 reconciliation backstop interval for active parent waits.
ORCHESTRATOR_WAIT_MAX_CHILDREN20Maximum same-project task IDs selected by one durable wait (hard ceiling: 90, preserving D1 bind headroom).
ORCHESTRATOR_WAIT_MAX_ACTIVE_PER_PROJECT100Maximum active durable parent waits per project.
ORCHESTRATOR_WAIT_MAX_DURATION_MS86400000Maximum finite wait deadline.
ORCHESTRATOR_WAIT_MAX_CANDIDATES_PER_ALARM10Maximum wait subscriptions reconciled by one ProjectData alarm.

ProjectData stores a single prompt-delivery queue and checkpoint episodes keyed by ACP session and prompt epoch. Sleeping-session prompts stay in that queue until strict restore succeeds, then use stable receipts for exactly-once acceptance. Task agents can register wait_for_subtasks for same-project tasks; terminal hooks provide low-latency nudges, bounded alarms reconcile missed writers, and one stable delivery ID wakes the caller exactly once. Automatic checkpoint preemption remains disabled.

Liveness-gated recovery. Stuck-task recovery for in_progress tasks (including task-mode work paused at the awaiting_followup execution step) is gated on task-scoped liveness — a live workspace, node reachability, and an active task-scoped ACP session. A healthy, recent D1 node mirror satisfies the reachability gate directly; a stale or unhealthy mirror is only suspect and must be contradicted by a successful bounded node-health probe before task-scoped ACP classification continues. A shared-node heartbeat or health response alone is never sufficient. Consequently, TASK_RUN_MAX_EXECUTION_MS is the point from which a task with no proven live runtime is failed; a task with a demonstrably live runtime is preserved past it, bounded only by TASK_RUN_ABSOLUTE_CEILING_MS as a runaway-cost backstop (there is no separate hard timeout). When liveness cannot be determined (including a probe timeout, transport error, or failed health response), the task is left untouched (fail-safe). A live verdict does not mean work is happening: the task’s own agent may simply be alive and idle, waiting for its user. Each preserve is logged as stuck_task.skipped_active_heartbeat with the liveness basis (livenessReason), the agent’s work state (idle, prompt_turn_active, runtime_work_active, or prompt_turn_unproven for a turn that has reported nothing past SESSION_ACTIVITY_STALE_THRESHOLD_MS, logged as a warning) with its ages, and, for an idle agent, whether a sleep is in flight to release it. One stuck_task_heartbeat_skip row per task and basis is kept in the observability database. Idle runtimes are released by automatic sleep, not by stuck-task recovery (apps/api/src/scheduled/stuck-task-live-runtime.ts).

A sleeping conversation is never failed. Sleeping is what deletes a workspace and destroys its node, so a slept session is indistinguishable from a dead one by those signals alone. Before any terminal verdict, recovery reads the chat session’s own session_snapshots row through the same predicate the wake path uses to authorize a restore (restorableOrInFlightSleepSnapshotPredicateSql in apps/api/src/services/session-snapshot-sleep-predicate.ts, via loadTaskSleepPreservation). A session that is asleep and restorable, or with a sleep capture still in flight, is preserved with no status change and no error message — a restorable one including past TASK_RUN_ABSOLUTE_CEILING_MS, because a released runtime holds no compute for the cost ceiling to bound. The preserve is bounded on both arms and cannot strand a task: the restorable arm expires with the snapshot’s own expires_at, and the in-flight arm with SESSION_SLEEP_IN_FLIGHT_MAX_AGE_MS after the sleep’s latest attempt. Because every retry renews that age, a sleep that is only in flight on still-allocated compute defers the ceiling for at most TASK_RUN_ABSOLUTE_CEILING_SLEEP_GRACE_MS, measured on runtime-generation age (evaluateRunawayCostCeiling in apps/api/src/scheduled/stuck-task-ceiling.ts); past it the ceiling terminalizes the task and the terminal gate does not re-defer to that sleep. Once either lapses, the ordinary liveness path terminalizes the task with its accurate reason. A failed snapshot read withholds the verdict until the next sweep.

VariableDefaultDescription
NODE_AGENT_READY_TIMEOUT_MS900000 (15 min)Wait for VM agent to report ready
NODE_AGENT_READY_POLL_INTERVAL_MS5000Poll interval for agent readiness
VM_AGENT_REQUIRED_VERSION(deploy-generated)Required vm-agent build for reusable VM nodes. Official deploys derive this from the last commit that changed a vm-agent build input, not from the deployment commit, so a Worker-only deploy does not make every running node ineligible for reuse. Build inputs are packages/vm-agent/** minus exactly four pathspecs — .claude/, the top-level AGENTS.md, **/*_test.go and *_test.go — so a commit touching only those deliberately leaves the release, and the node pool, untouched. Any other file under that directory rotates the release even if it reads as documentation; the resolver excludes rather than includes so that an unrecognised new file fails safe. Binaries are published under that immutable release key first, and cloud-init requests that exact release. Leave unset only for local/manual development or skip-agent deploys.
TASK_RUNNER_STEP_MAX_RETRIES3Max retries per TaskRunner step before failing the task
TASK_RUNNER_RETRY_BASE_DELAY_MS5000Base delay for TaskRunner retry backoff
TASK_RUNNER_RETRY_MAX_DELAY_MS60000Maximum delay for TaskRunner retry backoff
TASK_RUNNER_AGENT_POLL_INTERVAL_MS5000TaskRunner D1 poll interval while waiting for a fresh provisioned VM agent
TASK_RUNNER_AGENT_READY_TIMEOUT_MS900000 (15 min)TaskRunner max wait for a provisioned VM agent before failing the task
TASK_RUNNER_AGENT_READY_FRESHNESS_SKEW_MS30000Timestamp skew tolerated between TaskRunner wait start, node heartbeat, and /ready signals
TASK_RUNNER_WORKSPACE_DISPATCH_TIMEOUT_MS600000 (10 min)Max wait for VM-agent workspace dispatch acknowledgement
TASK_RUNNER_WORKSPACE_DISPATCH_BASE_DELAY_MS30000Base delay for workspace dispatch retry backoff
TASK_RUNNER_WORKSPACE_DISPATCH_MAX_DELAY_MS120000 (2 min)Maximum delay for workspace dispatch retry backoff
TASK_RUNNER_WORKSPACE_READY_TIMEOUT_MS1800000 (30 min)Max wait for workspace-ready callback
TASK_RUNNER_WORKSPACE_READY_POLL_INTERVAL_MS30000D1 poll interval during the TaskRunner workspace-ready step
TASK_RUNNER_PROVISION_POLL_INTERVAL_MS10000TaskRunner provision status poll interval
TASK_RUNNER_PROVISION_TIMEOUT_MS900000 (15 min)Max time for TaskRunner node provisioning before permanent failure
PROVISIONING_TIMEOUT_MS1800000 (30 min)Cron marks stuck workspaces as error
NODE_HEARTBEAT_STALE_SECONDS180Seconds without a heartbeat before a node is treated as stale

Each VM runs the SAM agent and the workspace containers side by side. systemd slices keep the agent from being starved by the work it is supervising: sam-infra.slice holds the agent, sam-workload.slice is Docker’s cgroup parent, and both sit under sam.slice.

Memory and CPU are protected differently on purpose. Memory exhaustion kills the agent, so it gets a hard reservation. CPU contention only slows it, so it gets a proportional share — which Linux applies only when something is actually competing, and therefore costs nothing on an idle machine.

VariableDefaultDescription
SAM_INFRA_SLICE_MEMORY_MIN_MB256systemd MemoryMin reserved for the agent slice
VM_AGENT_MEMORY_RESERVE_MB(set by deployment)Memory held back from the workload slice’s MemoryMax, leaving headroom for the agent
SAM_INFRA_SLICE_CPU_WEIGHT1000systemd CPUWeight for the agent slice (cgroup v2 range 1–10000). Higher than the workload weight so a saturated node cannot delay the heartbeat and get the node declared dead
SAM_WORKLOAD_SLICE_CPU_WEIGHT100systemd CPUWeight for the Docker workload slice — the cgroup v2 default

An out-of-range weight makes the unit fail to load, which would take the slice hierarchy and its memory reservation down with it, so cloud-init generation rejects anything outside 1–10000 rather than emitting it.

VariableDefaultDescription
DEPLOY_PAYLOAD_EXPIRY_SECONDS3600Signed deployment apply payload lifetime
DEPLOYMENT_ROUTE_PORT_BASE35000First node-local loopback port reserved for app routes
DEPLOYMENT_ROUTE_PORT_SPAN100Number of loopback ports reserved per deployment environment
AGENT_DEPLOYMENT_RESERVED_ENVIRONMENT_NAMESprod,productionComma-separated environment names agents cannot create through MCP
MAX_ENVIRONMENTS_PER_DEPLOYMENT_NODE5Maximum deployment environments to place on one deployment node
DEPLOYMENT_DEFAULT_VM_SIZEsmallDefault VM size for deployment nodes
DEPLOYMENT_MODEL_RUNNER_VM_SIZEmediumVM size for deployment nodes that need Docker Model Runner
DEPLOYMENT_DEFAULT_CPU_LIMIT_MILLIS250Per-service CPU reservation when the whole resource block is omitted
DEPLOYMENT_DEFAULT_MEMORY_LIMIT_MB256Per-service memory limit/reservation when the whole resource block is omitted
DEPLOYMENT_DEFAULT_ROOT_DISK_MB1024Per-service root-disk reservation used for deployment placement
DEPLOYMENT_LOG_MAX_SIZE10mDefault json-file log max-size for compose-publish releases
DEPLOYMENT_LOG_MAX_FILE3Default json-file log max-file for compose-publish releases
MCP_DEPLOYMENT_COMPOSE_PREVIEW_MAX_BYTES128000Max Compose YAML size accepted by deployment route preview MCP tool
BUILD_PUBLISH_TOOL_TIMEOUT_MS1260000Worker-to-VM proxy timeout for build_and_publish
DEPLOY_ACME_EMAIL(unset)Optional ACME contact email emitted into deployment-node Caddy config
DEPLOY_ACME_CA(unset)Optional ACME CA directory override, useful for Let’s Encrypt staging
DOH_RESOLVER_URLhttps://cloudflare-dns.com/dns-queryDNS-over-HTTPS resolver used to verify deployment custom domains
DOH_TIMEOUT_MS10000Timeout for deployment custom-domain DNS verification lookups
DEPLOY_COMPOSE_CMDdocker composeDocker Compose command used by the deployment engine
DEPLOY_HEALTH_TIMEOUT5mDeployment health-check timeout used by the VM agent
DEPLOY_RUNTIME_TIMEOUT15mVM-agent max time for deployment-node host dependency setup
GRACEFUL_SHUTDOWN_TIMEOUT30sVM-agent max time for graceful HTTP server shutdown after SIGTERM
SYSTEM_PROVISIONING_TIMEOUT15mVM-agent max time for workspace host provisioning before bootstrap
CF_IP_FETCH_TIMEOUT10sVM-agent timeout for Cloudflare IP range fetches during provisioning
BOOT_LOG_HTTP_TIMEOUT10sVM-agent timeout for boot-log callbacks to the control plane
MCP_SHORT_COMMAND_TIMEOUT10sVM-agent timeout for short MCP workspace command probes
MCP_DIFF_COMMAND_TIMEOUT30sVM-agent timeout for MCP diff-summary git commands
MCP_BUILD_PREPARE_TIMEOUT30sVM-agent timeout for MCP build/publish preparation probes
JWKS_FETCH_TIMEOUT10sVM-agent startup JWKS fetch timeout
ACP_CREDENTIAL_SYNC_TIMEOUT10sVM-agent ACP auth-file sync-back timeout during shutdown
ACP_RESTART_ATTEMPT_TIMEOUT5mVM-agent bound on one automatic ACP agent restart attempt
ACP_ACTIVITY_REPORT_TIMEOUT10sVM-agent timeout for each ACP activity callback attempt
ACP_USAGE_PROBE_TIMEOUT10sVM-agent bound on one post-turn provider usage probe (Codex rollout read or OpenCode Go usage request)
OPENCODE_GO_USAGE_URLhttps://opencode.ai/zen/go/v1/usageVM-agent OpenCode Go usage endpoint probed with the session’s key after each completed opencode-go turn Must be https unless the host is loopback (localhost, 127.0.0.1, ::1), because the probe sends the key as a bearer token; the probe never follows redirects.
DEVCONTAINER_CACHE_PUSH_TIMEOUT10mVM-agent best-effort devcontainer cache image push timeout
DEPLOY_PREFLIGHT_COMMAND_TIMEOUT15sVM-agent deployment preflight diagnostic command timeout
LOG_STREAM_PING_WRITE_TIMEOUT10sVM-agent log-stream WebSocket ping write deadline
DEPLOY_TEARDOWN_TIMEOUT2mVM-agent max time for deployment environment teardown (stop/start)
DEPLOY_APPLY_IDLE_TIMEOUT15mVM-agent idle watchdog for deployment apply (no-progress only)
DEPLOY_BUILD_PUBLISH_TIMEOUT20mVM-agent max time for host build + push + release publish
DEPLOY_ARTIFACT_DIAL_TIMEOUT30sVM-agent TCP dial timeout for artifact downloads
DEPLOY_ARTIFACT_TLS_HANDSHAKE_TIMEOUT15sVM-agent TLS handshake timeout for artifact downloads
DEPLOY_ARTIFACT_RESPONSE_HEADER_TIMEOUT60sVM-agent first-response-header timeout for artifact downloads
DEPLOY_ARTIFACT_IDLE_TIMEOUT2mVM-agent idle watchdog for artifact body-read progress
DEFAULT_RESOURCE_EVENT_BUFFER_SIZE64Capacity of each bounded pressure/Docker event queue; positive integer
DEFAULT_PSI_POLL_INTERVAL_SECONDS10Linux memory PSI sampling interval, in seconds
DEFAULT_CONTAINER_STATS_INTERVAL_SECONDS30Docker resource statistics sampling interval, in seconds
DEFAULT_PSI_MEMORY_SOME_WARNING_THRESHOLD25Warning threshold for the maximum PSI some-memory avg10/avg60 percentage
DEFAULT_PSI_MEMORY_SOME_CRITICAL_THRESHOLD50Critical threshold for the maximum PSI some-memory avg10/avg60 percentage
DEFAULT_PSI_MEMORY_FULL_WARNING_THRESHOLD10Warning threshold for the maximum PSI full-memory avg10/avg60 percentage
DEFAULT_PSI_MEMORY_FULL_CRITICAL_THRESHOLD25Critical threshold for the maximum PSI full-memory avg10/avg60 percentage
DEFAULT_EVICTION_DEBOUNCE_SECONDS30Minimum cooldown between ResourceGuard eviction attempts
DEFAULT_EVICTION_SNAPSHOT_TIMEOUT_SECONDS120VM-agent pre-stop ResourceGuard eviction snapshot deadline
DEFAULT_EVICTION_DOCKER_STOP_TIMEOUT_SECONDS10Grace period passed to docker stop --time during ResourceGuard eviction
DEFAULT_EVICTION_CALLBACK_RETRY_MAX_SECONDS300Backoff cap for durable eviction callback retries, in seconds; the operation lease is a lower bound and can exceed this cap. Delivery starts on a later heartbeat
DEFAULT_EVICTION_RESOLVE_TIMEOUT_SECONDS5Docker label resolution deadline before ResourceGuard eviction
RESOURCE_HISTORY_SAMPLE_INTERVAL5sVM-agent retained resource-history cgroup sampling cadence
RESOURCE_HISTORY_CHUNK_INTERVAL15mVM-agent retained resource-history chunk duration before upload
RESOURCE_HISTORY_SPOOL_DIR/var/lib/vm-agent/resource-historyNode-local retry spool for resource-history chunks
RESOURCE_HISTORY_SPOOL_MAX_BYTES20971520Max node-local resource-history retry spool bytes
RESOURCE_HISTORY_UPLOAD_TIMEOUT10sVM-agent deadline for one resource-history upload callback
RESOURCE_HISTORY_MAX_SAMPLES4096Max resource samples packed into one uploaded chunk
VariableDefaultDescription
MAX_NODES_PER_USER10Max nodes per user
TASK_RUN_NODE_CPU_SHARE_BUDGET_PERCENT100CPU millicore share budget per node for aggregate workspace reservations
TASK_RUN_NODE_HOST_MEMORY_RESERVE_MB512Memory reserved for the host/VM agent before admitting occupied-node packing
TASK_RUN_NODE_DISK_PRESSURE_THRESHOLD_PERCENT90Fresh node disk telemetry at or above this percent vetoes VM workspace reuse
TASK_RUN_NODE_METRICS_TTL_MS180000Freshness window for occupied-node resource telemetry
TASK_RUN_NODE_CPU_SCORE_WEIGHT_PERCENT40CPU weight in existing-node load scoring after load average is normalized by vCPU
TASK_RUN_NODE_MEMORY_SCORE_WEIGHT_PERCENT60Memory weight in existing-node load scoring
VM_ADMISSION_CONTROL_MODEenforceVM task/session admission mode: off, shadow, or enforce
VM_ADMISSION_LEASE_TTL_MS1200000 (20 min)Fenced provisioning-claim lease duration
VM_ADMISSION_RETRY_MIN_MS15000Minimum retry delay for tasks waiting on VM capacity
VM_ADMISSION_RETRY_MAX_MS60000Maximum retry delay for tasks waiting on VM capacity
VM_ADMISSION_WAIT_TIMEOUT_MS7200000 (2 h)Maximum visible wait for VM capacity before failing the task
VM_ADMISSION_PROVIDER_COOLDOWN_MS600000 (10 min)Cooldown after provider/account capacity errors such as Hetzner server limits
VM_ADMISSION_WAKE_BATCH_SIZE25Maximum waiting TaskRunner DOs nudged by one capacity event
VM_ADMISSION_DIAGNOSTIC_MESSAGE_MAX_LENGTH500Maximum provider diagnostic message length stored on admission records
MAX_AGENT_SESSIONS_PER_WORKSPACE10Max concurrent agent sessions
MAX_PROJECTS_PER_USER100Max projects per user
MAX_TASKS_PER_PROJECT10000Max ideas per project
MAX_TASK_MESSAGE_LENGTH16000Max task description and reserved prompt length
RESERVED_TASK_BRANCH_NAME_SEED_MAX_LENGTH512Max branch-name seed characters accepted by reserved task submissions
RESERVED_TASK_SOURCE_DISPLAY_NAME_MAX_LENGTH512Max source display-name characters accepted by reserved task submissions
RESERVED_TASK_REPOSITORY_ACCESS_FLOW_MAX_LENGTH512Max repository-access audit flow characters accepted by reserved submissions
RESERVED_TASK_INITIAL_STATUS_REASON_MAX_LENGTH1024Max initial status reason characters accepted by reserved task submissions
RATE_LIMIT_CALLBACK_TOKEN_RENEWAL12Workspace callback-token renewals accepted per workspace in each window; past it the renewal route answers 429 with Retry-After
RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS3600Window for RATE_LIMIT_CALLBACK_TOKEN_RENEWAL
VariableDefaultDescription
MAX_SESSIONS_PER_PROJECT10000Max chat sessions per project
MAX_MESSAGES_PER_SESSION100000Max messages per chat session
COMMENT_BODY_MAX_LENGTH8000Max characters per message-anchored comment or reply body
COMMENT_QUOTE_MAX_LENGTH2000Max characters preserved from quoted message text
COMMENT_IDEMPOTENCY_KEY_MAX_LENGTH200Max clientMutationId length for message-anchored comment writes
COMMENT_LIST_LIMIT_DEFAULT100Default page size for comment thread lists
COMMENT_LIST_LIMIT_MAX500Max page size for comment thread lists
COMMENT_THREADS_PER_SESSION_MAX1000Max message-anchored comment threads per chat session
COMMENT_REPLIES_PER_THREAD_MAX200Max replies per message-anchored comment thread
PROJECT_COMMENT_LIST_LIMIT100Page size for the project-wide comment inbox
PROJECT_COMMENT_LIST_MAX300Max page size for the project-wide comment inbox
PROJECT_COMMENT_LIST_MAX_BYTES4000000Byte budget for one project-wide comment inbox response, so a few very long threads cannot exhaust the Durable Object RPC limit
DOCUMENT_CARD_RAW_OUTPUT_MAX_BYTES16384Max compact metadata bytes preserved for library document cards
PROJECT_DATA_TOOL_METADATA_MAX_BYTES131072Max stored tool_metadata bytes per message before oversized tool content is stripped into bounded metadata
PROJECT_DATA_STORAGE_TELEMETRY_ENABLEDtrueEnables ProjectData databaseSize alarm measurement and D1 telemetry writes
PROJECT_DATA_STORAGE_LIMIT_BYTES10000000000Cloudflare SQLite-backed Durable Object storage limit used for ProjectData usage classification
PROJECT_DATA_STORAGE_MEASURE_INTERVAL_MS3600000Interval between full per-object ProjectData storage measurements. Each one appends a telemetry history row and evaluates the storage alerts; cleanup passes publish the latest size but never postpone it (measureAndPersistProjectDataStorage)
PROJECT_DATA_STORAGE_ALERT_INTERVAL_MS21600000Minimum interval between repeated ProjectData storage observability alerts of the same kind and status. Threshold (warning/critical/degraded) and cleanup-target-unreachable alerts are throttled separately (maybePersistProjectDataStorageAlert)
PROJECT_DATA_STORAGE_NOTICE_RATIO0.6ProjectData storage usage ratio classified as notice
PROJECT_DATA_STORAGE_WARNING_RATIO0.8ProjectData storage usage ratio classified as warning
PROJECT_DATA_STORAGE_CRITICAL_RATIO0.9ProjectData storage usage ratio classified as critical
PROJECT_DATA_STORAGE_DEGRADED_RATIO0.95ProjectData storage usage ratio classified as degraded
PROJECT_DATA_STORAGE_EMERGENCY_TARGET_RATIO0.9Target usage ratio for explicit superadmin ProjectData emergency purge calls
PROJECT_DATA_STORAGE_EMERGENCY_BATCH_ROWS500Oldest activity_events and acp_session_events rows deleted per table per emergency purge batch
PROJECT_DATA_STORAGE_EMERGENCY_MAX_BATCHES4Maximum emergency purge batches per explicit call
PROJECT_DATA_STORAGE_GROWTH_LOOKBACK_DAYS7Lookback window used to estimate ProjectData bytes/day growth and days to storage limit
PROJECT_DATA_STORAGE_TELEMETRY_LIST_LIMIT_DEFAULT50Default row count for admin ProjectData storage telemetry and history lists
PROJECT_DATA_STORAGE_TELEMETRY_LIST_LIMIT_MAX200Max accepted row count for admin ProjectData storage telemetry and history lists
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_ENABLEDtrueEnables automatic ProjectData cleanup that archives expandable tool_metadata.content payloads to private R2 before stripping them from old message rows. NOT sufficient on its own: setting PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_CUTOFF_CREATED_AT arms the stricter approved-manifest plan, which additionally requires _PLAN_ID, _MANIFEST_KEY, _MANIFEST_SHA256 and all four _MAX_TOTAL_* ceilings. With the cutoff set and any of those blank, cleanup is enabled but builds no plan; the refusal is logged as project_data.tool_payload_cleanup.config_refused with the unmet field names.
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_PLAN_ID(empty)Immutable operator plan identifier; required with a fixed cleanup cutoff and persisted with continuation state to reject configuration drift
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MANIFEST_KEY(empty)Immutable R2 root key for the verified row-level target manifest; required with a fixed cleanup cutoff
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MANIFEST_SHA256(empty)Expected SHA-256 of the approved target-manifest root; required with a fixed cleanup cutoff
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_BATCH_MANIFEST_MAX_BYTES2000000Verified approved-plan batch-manifest read/write ceiling; included in the immutable exact-plan fingerprint
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_ROOT_MANIFEST_MAX_BYTES1000000Verified approved-plan root-manifest read/write ceiling; included in the immutable exact-plan fingerprint
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_ROWS(empty)Hard cumulative approved target-row ceiling; required with a fixed cleanup cutoff
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_BYTES(empty)Hard cumulative approved projected-reclaim ceiling; required with a fixed cleanup cutoff
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_R2_OPERATIONS(empty)Hard cumulative R2 operation ceiling including manifest verification; exact execution transactionally charges the full per-pass allowance before external work and never refunds it
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_WALL_TIME_MS(empty)Hard cumulative cleanup wall-time ceiling; exact execution transactionally charges the full per-pass allowance before external work and never refunds it
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_PROJECT_IDS(empty)Optional comma-separated project allowlist for automatic cleanup; empty preserves cleanup eligibility for every project, while an emergency rollout can scope elevated budgets to exact projects. One variable, two meanings — see the note below
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_CUTOFF_CREATED_AT(empty)Optional fixed exclusive tool-message cutoff in epoch milliseconds; when set, the immutable verified manifest, cumulative ceilings, plan ID, and exactly one allowlisted project are required; invalid values fail closed
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_TRIGGER_RATIO0.8ProjectData storage usage ratio that starts automatic tool payload archival cleanup even before the retention cadence is due
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_TARGET_RATIO0.75ProjectData storage usage ratio below which automatic tool payload cleanup stops
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_BATCH_ROWS500Maximum eligible tool-message candidates processed by one automatic cleanup alarm batch; ordinary selection may examine a larger physical row window bounded by PROJECT_DATA_STORAGE_RELIEF_MEASURE_MAX_BATCH_ROWS
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_BATCH_BYTES2097152Legacy tool_metadata read budget per automatic pass; an ordinary retention pass may admit its first oversized candidate up to the archive metadata ceiling, while a fixed exact plan requires its row ceiling to fit this budget
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_ROW_BYTES1048576Maximum single legacy tool_metadata row bytes read into JS by archival cleanup; larger rows fail closed unless the operator deliberately raises this limit
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MIN_SESSION_AGE_DAYS7Legacy terminal-session age guard retained for storage telemetry compatibility; tool payload archival uses PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_RETENTION_DAYS
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_RECHECK_MS86400000Delay before the next automatic cleanup alarm batch when more candidates remain; daily by default
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_SESSIONS_PER_ALARM25Deprecated legacy terminal-session cleanup knob retained for env compatibility; archival cleanup scans tool-message rows directly and is bounded by rows, bytes, and wall time instead
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_WALL_TIME_MS20000Absolute wall-clock deadline shared by candidate processing and every R2 write/read-back operation in one cleanup pass
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_MAX_BATCH_ROWS500Hard row cap for one explicit superadmin manual ProjectData tool payload archival cleanup pass
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_MAX_BATCH_BYTES2097152Maximum configurable legacy tool_metadata read budget for one explicit manual pass; the ordinary-path first-row exception applies unless a fixed exact plan binds the row ceiling within this budget
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_MAX_WALL_TIME_MS20000Hard wall-clock cap for one explicit manual cleanup pass
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_RECHECK_MS86400000Persisted project-scoped cooldown after an explicit manual cleanup pass; daily by default
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_RETENTION_DAYS5Message age before expandable tool payload JSON may be archived to private R2 and stripped from the ProjectData DO
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_INTERVAL_MS86400000Cadence for the retention-driven ProjectData tool payload archival scan
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_R2_PREFIXproject-data/tool-payloadsPrivate R2 prefix used for archived ProjectData tool payload JSON objects
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_WRITE_TIMEOUT_MS5000Per-R2 write/read-back verification timeout for automatic archival cleanup; timeout leaves the original payload in ProjectData and defers retry
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_MAX_OPERATIONS1500Maximum R2 PUT, GET, and body-read operations reserved by one cleanup pass
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_RETRY_DELAY_MS300000Row-level retry deferral after retryable archive/write failures
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_CHUNK_BYTES524288R2 chunk size used when archiving legacy tool payload metadata larger than one archive object slice
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_MAX_METADATA_BYTES1900000Absolute bounded read cap for legacy oversized tool payload metadata; larger rows fail closed and remain in ProjectData
PROJECT_DATA_STORAGE_RELIEF_MEASURE_BATCH_ROWS500Default row budget for superadmin ProjectData relief measurement slices
PROJECT_DATA_STORAGE_RELIEF_MEASURE_MAX_BATCH_ROWS5000Maximum physical row window accepted for superadmin/preflight relief measurements and examined by one ordinary cleanup candidate-selection slice
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_ENABLEDfalseEnables the scheduled, read-only, fixed-cutoff ProjectData tool-payload relief preflight; requires the exact plan, project, and cutoff variables below
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_PLAN_ID(empty)Immutable operator plan identifier used to resume one preflight without mixing evidence from another run
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_PROJECT_ID(empty)Exact ProjectData project targeted by the enabled preflight
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_CUTOFF_CREATED_AT(empty)Fixed exclusive tool-message creation cutoff in epoch milliseconds; required while preflight is enabled
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_BATCH_ROWS5000Maximum physical chat_messages rowid window examined by one preflight slice before eligibility filtering
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_INTERVAL_MS300000Persisted minimum interval between preflight slices
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_BATCHES100Overall claimed-attempt ceiling; an incomplete successful scan becomes truncated at the ceiling, while the final failed attempt becomes failed
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_ROWS500000Overall physical chat_messages rows-examined ceiling for one preflight plan before eligibility filtering
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_BYTES2000000000Overall projected net reclaimable-byte evidence ceiling for one preflight plan
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_LEASE_MS60000D1 claim lease that prevents overlapping slices; must exceed the wall-time budget by the configured lease margin
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_WALL_TIME_MS20000Absolute per-slice deadline shared by bounded ProjectData measurement and verified R2 manifest writes/read-backs; failures retain the prior cursor and fail closed for retry
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_SLICES_PER_RUN1Maximum sequential, separately leased slices in one scheduled invocation; later slices bypass only that invocation’s persisted cadence
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_RUN_WALL_TIME_MS25000Aggregate invocation admission budget used to decide whether another full sequential slice can start; must exceed the per-slice wall-time budget by the return margin
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_LEASE_MARGIN_MS5000Required lease headroom above the per-slice wall-time budget
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_RETURN_MARGIN_MS500Required return headroom inside measurement, slice, and aggregate run budgets
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MEASUREMENT_WALL_TIME_MS10000Explicit ProjectData measurement budget inside one slice; plus the return margin it must not exceed the slice wall budget
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_STATE_BYTES1750000Combined D1 JSON byte ceiling for accumulated session and target-batch proof state in one preflight row
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_ERROR_MAX_LENGTH1000Character ceiling for a persisted preflight failure diagnostic
PROJECT_DATA_ARCHIVE_SHARDING_ENABLEDfalseFeature switch for exact archive read routing. It is disabled when unset; enabling it alone does not run the unscoped scheduled sweep
PROJECT_DATA_ARCHIVE_GLOBAL_SWEEP_ENABLEDfalseSeparate feature switch for the unscoped scheduled archive-sharding sweep. It is disabled when unset; scheduled global migration also requires exact archive routing to be enabled
PROJECT_DATA_ARCHIVE_GLOBAL_SWEEP_INTERVAL_MS86400000Persisted cadence gate for unscoped scheduled archive-sharding sweeps; the five-minute handler claims the cadence row only when due. Code fallback is daily; the checked-in wrangler.toml ships 1080000 (18 min), which lands a claim on the fourth five-minute tick (occasionally the fifth, when the archive step starts over two minutes earlier than at the previous claim). A claim can migrate zero or more sessions (usually one, because a real candidate outlasts the PROJECT_DATA_ARCHIVE_WALL_TIME_MS gate checked between candidates), subject to the per-sweep session and message budgets and PROJECT_DATA_ARCHIVE_DAILY_WRITE_BUDGET. The interval allows roughly 72 installation-wide claims/day under the observed cron timing, so cadence alone does not set the drain rate; refused, failed or budget-limited claims publish nothing
PROJECT_DATA_ARCHIVE_SHARD_COUNT128Deterministic archive-shard fanout used when assigning terminal sessions to ProjectData archive Durable Objects
PROJECT_DATA_ARCHIVE_SWEEP_PROJECTS1Maximum projects selected by one archive-sharding cron pass
PROJECT_DATA_ARCHIVE_SWEEP_SESSIONS10Hard ceiling on sessions (in-flight plus new candidates) one archive-sharding pass may process; the checked-in wrangler.toml ships 8
PROJECT_DATA_ARCHIVE_SWEEP_MESSAGE_BUDGET20000Cumulative session_summaries.message_count of new candidates one pass may journal; candidates are selected largest first, and a single session above the budget is still selected alone; the checked-in wrangler.toml ships 10000
PROJECT_DATA_ARCHIVE_SESSION_GRACE_MS604800000Minimum terminal-session age before archive-sharding may consider a session
PROJECT_DATA_ARCHIVE_PRECOPY_REFUSAL_RETRY_MS604800000How long a session the root object refused at prepare (pre-copy invariant, e.g. active_session_state) stays out of unscoped sweep selection; journal frozen/precopy_refused, location back at root; a named-session canary bypasses it
PROJECT_DATA_ARCHIVE_FAILED_RETRY_DELAY_MS3600000Minimum age of a failed archive journal before an unscoped sweep reclaims it, so PROJECT_DATA_ARCHIVE_POISON_AFTER_ATTEMPTS spans hours rather than consecutive cadence ticks; 0 disables it and a named-session canary bypasses it
PROJECT_DATA_ARCHIVE_CHUNK_ROWS500Maximum rows exported in one idempotent transcript archive chunk
PROJECT_DATA_ARCHIVE_CHUNK_BYTES16777216Maximum bytes exported in one archive chunk; clamped below Cloudflare’s 32MiB RPC ceiling
PROJECT_DATA_ARCHIVE_HASH_PAGE_ROWS500Rows read per statement while streaming a session’s terminal-version hash inside the ProjectData Durable Object; bounds object memory by page size instead of session size
PROJECT_DATA_ARCHIVE_LEASE_MS300000D1 CAS journal lease duration for archive-sharding work
PROJECT_DATA_ARCHIVE_WALL_TIME_MS5000Soft wall-clock budget for one archive-sharding cron pass, checked between candidates (one large session may run past it). Code fallback 5 s; the checked-in wrangler.toml ships 10000 after the 2026-09-08 billing firebreak
PROJECT_DATA_ARCHIVE_ROLLOUT_LIST_LIMIT_DEFAULT25Default row limit for superadmin archive-sharding rollout inspection and failed/poisoned/frozen migration list endpoints
PROJECT_DATA_ARCHIVE_ROLLOUT_LIST_LIMIT_MAX100Maximum accepted row limit for D1-only superadmin archive-sharding rollout inspection endpoints. Frozen-intent DO inspection is capped separately.
PROJECT_DATA_ARCHIVE_FROZEN_INTENT_INSPECTION_LIMIT_DEFAULT5Default row limit for GET /archive-sharding/frozen-intents; each row may perform source and target ProjectData DO RPCs, so this endpoint intentionally has a smaller hot-object fan-out budget
PROJECT_DATA_ARCHIVE_FROZEN_INTENT_INSPECTION_LIMIT_MAX10Maximum accepted frozen-intent detail inspection limit; configured values above 10 are clamped to the reviewed hot-DO fan-out ceiling and request limits above the configured max are rejected
PROJECT_DATA_ARCHIVE_MANUAL_CANARY_MAX_SESSIONS5Maximum sessions a superadmin scoped manual archive-sharding canary request may select; default requests select one session and dry-run by default
PROJECT_DATA_ARCHIVE_MANUAL_CANARY_MAX_WALL_TIME_MS15000Maximum wall-clock budget accepted by the superadmin scoped manual archive-sharding canary endpoint
PROJECT_DATA_ARCHIVE_ROLLOUT_WARNING_EXAMPLES_MAX5Maximum malformed-row warning examples returned by archive rollout read/list endpoints
PROJECT_DATA_ARCHIVE_ROLLOUT_WARNING_REASON_MAX_LENGTH300Maximum characters per malformed-row warning reason returned by archive rollout read/list endpoints
PROJECT_DATA_ARCHIVE_POISON_AFTER_ATTEMPTS3Failed archive-sharding attempts before the migration is poisoned and the project circuit breaker opens. Each opening sends every active superadmin one Operational Failure notification and records one /admin/errors entry (alertProjectDataArchiveBreakerOpened)
PROJECT_DATA_ARCHIVE_R2_PREFIXproject-data/session-archivesPrivate R2 prefix for terminal-session archive recovery chunks and manifests
PROJECT_DATA_ARCHIVE_SEARCH_MAX_OWNERS4Archive-shard owner batch size for one project-wide message-search continuation (maximum 64; invalid/out-of-range values use 4); callers continue until coverage is complete
PROJECT_DATA_ARCHIVE_SEARCH_CONCURRENCY4Concurrent archive-owner queries in one batch (maximum 16; invalid/out-of-range values use 4)
PROJECT_DATA_ARCHIVE_SEARCH_REPAIR_SESSIONS1Incomplete published archive sessions advanced per owner query (maximum 16); remaining repair stays explicit and resumable
PROJECT_DATA_ARCHIVE_SEARCH_REPAIR_CHUNKS1Immutable compact R2 chunks consumed per session repair step (maximum 64)
PROJECT_DATA_ARCHIVE_SEARCH_CONTINUATION_TTL_MS900000Lifetime of a signed, query-bound complete-history continuation (maximum 24 hours)
PROJECT_DATA_ARCHIVE_SEARCH_CURSOR_MAX_BYTES1048576Maximum continuation bytes accepted before decoding or emitted after signing (maximum 4 MiB)
PROJECT_DATA_ARCHIVE_SEARCH_ERROR_LIMIT20Maximum entries retained independently in each public execution- and index-error collection (maximum 100); full details remain in server logs
PROJECT_DATA_SEARCH_FTS_CANDIDATE_LIMIT2000Newest full-text matches ranked per ProjectData message search; when more match, results report rootSearch.ftsCandidatesTruncated
PROJECT_DATA_SEARCH_FTS_SCAN_LIMIT20000Full-text index entries inside a session’s rowid span that a session-scoped search examines, newest first, to collect that session’s newest matches; never smaller than PROJECT_DATA_SEARCH_FTS_CANDIDATE_LIMIT
PROJECT_DATA_SEARCH_KEYWORD_SCAN_ROW_LIMIT50000Newest raw messages the keyword fallback scans for not-yet-indexed text; older rows are reported as rootSearch.keywordScanTruncated
PROJECT_DATA_ALARM_SECTION_GATING_ENABLEDtrueRun only the ProjectData alarm sections that are due each tick; false runs every section every tick
PROJECT_DATA_ALARM_FULL_RUN_INTERVAL_MS900000Maximum interval between ProjectData alarm ticks that run every section regardless of due times
PROJECT_DATA_ALARM_DUE_TOLERANCE_MS2000A ProjectData alarm section due within this many ms of the tick runs in that tick
PROJECT_DATA_ALARM_SLOW_SECTION_MS1000ProjectData alarm sections at or above this wall time log project_data.alarm.section_slow
WORKSPACE_RESOURCE_RAW_RETENTION_DAYS90Retention for immutable raw resource-history gzip chunks in the private archive R2 binding
WORKSPACE_RESOURCE_SUMMARY_RETENTION_DAYS180Retention for bounded D1 workspace resource summary rows
WORKSPACE_RESOURCE_UNCOMPRESSED_MAX_BYTES8388608Max decoded resource chunk JSON bytes accepted/read before rejecting detail payloads
WORKSPACE_RESOURCE_METADATA_MAX_BYTES8192Max summary or completeness JSON bytes stored in D1 for one resource chunk/summary
WORKSPACE_RESOURCE_TOOL_NAME_MAX_BYTES256Max UTF-8 bytes retained and returned for one resource-history tool name; longer metadata names are truncated at a complete Unicode character
WORKSPACE_RESOURCE_UPLOAD_MAX_BYTES2097152Max compressed resource-history chunk bytes accepted by the callback upload route
WORKSPACE_RESOURCE_DETAIL_MAX_POINTS720Max samples returned by one raw detail read after spike-preserving downsampling
WORKSPACE_RESOURCE_LIST_LIMIT24Max resource-history chunk index rows returned for one contextual read
WORKSPACE_RESOURCE_TIMELINE_MAX_CHUNKS1000Max chunks the session resource timeline lists; the oldest beyond it are left out and disclosed as omittedChunkCount
WORKSPACE_RESOURCE_ROLLUP_BUCKET_MS60000Width of each per-chunk rollup bucket computed on upload
WORKSPACE_RESOURCE_ROLLUP_MAX_BUCKETS60Max rollup buckets stored per chunk; the bucket width widens in whole multiples to fit
WORKSPACE_RESOURCE_CLEANUP_BATCH_SIZE50Max expired resource-history chunks and summaries processed per scheduled cleanup sweep
WORKSPACE_RESOURCE_OBJECT_CLEANUP_LIMIT5000Max resource-history R2 objects deleted when a project or workspace is deleted; deletion paginates until the prefix is empty or this safety budget is reached

A completed fixed tool-payload plan writes an explicit terminal marker rather than an indefinite timestamp. Later alarm or manual entry points no-op after validating the same immutable fingerprint. A different fingerprint also fails closed; there is deliberately no automatic reset or supersession, so an emergency runbook must include a reviewed follow-up reset before ordinary cleanup is expected to resume.

Archive-sharding rollout is deliberately staged. The scheduled coordinator in apps/api/src/scheduled/project-data-archive-sharding.ts still does no work unless PROJECT_DATA_ARCHIVE_GLOBAL_SWEEP_ENABLED=true and exact archive routing is active through PROJECT_DATA_ARCHIVE_SHARDING_ENABLED=true. Both switches are ordinary Worker [vars]: a GitHub Environment variable of the same name replaces the checked-in wrangler.toml value at deploy time, and the deploy log prints every such override that differs from wrangler.toml. When the sweep is enabled, each cron.completed log carries projectDataArchiveShardingSkipReason (disabled, exact_routing_disabled, missing_r2_binding, cadence_not_due, cadence_running, cadence_unavailable) or the tick’s selected/migrated/refused/failed counts. A session the root object refuses before any copy (for example one that still holds an active session_state row) is returned to root in the same tick with a frozen/precopy_refused journal and is skipped by the sweep for PROJECT_DATA_ARCHIVE_PRECOPY_REFUSAL_RETRY_MS.

The shipped cadence (1080000, 18 minutes) and wall budget (10000) were sized against the SAM root object (~3,350 backlog sessions plus 4-233 newly terminal sessions per day): one tick moves one non-trivial session because the wall-time break fires between candidates, so a daily tick can never keep up. Idle ticks cost two D1 statements. Per-tick copy work is bounded by the session ceiling and the message budget; the wall budget is a soft gate checked between candidates, so one archive can run past it. The sweep runs after every lifecycle sweep in the cron chain so that budget cannot delay them. failed journals wait PROJECT_DATA_ARCHIVE_FAILED_RETRY_DELAY_MS before a sweep retries them, so the three-attempt poison budget still spans hours at the shorter cadence. The frozen-intent inspection route skips precopy_refused rows (nothing exists on either object to inspect) and the problem-migrations list sorts them after rows that need a human. Before changing the global sweep switch, operators use the superadmin routes in apps/api/src/routes/admin/project-data-storage.ts to inspect one project’s D1 journal/location/breaker state and run a scoped dry-run canary for a specific project and optional session. Dry-runs work while both rollout switches are false and never call source deletion RPCs. Non-dry scoped canaries fail closed unless exact archive routing is active, because publishing or deleting source rows while exact reads still resolve to root can render conversations empty. Failed, poisoned, or frozen rows are inspected through the frozen-intent route (and listed under Admin → Storage → Problem migrations). A migration that never reached source deletion is cleared with Abandon (the page’s button, or POST .../migrations/:migrationId/abandon); one past source deletion is recovered through the copy-back helper with exact archive routing enabled. Both take an explicit operator reason, and neither closes the project’s circuit breaker.

Safe operator sequence for ProjectData storage relief:

  1. Inspect current telemetry with GET /api/admin/project-data/storage?projectId=<projectId>. Before building an approval manifest, freeze competing source mutations: disable tool-payload, event-log, and grouped/FTS cleanup, and leave both archive-sharding switches off.
  2. For an emergency exact plan, configure the immutable PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_{PLAN_ID,PROJECT_ID,CUTOFF_CREATED_AT} scope plus row, byte, attempt, slice, run-admission, state, and manifest limits. Enable the scheduled preflight only for that scope.
  3. Wait for a terminal project_data_storage_relief_preflights row. Accept only a complete or deliberately bounded truncated result with exact totals and a non-null target_manifest_key, byte count, and SHA-256. The preflight reads every batch and root manifest back from R2 and verifies bytes and SHA-256 before persisting these proofs; independently inspect the root and referenced batch proofs before approval.
  4. Disable preflight and keep cleanup frozen. Present the exact project, cutoff, row targets, manifest keys/hashes, cumulative ceilings, switches, failure behavior, expected relief, observation window, and stop conditions for explicit human approval. Broad rollout approval does not substitute for this plan approval.
  5. Only after approval, configure cleanup with the same plan/project/cutoff, exact root key/SHA-256, and cumulative row/byte/R2-operation/wall-time ceilings. Execute only approved bounded passes, using POST /api/admin/project-data/storage/:projectId/tool-payload-cleanup with a non-empty reason and unique idempotencyKey when an explicit manual slice is required.
  6. Each payload archive object/chunk is written to an immutable content-addressed R2 key and read back for byte/SHA-256 verification before a guarded transaction changes only chat_messages.tool_metadata; chat_messages.content is never changed. Stop on any manifest, archive, source-hash, scope, cap, timeout, or circuit-breaker uncertainty.
  7. After every pass, inspect terminationReason, before/after databaseSize, reclaimed bytes, row counts, cursor/recheck state, and cooldown, then re-measure with POST /api/admin/project-data/storage/:projectId/measure. Stop at the approved target or any approved stop condition.
  8. A completed fixed plan is terminal. Before ordinary retention cleanup resumes, use a separately reviewed follow-up reset; do not reuse or silently supersede the emergency plan fingerprint.
VariableDefaultDescription
PROJECT_DATA_MATERIALIZATION_PAGE_ROWS500Tokens read into memory by one chat-search materialization SELECT, bounding Durable Object isolate memory
PROJECT_DATA_MATERIALIZATION_MAX_ROWS_PER_PASS5000Tokens one chat-search materialization pass indexes before leaving the rest to the next pass
PROJECT_DATA_MATERIALIZATION_MAX_GROUP_CHARS65536Grouped-row size past which a continuing run starts a new row instead of rewriting the accumulated content
PROJECT_DATA_MATERIALIZATION_SWEEP_LIMIT50Sessions indexed by one chat-search materialization backfill call
PROJECT_DATA_MATERIALIZATION_SWEEP_SCAN_LIMIT500Sessions examined by one chat-search materialization backfill call
PROJECT_DATA_GROUPED_FTS_CLEANUP_ENABLEDfalseEnables cleanup of old terminal-session grouped message rows and their FTS entries under storage pressure. wrangler.toml ships true, so this is ON; a cleaned session falls back to keyword search and is never re-indexed
PROJECT_DATA_GROUPED_FTS_CLEANUP_TRIGGER_RATIO0.9ProjectData usage ratio that starts grouped/FTS derived-data cleanup
PROJECT_DATA_GROUPED_FTS_CLEANUP_TARGET_RATIO0.85ProjectData usage ratio below which grouped/FTS cleanup stops
PROJECT_DATA_GROUPED_FTS_CLEANUP_BATCH_SESSIONS2Maximum terminal sessions examined by one grouped/FTS slice; candidate selection pages IDs first and reads at most one extra session for continuation
PROJECT_DATA_GROUPED_FTS_CLEANUP_BATCH_ROWS1000Maximum grouped message rows deleted by one grouped/FTS canary slice
PROJECT_DATA_GROUPED_FTS_CLEANUP_BATCH_BYTES4194304Maximum grouped message content bytes deleted by one grouped/FTS canary slice
PROJECT_DATA_GROUPED_FTS_CLEANUP_MIN_SESSION_AGE_DAYS7Minimum terminal-session age before grouped/FTS derived rows may be cleaned
PROJECT_DATA_GROUPED_FTS_CLEANUP_RECHECK_MS300000Delay before the next grouped/FTS cleanup slice when more candidates remain. Also the back-off after an overload/reset error, measured from when that error was recorded
PROJECT_DATA_GROUPED_FTS_CLEANUP_WALL_TIME_MS5000Soft wall-clock budget for one grouped/FTS cleanup slice
PROJECT_DATA_GROUPED_FTS_CLEANUP_WALL_UNSAFE_RATIO0.98Refuses grouped/FTS cleanup writes when the object is too close to the configured storage limit
PROJECT_DATA_GROUPED_FTS_CLEANUP_WEAK_RECLAIM_BYTES1Stops grouped/FTS cleanup if a slice deletes rows but databaseSize does not drop by at least this many bytes
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_MAX_ROWS10000Ceiling on grouped rows one superadmin grouped/FTS wall-recovery call may prune
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_MAX_BYTES33554432Ceiling on grouped content bytes one grouped/FTS wall-recovery call may prune
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_MAX_SESSIONS500Ceiling on sessions one grouped/FTS wall-recovery call may consider, and on its skipSessionIds length
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_TRANSACTION_ROWS500Grouped rows pruned per wall-recovery transaction
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_TRANSACTION_BYTES8388608Grouped content bytes held in memory per wall-recovery transaction
PROJECT_DATA_EVENT_LOG_CLEANUP_ENABLEDtrueEnables automatic deletion of old low-value terminal-session activity_events and terminal ACP event history when storage remains above the cleanup target
PROJECT_DATA_EVENT_LOG_CLEANUP_BATCH_ROWS500Maximum terminal activity_events rows and terminal acp_session_events rows deleted per automatic cleanup alarm batch
PROJECT_DATA_EVENT_LOG_CLEANUP_MIN_SESSION_AGE_DAYS7Minimum terminal-session age before automatic event-log cleanup may delete its activity/ACP event history
PROJECT_DATA_EVENT_LOG_CLEANUP_RECHECK_MS86400000Delay before the next terminal event-log cleanup alarm batch when more candidates remain; daily by default
PROJECT_EVENT_MAX_ACTIVE_SUBSCRIPTIONS_PER_PROJECT200Maximum active durable event subscriptions in one ProjectData object
PROJECT_EVENT_FILTER_MAX_VALUES_PER_FIELD20Maximum exact/set values accepted for one v1 event filter field
PROJECT_EVENT_FILTER_MAX_MATCH_KEYS100Maximum deterministic match keys compiled for one event subscription
PROJECT_EVENT_FILTER_MAX_STRING_BYTES160Maximum bytes for event source/type/subject, filter values, delivery keys, fingerprints, and event idempotency strings
PROJECT_EVENT_METADATA_MAX_BYTES8192Maximum normalized event metadata JSON bytes stored in ProjectData
PROJECT_EVENT_METADATA_MAX_DEPTH4Maximum nesting depth for normalized event metadata
PROJECT_EVENT_METADATA_MAX_KEYS64Maximum total object keys in normalized event metadata
PROJECT_EVENT_METADATA_MAX_ARRAY_ITEMS64Maximum array items in normalized event metadata
PROJECT_EVENT_DISPLAY_MAX_BYTES4096Maximum deterministic untrusted event display JSON bytes
PROJECT_EVENT_DISPLAY_MAX_LABELS12Maximum labels accepted in event display data
PROJECT_EVENT_RAW_PAYLOAD_REF_MAX_BYTES512Maximum optional raw-payload reference JSON bytes; raw webhook bodies are not persisted in ProjectData
PROJECT_EVENT_REASON_MAX_BYTES1024Maximum diagnostic reason bytes on subscriptions, batches, and attempts
PROJECT_EVENT_MAX_MATCHES_PER_EVENT100Maximum durable subscription matches written for one admitted normalized event
PROJECT_EVENT_DELIVERY_BATCH_MAX_EVENTS50Maximum event matches in one durable delivery batch
PROJECT_EVENT_DELIVERY_ATTEMPT_MAX_PER_BATCH10Maximum recorded delivery attempts for one event delivery batch
PROJECT_EVENT_SCHEDULE_MAX_SCHEDULES128Maximum active schedules per project
PROJECT_EVENT_SCHEDULE_MAX_WATCHES64Maximum active or paused standing watches per project
PROJECT_EVENT_SCHEDULE_MAX_RETAINED_SCHEDULES4096Maximum retained schedule records of all states per project; cancellation does not free retained capacity
PROJECT_EVENT_SCHEDULE_MAX_RETAINED_WATCHES256Maximum retained watch records of all states per project; revocation does not free retained capacity
PROJECT_EVENT_SCHEDULE_PROMPT_MAX_BYTES32768Maximum scheduled action prompt size in UTF-8 bytes; start_session admission also validates MAX_TASK_MESSAGE_LENGTH
PROJECT_EVENT_SCHEDULE_MAX_HORIZON_MS2592000000Maximum future scheduling horizon (30 days)
PROJECT_EVENT_SCHEDULE_LATE_GRACE_MS86400000Maximum late admission grace after due time
PROJECT_EVENT_SCHEDULE_DELIVERY_TTL_MS86400000Maximum durable message lifetime, also bounded by schedule expiry
PROJECT_EVENT_SCHEDULE_SWEEP_BATCH_SIZE20Maximum schedule or watch candidates processed per alarm pass
PROJECT_EVENT_SCHEDULE_CLAIM_LEASE_MS60000Finite admission/submission claim lease
PROJECT_EVENT_SCHEDULE_RETRY_BASE_MS30000Retry and accepted execution reconciliation interval
PROJECT_EVENT_SCHEDULE_MAX_ATTEMPTS8Maximum uncertain task submission attempts
PROJECT_EVENT_SCHEDULE_MAX_DEFERRAL_MS86400000Maximum deferral before a new task may initially start
PROJECT_EVENT_WATCH_COOLDOWN_MIN_MS60000Minimum time between standing watch executions
PROJECT_EVENT_WATCH_MAX_EXECUTIONS100Maximum configurable executions per standing watch
PROJECT_EVENT_WATCH_MAX_CONCURRENT3Maximum configurable concurrent actions per standing watch
PROJECT_EVENT_CHANNEL_MAX_CHANNELS128Maximum catalog generations per project
PROJECT_EVENT_CHANNEL_MESSAGE_MAX_BYTES4096Maximum channel message UTF-8 bytes, also subject to canonical metadata limits
PROJECT_EVENT_CHANNEL_NAME_MAX_BYTES64Maximum channel name bytes
PROJECT_EVENT_CHANNEL_PUBLISH_WINDOW_MS60000Per-project fixed publish window in milliseconds
PROJECT_EVENT_CHANNEL_PUBLISH_MAX_PER_WINDOW120Maximum newly committed channel publishes per project window; retained replays do not consume quota
PROJECT_EVENT_CHANNEL_CURSOR_TTL_MS3600000History cursor and unfinished catch-up lifetime in milliseconds; continuation never extends it
PROJECT_EVENT_CHANNEL_CATALOG_IDLE_TTL_MS2592000000Minimum idle time before reclaiming an empty catalog generation without live catch-up
AGENT_MESSAGE_CHANNELS_ENABLEDfalseCode fallback is off; managed deployment sets true. Send notify/deliver agent messages over SAM-managed agent-dm.* pair channels. Effective only while PROJECT_EVENT_WAKE_ENABLED and durable prompt delivery are on
AGENT_MESSAGE_CHANNEL_MAX_CHANNELS1024Maximum agent-dm.* pair channels per project, separate from PROJECT_EVENT_CHANNEL_MAX_CHANNELS
AGENT_MESSAGE_SUBSCRIPTION_ROTATION_GRACE_MS300000Replace a managed pair subscription this close to the end of its wake lifetime when it owes no pending wake
AGENT_MESSAGE_MAX_ACTIVE_SUBSCRIPTIONS100Share of PROJECT_EVENT_MAX_ACTIVE_SUBSCRIPTIONS_PER_PROJECT that SAM-managed agent-message subscriptions may hold; idle ones on other pairs are released first
PROJECT_EVENT_LIST_LIMIT50Default ProjectData event-subscription list/status page size
PROJECT_EVENT_LIST_MAX200Maximum ProjectData event-subscription list/status page size
PROJECT_EVENT_SUBSCRIPTION_EVENT_CURSOR_MAX_LENGTH512Maximum opaque list_subscription_events cursor length accepted by ProjectData pull delivery
PROJECT_EVENT_RECENT_STATUS_LIMIT50Maximum rows returned per section in recent event-subscription status inspection
PROJECT_EVENT_RETENTION_DAYS30Retention window for terminal ProjectData event-subscription records
PROJECT_EVENT_RETENTION_BATCH_ROWS500Maximum old rows pruned per event-subscription retention category per pass
PROJECT_EVENT_SOURCE_OUTBOX_BATCH_ROWS25Maximum producer-side outbox row mutations per trigger cleanup pass. Must be an integer of at least 2 for claim plus settlement; invalid values reject before any writes. Per-call budget overrides have the same minimum.
PROJECT_EVENT_SOURCE_OUTBOX_MAX_ATTEMPTS8Maximum ProjectData admission attempts before a source intent is marked permanent_failed
PROJECT_EVENT_SOURCE_OUTBOX_TTL_MS86400000Maximum lifetime for a producer-side event admission intent
PROJECT_EVENT_SOURCE_OUTBOX_RETRY_BASE_MS30000Initial retry delay for failed source event admission
PROJECT_EVENT_SOURCE_OUTBOX_RETRY_MAX_MS1800000Maximum retry delay for failed source event admission
PROJECT_EVENT_SOURCE_OUTBOX_PROCESSING_LEASE_MS60000Lease before an in-flight source event admission intent can be claimed by reconciliation
PROJECT_EVENT_RETENTION_INTERVAL_MS86400000Default delay before the next ProjectData event retention alarm after a complete pass; daily by default
PROJECT_EVENT_RETENTION_MIN_ALARM_DELAY_MS60000Minimum delay before a follow-up event retention alarm when more bounded work remains or a due checkpoint is re-armed
PROJECT_EVENT_WAKE_ENABLEDfalseEnables same-chat ProjectData event wake materialization for v2 existing_session_prompt and runtime_interrupt subscriptions; opt in with true; unset or false leaves event delivery pull-only
PROJECT_EVENT_WAKE_MATERIALIZATION_MIN_ALARM_DELAY_MS1000Minimum delay before a ProjectData alarm retries wake materialization after due matches are present
PROJECT_EVENT_WAKE_MATERIALIZATION_BACKOFF_BASE_MS5000Initial persisted backoff for a failed event-wake materialization or retention alarm phase
PROJECT_EVENT_WAKE_MATERIALIZATION_BACKOFF_MAX_MS300000Maximum persisted backoff for repeated event-wake materialization or retention alarm failures
PROJECT_EVENT_WAKE_PROMPT_TTL_MS86400000Hard lifetime for queued same-chat event wake prompts before physical delivery is rejected; also bounds how long an undelivered wake holds its chat
PROJECT_EVENT_WAKE_READ_GRACE_MS86400000Read/ack grace retained on accepted event-wake batches after natural subscription expiry
PROJECT_EVENT_WAKE_TARGET_COOLDOWN_MS30000Delay before retrying an event wake for a chat whose transcript or project mailbox is at capacity. A chat with an undelivered wake is skipped until that wake is accepted, without this cooldown
PROJECT_EVENT_WAKE_SUBSCRIPTION_COOLDOWN_MS30000Per-subscription cooldown after a wake batch is queued
PROJECT_EVENT_WAKE_SUBSCRIPTION_LIFETIME_MS86400000Maximum operational lifetime for same-chat wake delivery on a v2 event subscription when the requested subscription expiry is absent or farther out
PROJECT_EVENT_WAKE_MAX_PER_SUBSCRIPTION50Maximum same-chat wake prompt batches materialized for one v2 event subscription
PROJECT_EVENT_SOURCE_OUTBOX_SWEEP_WALL_MS2500Per-pass wall-clock budget for producer-side outbox reconciliation before the sweep yields
PROJECT_EVENT_SOURCE_OUTBOX_ADMISSION_TIMEOUT_MS5000Per-intent timeout for the ProjectData admission call; timed-out calls are treated as ambiguous and retried through the fenced outbox
PROJECT_EVENT_SOURCE_OUTBOX_TERMINAL_RETENTION_MS604800000Retention window for terminal outbox rows before the bounded cleanup removes admitted, expired, and permanent_failed history
MESSAGE_SIZE_THRESHOLD102400Max message size in bytes
ACTIVITY_RETENTION_DAYS90Days to retain activity events
SESSION_IDLE_TIMEOUT_MINUTES60Idle session timeout
SESSION_ACTIVITY_STALE_THRESHOLD_MS300000 (5 min)Threshold before stale working activity is checked against authoritative SessionHost inventory
SESSION_ACTIVITY_PROBE_TIMEOUT_MS5000 (5 s)Timeout for the vm-agent session-activity probe. Background control-loop budget — deliberately far below the interactive node-agent timeout
SESSION_ACTIVITY_PROBE_MAX_ATTEMPTS3Consecutive unreachable probes after which a stale working state is quarantined until an authoritative report refreshes it
SESSION_ACTIVITY_PROBE_MAX_CANDIDATES10Stale-activity candidates probed per ProjectData alarm pass
DO_SUMMARY_SYNC_DEBOUNCE_MS5000Debounce for DO-to-D1 summary sync
SESSION_INDEX_MAX_ROWS1000Sessions mirrored into the D1 session_summaries index per project. A project holding more is recorded as incomplete and its chat sidebar reads fall back to the Durable Object
SESSION_INDEX_MAX_STALENESS_MS900000 (15 min)How stale the session index may be before the per-project sidebar list stops trusting it and falls back to the Durable Object
CREDENTIAL_LIMIT_WARNING_PERCENT75Advisory warning threshold for credential quota utilization samples
CREDENTIAL_LIMIT_CRITICAL_PERCENT90Advisory critical threshold for credential quota utilization samples
CREDENTIAL_LIMIT_MAX_OBSERVATIONS_PER_REPORT16Maximum credential-limit observations accepted from one VM usage callback or proxy report
CREDENTIAL_LIMIT_TRANSITION_RECOMPUTE_ATTEMPTS4Maximum predecessor-CAS recomputes for one credential-limit observation when concurrent samples update the same window
CREDENTIAL_LIMIT_ADMISSION_MAX_ACTIVE_PER_PROJECT1000Maximum active credential-limit event admissions per project (range: 1–100000)
CREDENTIAL_LIMIT_READ_MAX_ROWS200Maximum credential usage-limit window rows returned by one read request (GET /api/projects/:id/credential-limits, GET /api/credentials/limits, MCP get_credential_limits)
CREDENTIAL_LIMIT_ADMISSION_RETRY_BATCH_SIZE25Maximum pending credential-limit event admissions retried per batch (range: 1–500)
CREDENTIAL_LIMIT_ADMISSION_RETENTION_DAYS30Retention of inactive credential-limit window observations in days (range: 1–365)
CREDENTIAL_LIMIT_USAGE_CALLBACK_MAX_BODY_BYTES32768Maximum raw JSON bytes accepted for one authenticated VM usage callback before schema validation
CREDENTIAL_LIMIT_USAGE_CALLBACK_RATE_LIMIT_RPM120Authenticated VM usage callbacks accepted per session in each rate-limit window
CREDENTIAL_LIMIT_USAGE_CALLBACK_RATE_LIMIT_WINDOW_SECONDS60Usage callback rate-limit window in seconds
CREDENTIAL_LIMIT_OBSERVATION_MAX_AGE_MS86400000Oldest provider observation timestamp accepted for credential-limit telemetry
CREDENTIAL_LIMIT_OBSERVATION_FUTURE_SKEW_MS300000Future clock skew accepted for credential-limit observation timestamps
CREDENTIAL_LIMIT_RESET_MAX_FUTURE_MS691200000Maximum future provider reset timestamp accepted for credential-limit telemetry
CREDENTIAL_LIMIT_SUPPORTED_PROVIDERSanthropic,openai,opencodeComma-separated allowlist of providers accepted by credential-limit telemetry
CREDENTIAL_LIMIT_SUPPORTED_SOURCESbuilt-in sourcesComma-separated allowlist of VM-agent and AI-proxy telemetry source identifiers accepted by credential-limit telemetry
CREDENTIAL_LIMIT_SUPPORTED_WINDOW_TYPESbuilt-in windowsComma-separated allowlist of provider quota window identifiers accepted by credential-limit telemetry
AI_PROXY_REQUEST_BODY_MAX_BYTES1048576Maximum raw JSON bytes accepted by OpenAI-compatible, Anthropic-native, and passthrough AI proxy request endpoints before request validation
AI_PROXY_ALLOWED_MODELSplatform catalogComma-separated models the AI proxy serves on every route that spends platform credentials, the native Anthropic endpoint included; any other model returns 400

Ordinary ProjectData storage alarms record O(1) databaseSize telemetry and bounded cleanup row/byte counters. Category breakdown scans are reserved for explicit/admin measurement paths so hot ProjectData alarms do not delay lifecycle bookkeeping.

VariableDefaultDescription
DO_RETRY_MAX_ATTEMPTS8Max attempts for transient Durable Object RPC reset/overload errors
DO_RETRY_BASE_DELAY_MS100Base retry delay in milliseconds for transient Durable Object RPC failures
DO_RETRY_MAX_DELAY_MS250Max per-attempt retry delay for transient Durable Object RPC failures
DO_RETRY_CONNECTION_LOST_MAX_ATTEMPTS3Attempts (capped at DO_RETRY_MAX_ATTEMPTS) for an idempotent ProjectData read whose connection to the object was lost or that encounters SQLITE_NOMEM. Mutations never retry SQLITE_NOMEM. An idempotent read that exhausts its attempts on either error or a CPU-limit reset returns 503 PROJECT_DATA_UNAVAILABLE
PROJECT_DATA_ENSURE_MEMO_MAX_ENTRIES2000Max ProjectData Durable Objects one Worker isolate remembers as already having a persisted projectId, so ensureProjectId costs one RPC per isolate instead of one before every DO call
VariableDefaultDescription
MAX_PROJECT_RUNTIME_ENV_VARS_PER_PROJECT150Max env vars per project
MAX_PROJECT_RUNTIME_FILES_PER_PROJECT50Max files per project
MAX_PROJECT_RUNTIME_ENV_VALUE_BYTES8192Max bytes per env var value
MAX_PROJECT_RUNTIME_FILE_CONTENT_BYTES131072Max bytes per file content
MAX_PROJECT_RUNTIME_FILE_PATH_LENGTH256Max file path length
MAX_DEPLOYMENT_ENV_VARS_PER_ENVIRONMENT100Max deployment config vars per environment
MAX_DEPLOYMENT_ENV_VALUE_BYTES65536Max bytes per deployment config value
MAX_DEPLOYMENT_ENV_TOTAL_BYTES262144Max aggregate deployment config env size
MAX_MCP_CONNECTIONS_PER_SCOPE25Max bring-your-own MCP servers per scope
MCP_CONNECTION_URL_MAX_BYTES2048Max MCP endpoint URL size
MCP_CONNECTION_TOKEN_MAX_BYTES8192Max MCP bearer token size
MAX_MCP_CONNECTION_HEADERS10Max custom headers per MCP server
MCP_CONNECTION_HEADER_VALUE_MAX_BYTES8192Max bytes per MCP custom header value
VariableDefaultDescription
HETZNER_API_TIMEOUT_MS30000Hetzner API request timeout
HETZNER_CAPACITY_RETRY_INITIAL_DELAY_MS15000Initial delay for transient Hetzner capacity retry backoff
HETZNER_CAPACITY_RETRY_MAX_DELAY_MS120000Maximum delay per transient Hetzner capacity retry wait
HETZNER_CAPACITY_RETRY_MAX_ATTEMPTS10Maximum transient Hetzner capacity retry attempts
HETZNER_CAPACITY_RETRY_BUDGET_MS300000Total transient Hetzner capacity retry budget
HETZNER_MAX_LIST_PAGES100Maximum pages per Hetzner list request
CF_API_TIMEOUT_MS30000Cloudflare API request timeout
GCP_API_TIMEOUT_MS30000GCP OAuth, IAM, and Compute request timeout
NODE_AGENT_REQUEST_TIMEOUT_MS30000VM Agent request timeout
DIGITALOCEAN_API_TIMEOUT_MS30000DigitalOcean API request timeout
DIGITALOCEAN_IP_POLL_TIMEOUT_MS20000Bounded best-effort public IPv4 poll budget
DIGITALOCEAN_IP_POLL_INTERVAL_MS3000Public IPv4 poll interval
DIGITALOCEAN_ACTION_POLL_TIMEOUT_MS60000Block Storage action completion budget
DIGITALOCEAN_ACTION_POLL_INTERVAL_MS1000Block Storage action poll interval
DIGITALOCEAN_MAX_LIST_PAGES20Maximum pages per DigitalOcean list request
DIGITALOCEAN_REGIONfra1Default DigitalOcean region
DIGITALOCEAN_IMAGEubuntu-24-04-x64Default Droplet image slug
CF_CONTAINER_CREATE_WORKSPACE_TIMEOUT_MS120000Instant-session create-workspace budget (includes in-container clone)
VariableDefaultDescription
OBSERVABILITY_ERROR_RETENTION_DAYS30Error log retention
OBSERVABILITY_ERROR_MAX_ROWS100000Max stored error rows
OBSERVABILITY_ERROR_BATCH_SIZE25Error ingestion batch size
OBSERVABILITY_ERROR_MESSAGE_MAX_LENGTH2048Maximum persisted message length
OBSERVABILITY_ERROR_STACK_MAX_LENGTH4096Maximum persisted stack length
OBSERVABILITY_ERROR_USER_AGENT_MAX_LENGTH512Maximum persisted user-agent length
OBSERVABILITY_LOG_QUERY_RATE_LIMIT30Log queries per minute per admin
VariableDefaultDescription
VM_AGENT_PROTOCOLhttpsProtocol for VM agent communication
VM_AGENT_PORT8443VM agent listening port
VM_AGENT_MEMORY_RESERVE_MB512Optional Docker workload-slice memory reserve for VM-agent reachability headroom
SAM_INFRA_SLICE_MEMORY_MIN_MB256systemd MemoryMin for the VM-agent/system-services slice
DOCKER_MEMORY_MIN_MB512Minimum Docker MemoryMax retained when VM_AGENT_MEMORY_RESERVE_MB is enabled
SAM_INFRA_SLICE_CPU_WEIGHT1000systemd CPUWeight for the VM-agent slice (cgroup v2 range 1–10000)
SAM_WORKLOAD_SLICE_CPU_WEIGHT100systemd CPUWeight for the Docker workload slice (cgroup v2 default)
HEARTBEAT_DOCKER_STATS_TIMEOUT2sVM-agent timeout for heartbeat Docker stats used by workspace memory telemetry
HEARTBEAT_WORKSPACE_METRICS_MAX_CONTAINERS8Maximum workspace containers measured by one heartbeat
HEARTBEAT_WORKSPACE_METRICS_MAX_OUTPUT_BYTES65536Maximum bytes read from each heartbeat Docker metric command
ORIGIN_CA_CERT_VALIDITY_DAYS7Validity for per-node Origin CA certificates signed by the API Worker

New nodes generate /etc/sam/tls/origin-ca-key.pem locally in cloud-init and fetch only the signed certificate from POST /api/nodes/:id/origin-ca-certificate (packages/cloud-init/src/template.ts, apps/api/src/routes/node-lifecycle.ts). Legacy ORIGIN_CA_CERT and ORIGIN_CA_KEY Worker secrets are not required for new node provisioning.

VM workspace admission uses persisted workspaces.resolved_reservation_json snapshots and provider capacity fields in a final single-statement D1 reservation (apps/api/src/services/workspace-placement.ts). Aggregate CPU, memory, and disk reservations plus exclusiveNode control packing. Fresh telemetry remains mandatory on occupied nodes; disk pressure and CPU saturation veto reuse, while live memory percentage is a scoring signal rather than a second capacity gate (apps/api/src/services/workspace-resource-capacity.ts, apps/api/src/durable-objects/task-runner/node-selection.ts). There is no workspace-count or co-tenant cap: MAX_WORKSPACES_PER_NODE and maxCoTenants were removed, and legacy reservation rows that still carry maxCoTenants are admitted on their resource reservations alone. Live memory-threshold settings remain scoring inputs, not gates. VM agents report optional per-workspace memory telemetry in heartbeat metrics when Docker stats can be collected within the configured bounds (packages/vm-agent/internal/server/health.go, packages/vm-agent/internal/sysinfo/docker_metrics.go).

For managed VM nodes, the effective host-memory reserve contract is shared by admission and cloud-init: project scaling overrides win, then TASK_RUN_NODE_HOST_MEMORY_RESERVE_MB, then VM_AGENT_MEMORY_RESERVE_MB, then the 512 MB default. Cloud-init applies the cgroup hierarchy only on newly provisioned nodes. Existing nodes need a drain/recreate, a VM-agent/bootstrap upgrade flow, or a manual in-place systemd/Docker reconfiguration before they can be treated as protected by the workload-slice cap; SAM does not destructively evict existing workspaces to retrofit this.

Applied via cloud-init on each node:

SettingDefaultDescription
SystemMaxUse500MMax disk space for journal
SystemKeepFree1GMinimum free disk to maintain
MaxRetentionSec7dayMax log retention period
StoragepersistentPersist logs across reboots
CompressyesCompress stored entries
VariableDefaultDescription
FILE_UPLOAD_MAX_BYTES52428800 (50 MB)Max size per uploaded file
FILE_UPLOAD_BATCH_MAX_BYTES262144000 (250 MB)Max total size per upload batch
FILE_UPLOAD_TIMEOUT120sUpload timeout (VM agent)
FILE_UPLOAD_TIMEOUT_MS120000 (120s)Upload proxy timeout (Worker)
FILE_DOWNLOAD_TIMEOUT_MS60000 (60s)Download proxy timeout
FILE_DOWNLOAD_MAX_BYTES52428800 (50 MB)Max download file size
VariableDefaultDescription
FILE_PROXY_TIMEOUT_MS15000File proxy request timeout
FILE_PROXY_MAX_RESPONSE_BYTES2097152 (2 MB)Max file proxy response size
FILE_RAW_MAX_SIZE52428800 (50 MB)Max raw binary file size (VM agent)
FILE_RAW_TIMEOUT60sRaw file streaming timeout (VM agent)
FILE_RAW_PROXY_MAX_BYTES52428800 (50 MB)Max raw file proxy size (Worker)
VariableDefaultDescription
REPO_BROWSE_MAX_INLINE_BYTES1000000 (1 MB)Max bytes to inline as text in the file viewer; larger stream raw
REPO_BROWSE_MAX_COMPARE_FILES300Max changed files in an Artifacts diff before truncation
VariableDefaultDescription
MCP_IDEA_CONTEXT_MAX_LENGTH500Max characters of idea context shown to agents
MCP_IDEA_LIST_LIMIT20Default page size for list_ideas
MCP_IDEA_LIST_MAX100Max page size for list_ideas
MCP_IDEA_SEARCH_MAX20Max results from search_ideas
MCP_RELATED_IDEA_SEARCH_LIMIT10Default results from find_related_ideas
MCP_TASK_LIST_LIMIT10Default page size for list_tasks
MCP_TASK_LIST_MAX50Max page size for list_tasks
MCP_TASK_SEARCH_LIMIT10Default page size for search_tasks
MCP_TASK_SEARCH_MAX20Max results from search_tasks
MCP_TASK_DETAIL_RECENT_MESSAGE_LIMIT5Recent assistant messages returned by get_task_details
MCP_TASK_DETAIL_MESSAGE_SNIPPET_LENGTH2000Max characters per assistant message snippet in get_task_details
MCP_MESSAGE_SEARCH_LIMIT10Default page size for search_messages
MCP_MESSAGE_SEARCH_MAX20Max results from search_messages
SEARCH_QUERY_MAX_LENGTH4096Max total UTF-8 bytes retained by idea, task, knowledge, and message search as a DoS guard; responses disclose truncation
SEARCH_QUERY_MAX_TERM_LENGTH48Max LIKE-safe UTF-8 bytes retained per search term; higher values clamp to SQLite’s safe pattern ceiling
SEARCH_QUERY_MAX_TERMS40Max whitespace-delimited terms retained by those search surfaces; higher values clamp to the safe D1 parameter ceiling and responses disclose truncation
MCP_MESSAGE_LIST_LIMIT50Default page size for get_session_messages
MCP_MESSAGE_LIST_MAX200Max messages per get_session_messages request
MCP_ARCHIVED_TOOL_PAYLOAD_LIST_LIMIT10Default page size for get_archived_tool_payloads
MCP_ARCHIVED_TOOL_PAYLOAD_LIST_MAX50Max archived payloads per get_archived_tool_payloads
MCP_COMMENT_LIST_LIMIT10Default page size for list_message_comment_threads
MCP_COMMENT_LIST_MAX25Max threads per list_message_comment_threads request
MCP_COMMENT_BODY_MAX_LENGTH4000Max comment/reply body characters accepted through MCP
MCP_COMMENT_QUOTE_MAX_LENGTH1000Max quoted source-message characters returned to agents
COMMENT_DIRECTIVE_CONTEXT_MAX_LENGTH6000Max send-to-agent comment directive prompt length
MCP_TRIGGER_LIST_LIMIT20Default page size for list_triggers
MCP_TRIGGER_LIST_MAX100Max triggers per list_triggers request
MCP_INCIDENT_LIST_LIMIT10Default page size for private list_incident_queue
MCP_INCIDENT_LIST_MAX50Max private incidents per list_incident_queue request

Project event MCP tools use the ProjectData event limits above: PROJECT_EVENT_LIST_LIMIT, PROJECT_EVENT_LIST_MAX, and PROJECT_EVENT_SUBSCRIPTION_EVENT_CURSOR_MAX_LENGTH.

VariableDefaultDescription
VITE_FILE_PREVIEW_INLINE_MAX_BYTES10485760 (10 MB)Images below this size render inline automatically
VITE_FILE_PREVIEW_LOAD_MAX_BYTES52428800 (50 MB)Images below this size show click-to-load; above shows download link
VITE_ANALYTICS_MAX_QUEUE_SIZE100Max client-side analytics events retained before oldest events drop
VITE_ANALYTICS_FLUSH_THRESHOLD10Client event count that triggers an immediate analytics flush
VITE_ANALYTICS_FLUSH_INTERVAL_MS5000Client analytics background flush interval in milliseconds
VITE_DEBUG_DIAGNOSIS_EVENT_MAX_PAGES100Max paginated diagnosis-event pages loaded per browser request
VITE_CHAT_DELTA_MAX_PAGES50Max newer-message pages one chat refresh drains before failing visibly
VITE_CHAT_TIMELINE_MAX_PAGES200Max pages fetched per loop when the chat timeline drawer opens
VITE_CHAT_LOAD_UNTIL_MAX_PAGES400Max older-message pages chased while resolving a timeline jump
VITE_PROJECT_LIST_LIMIT50Projects loaded into each shared list-cache entry
VITE_PROJECT_POLL_INTERVAL_MS30000Project-list page refresh cadence in milliseconds; 0 disables
VITE_SIDEBAR_PROJECT_POLL_INTERVAL_MS60000App-shell project-list refresh cadence in milliseconds; 0 disables
VITE_WORKSPACE_PORTS_POLL_MS10000Workspace forwarded-port base refresh cadence in milliseconds
VITE_WORKSPACE_PORTS_BACKOFF_MAX_MS120000Maximum backoff between forwarded-port readiness polls
VITE_WORKSPACE_PORTS_FAILURE_BUDGET6Consecutive unavailable port-list responses before circuit cooldown
VITE_WORKSPACE_PORTS_BACKOFF_JITTER_RATIO0.2+/- jitter ratio applied to forwarded-port readiness backoff delays
VITE_WORKSPACE_PORTS_CIRCUIT_RESET_MS300000Open-circuit cooldown before probing forwarded-port readiness again
VITE_SESSION_INFRA_RETRY_DELAYS_MS2000,5000,10000Comma-separated retry delays for chat Details workspace/node fetches
VITE_PROJECT_PREFETCH_DELAY_MS120Mouse dwell before project-detail prefetch; focus/touch are immediate
VITE_BACKGROUND_FETCH_DELAY_MS150Delay before background query activity is shown and announced
VITE_ACP_PERMISSION_POLL_MS2000Pending ACP permission snapshot refresh cadence
VITE_ACP_PERMISSION_RECOVERY_POLL_MS30000Empty-snapshot recovery cadence after a missed realtime attention event
VITE_ACP_PERMISSION_QUERY_RETRY_COUNT3Transient ACP permission snapshot retries; 0 disables retries
VITE_CHUNK_LOAD_RETRY_DELAY_MS350Wait before retrying a failed lazy route-chunk import
VITE_CHUNK_RELOAD_COOLDOWN_MS15000Minimum gap between chunk-recovery reloads; guards against a reload loop
VITE_ROUTE_FALLBACK_REVEAL_DELAY_MS180Delay before the route loading spinner fades in, avoiding a flash
VITE_QUERY_PERSIST_MAX_AGE_MS86400000 (24 h)How long a persisted query-cache record may be restored after writing
VITE_QUERY_PERSIST_THROTTLE_MS1000Minimum gap between IndexedDB writes of the query cache
VITE_QUERY_PERSIST_RESTORE_TIMEOUT_MS250Budget for the initial cache restore before failing open to no cache
VITE_CHAT_TRANSCRIPT_CACHE_TTL_MS86400000 (24 h)How long a chat transcript stays cached after it was last used
VITE_CHAT_TRANSCRIPT_CACHE_MAX_SESSIONS20Most chat transcripts cached at once; older ones are evicted
VITE_CHAT_TRANSCRIPT_PERSIST_MAX_ROWS500Newest rows of each chat transcript written to IndexedDB
VITE_AGENT_CATALOG_STALE_TIME_MS300000Freshness window for the installable agent catalog query
VITE_PROVIDER_CATALOG_STALE_TIME_MS300000Freshness window for provider catalog size/location/price metadata
VITE_TRIAL_STATUS_STALE_TIME_MS60000Freshness window for trial availability status
VITE_CACHED_COMMANDS_STALE_TIME_MS300000Freshness window for cached slash-command registries
VITE_REPORT_ISSUE_CONFIG_STALE_TIME_MS300000Freshness window for the report-issue availability flag
VITE_PROJECT_CREATE_CONFIG_STALE_TIME_MS300000Freshness window for project-creation config flags
VITE_CREDENTIAL_LIMITS_STALE_TIME_MS30000Freshness window for credential usage-limit windows shown in the chat header and Settings → Advanced
VITE_CREDENTIAL_LIMITS_REFETCH_INTERVAL_MS60000Poll cadence for credential usage-limit windows while the tab is visible (paused in the background)

The control-plane UI writes an allowlisted slice of its query cache to IndexedDB so a full page reload paints from cache instead of refetching. Persisted slices are limited to allowlisted project summaries, stripped library indexes, and project-chat session messages. Credentials, admin diagnostics, node and workspace runtime details, file contents, signed URLs, and mutation state are never written to disk.

Records are namespaced by authenticated user and by a schema version, and are deleted on sign-out and on account switch, so one account can never be shown another account’s cached data. If IndexedDB is unavailable — private browsing, a storage quota failure, or a disabled store — the app degrades silently to its normal in-memory cache.

Project chat transcripts are kept for VITE_CHAT_TRANSCRIPT_CACHE_TTL_MS after they were last loaded or updated, capped at the VITE_CHAT_TRANSCRIPT_CACHE_MAX_SESSIONS most recently used. A chat opened inside that window renders from the cache at once and refreshes in the background; a chat that is not cached loads its newest page first, and older history loads as you scroll up. On disk each transcript keeps only its newest VITE_CHAT_TRANSCRIPT_PERSIST_MAX_ROWS rows, however far back it was read, so a chat restored after a reload pages older history back in the same way. The cache is only read after the sign-in check completes.

SAM uses first-party analytics ingestion for operational/product aggregates. Browser events are batched to /api/t; request analytics are written by API middleware when enabled. Analytics is best-effort and disabled paths preserve normal application behavior.

Client page/referrer fields follow a privacy normalization contract before enqueue: query strings, fragments, protocol, host/userinfo for page values, credentials, emails, UUIDs/ULIDs, long opaque tokens, common secret prefixes, repository/code file identifiers, and values after sensitive route markers are removed or replaced with [redacted]. Non-sensitive nested path shape, event names, durations, UTM source/medium/campaign, session ID, visitor/authenticated user ID, and explicit safe entity metadata are preserved for aggregate reporting.

VariableDefaultDescription
ANALYTICS_ENABLEDtrueEnable API middleware analytics; set false to skip request event writes
ANALYTICS_SKIP_ROUTES(built-in skip list)Comma-separated extra route prefixes/patterns excluded from middleware writes
ANALYTICS_DATASET(deployment-generated)Cloudflare Analytics Engine dataset name
ANALYTICS_SQL_API_URLhttps://api.cloudflare.com/client/v4/accountsAnalytics Engine SQL API base URL override
ANALYTICS_DEFAULT_PERIOD_DAYS30Default admin analytics query lookback in days
ANALYTICS_TOP_EVENTS_LIMIT50Max rows returned by top-events admin query
ANALYTICS_GEO_LIMIT50Max countries in geographic distribution view
ANALYTICS_RETENTION_WEEKS12Number of weeks for retention cohort analysis
ANALYTICS_WEBSITE_TRAFFIC_TOP_PAGES_LIMIT20Max top pages/referrers/events in website traffic sections
ANALYTICS_INGEST_ENABLEDtrueEnable browser event ingestion at /api/t; false returns success without writes
RATE_LIMIT_ANALYTICS_INGEST500Analytics ingest requests allowed per IP per hour
MAX_ANALYTICS_INGEST_BATCH_SIZE25Max browser events accepted per ingest request
MAX_ANALYTICS_INGEST_BODY_BYTES65536Max ingest request body size in bytes
MAX_ANALYTICS_DURATION_MS3600000Max accepted page-duration value; larger values are clamped

External analytics forwarding is off by default. When enabled, SAM forwards only analytics rows already accepted by first-party ingestion/middleware; it does not bypass the client-side URL normalization contract.

VariableDefaultDescription
ANALYTICS_FORWARD_ENABLEDfalseEnable external analytics event forwarding
ANALYTICS_FORWARD_EVENTSkey conversion eventsComma-separated list of events to forward
ANALYTICS_FORWARD_LOOKBACK_HOURS25Hours to look back for events
ANALYTICS_FORWARD_CURSOR_KEYanalytics-forward-cursorKV key used to remember forwarded progress
ANALYTICS_FORWARD_SQL_LIMIT10000Max rows fetched per forwarding run
ANALYTICS_SQL_FETCH_TIMEOUT_MS30000Timeout for Analytics Engine SQL fetches
SEGMENT_WRITE_KEY(unset)Segment Write Key for event forwarding
SEGMENT_API_URLhttps://api.segment.io/v1/batchSegment API endpoint
SEGMENT_MAX_BATCH_SIZE100Max events per Segment batch request
GA4_MEASUREMENT_ID(unset)Google Analytics 4 Measurement ID
GA4_API_SECRET(unset)Google Analytics 4 API secret
GA4_API_URLhttps://www.google-analytics.com/mp/collectGA4 Measurement Protocol endpoint
GA4_MAX_BATCH_SIZE25Max events per GA4 batch request

PROJECT_DATA_ARCHIVE_COMPACT_ENABLED=false is the default. Enabling it affects newly journaled terminal-session migrations only. Their format is pinned in D1 and the shard, so retries, reads and copy-back continue with that format after the flag is disabled. Legacy shards are not rewritten by deployment. Keep the existing global-sweep throttle while validating a compact canary.

Optional Worker variableDefaultMeaning
PROJECT_DATA_ARCHIVE_COMPACT_ENABLEDfalseWrite new archives as compact SQLite + compressed R2; existing formats remain readable.
PROJECT_DATA_ARCHIVE_DAILY_WRITE_BUDGET250000Installation-wide daily allowance of ESTIMATE UNITS (not billed rows) for compact migration attempts; 0 pauses admission. The checked-in wrangler.toml ships 2400000. Divide by 1000 + factor x units-per-session for the migrations/day ceiling: at the shipped factor of 2 and a measured ~10,250-unit session that is ~111/day, above the roughly 72 claim opportunities/day the 18-minute sweep cadence offers. Whichever of the two is smaller is what actually runs, so check both before tuning either.
PROJECT_DATA_ARCHIVE_WRITE_ESTIMATE_FACTOR32Safety multiplier applied to the row inventory estimateArchiveWrites already counts (rows, archived tool payloads, grouped rows, and 512-byte units of grouped FTS text, including source deletion) — NOT an independent amplification factor. The checked-in wrangler.toml ships 2. Production measurement 2026-09-14: two project_data_archive_write_budget samples an hour apart give ~65,001 estimated writes per migration, i.e. ~8,000 inventory units at the then-shipped factor of 8; four isolated archive ticks billed 5,950/6,436/7,942/9,823 Durable Object rows. Billed rows per inventory unit therefore spans 0.74-1.23, and 2 exceeds the worst observed ratio by ~63%. Raising it does not make the work safer, only more expensive in budget terms: at 8 the same allowance bought 12 migrations/day while the hourly cadence allowed 24, so half the sweep ticks reclaimed nothing. Sample size is small (2 budget samples, 4 telemetry buckets) — re-measure before changing it.
PROJECT_DATA_ARCHIVE_R2_TIMEOUT_MS10000R2 I/O deadline shared across each compact read/export/seal operation, or one chunk write, in milliseconds.
PROJECT_DATA_ARCHIVE_BUDGET_RECEIPT_RETENTION_MS604800000Unused reservation receipt retention (7 days); clamped to at least one UTC budget day.
PROJECT_DATA_ARCHIVE_BUDGET_RECEIPT_CLEANUP_LIMIT100Maximum expired receipts removed per unused release; zero disables cleanup.

Compact archives keep session metadata, a dedicated derived search projection, consolidated conversation text, grouped FTS and existing tool-archive pointers in SQLite. Original message rows (IDs, timestamps, order, origins and full tool metadata) live in immutable gzip R2 chunks in the private PROJECT_DATA_ARCHIVE_R2 binding. SQL chunk references contain sizes, SHA-256, role counts, time bounds and message IDs. Exact history and tool expansion fetch and verify those chunks. Missing/corrupt objects fail the request; they do not silently produce a truncated history. The original terminal hash and search-projection count/hash coverage must match before publication and source deletion. Already-published archives without current coverage are repaired through durable raw/grouped cursors: compact raw history advances by configured R2 chunks, while legacy raw rows and grouped fallback advance by the configured hash-page size. The resumable projection commitment and FTS counts are verified before coverage becomes complete. This is one-time backfill work and is reported separately from steady-state query coverage. Version 2 recovery manifests include the compressed-object references; legacy version 1 manifests remain unchanged.

Compact chunk exports use the smaller of PROJECT_DATA_ARCHIVE_CHUNK_BYTES and 2 MiB, with an 8 MiB serialized/decompressed object ceiling. A single valid row above the configured chunk target travels alone up to the absolute Durable Object RPC ceiling; it is never truncated or allowed to livelock the cursor. Rows above that platform ceiling still fail closed and remain on the source. R2 objects have no automatic expiry: deleting them destroys archive history and recovery data.

Everything in this section applies only when PROJECT_DATA_ARCHIVE_COMPACT_ENABLED is true: the write-budget reservation and the derived selection ceiling are both gated on it (selectCandidates / processArchiveMigrationBatch in apps/api/src/scheduled/project-data-archive-sharding.ts). Legacy (non-compact) archiving reserves nothing and is bounded only by PROJECT_DATA_ARCHIVE_SWEEP_MESSAGE_BUDGET.

The write allowance is a durable admission estimate, not a hard invoice cap. Each attempt reserves 1,000 writes plus the estimate factor times the sum of raw rows, grouped rows, tool-pointer rows and grouped UTF-8 text bytes rounded up to 512-byte units. The installation-wide D1 reservation is atomic, resets on the next UTC day and is never refunded after an interrupted attempt. Contenders that definitively lose journal creation or lease acquisition before doing archive writes release their reservation with an idempotent D1 receipt. The pool covers this SAM installation; other installations on the same Cloudflare account need separate headroom. Source inventory is rechecked under the transcript lock before creating its intent. New sessions that cannot reserve remain readable on root; existing interrupted migrations retain their existing recovery fence and can retry when allowance is available. Candidate selection is bounded by the SMALLER of PROJECT_DATA_ARCHIVE_SWEEP_MESSAGE_BUDGET and a ceiling derived from the daily allowance itself. The derivation floors twice, in two steps: archiveAffordableWriteUnits() computes floor((allowance - 1000) / factor) — the largest estimate the allowance can ever admit — and archiveAffordableMessageCeiling() then computes floor(units / (1 + PROJECT_DATA_ARCHIVE_SWEEP_UNIT_OVERHEAD_PERCENT / 100)) (both in apps/api/src/project-data-archive/write-budget.ts). Collapsing the two floors into one expression can differ by one at some factor/overhead combinations. Deriving the second half means a selector can never offer a candidate the write budget must refuse on every attempt. Two independently configured ceilings previously had to agree by hand, and when the deployed allowance was lowered without lowering the message budget, largest-first selection re-picked the same unaffordable session on every pass and reclaimed nothing for four days while still reporting success.

The derived ceiling assumes an overhead; it does not measure one. A session whose real tool-payload, grouped-row and FTS-unit cost exceeds the assumption is still refused at reservation time, and the pass then descends to the next-smaller candidate (PROJECT_DATA_ARCHIVE_SWEEP_FALLTHROUGH_DEPTH controls how many spare candidates it reads for that purpose). A refused candidate opens no migrating fence, stays readable on root, and consumes neither a session slot nor any of the cumulative message budget.

The two refusals are distinguished because they need different responses. exceeds_allowance means the session costs more than the entire daily pool and waiting cannot help; window_exhausted means the day’s pool is spent and the next UTC window refills it. After PROJECT_DATA_ARCHIVE_BUDGET_STALL_ALERT_SWEEPS consecutive passes that migrate nothing and see only the former, the sweep cadence row reports partial with an actionable last_error instead of succeeded. Sessions above every ceiling require an explicitly sized migration plan; increasing the daily budget alone does not lift the message cap.

Whole-session deletion and SQLite/FTS index maintenance still consume writes. project_data_archive_sql_usage reports actual cursor writes. project_data_archive_candidate_migrated.sourceFinalization reports source rows, before/after bytes, duration, and reclaimed bytes; copy checkpoints retain the last operation, operation ID, ordinal, byte totals, lease epoch, and timing for reset attribution. These invocation timings are diagnostic only: billed Durable Object duration comes from Cloudflare durableObjectsPeriodicGroups.sum.duration. Compare aligned daily metrics with these events before raising throughput, and separate one-time repair/backfill from steady-state search. At the default 250,000 estimated writes/day, 30 days admits at most 7.5 million estimated migration writes. Normal application traffic, legacy migrations, operator copy-back and other Workers are outside this pool, so operators must reserve account headroom separately. A zero allowance pauses new compact attempts without breaking reads or completed crash-gap publication. Source deletion is still one whole-session operation, not an interruptible per-row spending limit.

After any compact archive is published, rollback must retain compact readers (disable the writer flag to pause new migrations). Deploying a binary from before compact-reader support would read empty raw SQL tables for those sessions; copy them back and verify root ownership before considering such a downgrade.

Certificate issuance retries transport failures, HTTP 429 and HTTP 5xx with capped exponential backoff. Cloud-init also retries its certificate fetch and reports a fixed boot-failure reason when bootstrap cannot continue. Tasks replace a failed fresh VM only after its deletion is confirmed, before any workspace execution. The default is one replacement; a second boot failure ends with its specific reason. Wrong agent versions fail immediately, while a VM without any heartbeat gets six minutes by default.

  • ORIGIN_CA_RETRY_MAX_ATTEMPTS — Maximum upstream certificate attempts including the first; retries transport, 429 and 5xx only (default: 3).
  • ORIGIN_CA_RETRY_BASE_DELAY_MS — Initial upstream certificate retry delay (default: 500).
  • ORIGIN_CA_RETRY_MAX_DELAY_MS — Cap on upstream certificate exponential backoff (default: 2000).
  • ORIGIN_CA_REQUEST_TIMEOUT_MS — Deadline per upstream certificate request, including reading its body (default: 10000).
  • CLOUD_INIT_AGENT_DOWNLOAD_TIMEOUT_SECONDS — Deadline for the agent binary download, including DNS and connection time (default: 60). Failure triggers the best-effort boot-failure callback.
  • CLOUD_INIT_CERT_MAX_ATTEMPTS — Maximum cloud-init CSR POST attempts including the first (default: 3).
  • CLOUD_INIT_CERT_BASE_DELAY_SECONDS — Initial cloud-init certificate retry delay (default: 2).
  • CLOUD_INIT_CERT_MAX_DELAY_SECONDS — Cap on cloud-init certificate exponential backoff (default: 8).
  • CLOUD_INIT_CERT_REQUEST_TIMEOUT_SECONDS — Deadline per cloud-init certificate request and best-effort boot-failure report; must exceed the upstream API retry budget (default: 45).
  • TASK_RUNNER_FIRST_HEARTBEAT_TIMEOUT_MS — Fresh VM first-heartbeat deadline, capped by TASK_RUNNER_AGENT_READY_TIMEOUT_MS (default: 360000).
  • TASK_RUNNER_BOOT_MAX_REPLACEMENTS — Maximum fresh VM boot replacements per task run; 0 disables replacement; workspace execution is never replayed (default: 1).

Set these optional variables in the GitHub deployment Environment; the deployment pipeline forwards them to the Worker. No new secrets are required. Certificate request and retry budgets should remain below the first-heartbeat deadline.

The optional Worker variables CLI_RECEIPT_REQUEST_MAX_BYTES (default 262144) and CLI_RECEIPT_RESPONSE_MAX_BYTES (default 65536) accept positive byte counts. Keyed requests exceeding the request limit are rejected before reservation; replies exceeding the response limit leave the receipt pending for reconciliation. These are runtime configuration overrides, not credentials or required deployment secrets.