Configuration Reference
SAM uses environment variables for platform configuration. User-specific settings (cloud provider tokens, agent API keys) are stored encrypted in the database, not as environment variables.
Platform Secrets
Section titled “Platform Secrets”These are Cloudflare Worker secrets, set during deployment. Pulumi auto-generates security keys on first deploy.
| Secret | Description |
|---|---|
ENCRYPTION_KEY | AES-256-GCM master key. Used for BetterAuth session cookies and user credential encryption unless a purpose-specific override below is set (auto-generated) |
BETTER_AUTH_SECRET | Optional purpose-specific override for BetterAuth session cookie signing/encryption. Falls back to ENCRYPTION_KEY when unset (apps/api/src/lib/secrets.ts) |
CREDENTIAL_ENCRYPTION_KEY | Optional purpose-specific override for AES-GCM encryption of user cloud/agent credentials. Falls back to ENCRYPTION_KEY when unset (apps/api/src/lib/secrets.ts) |
JWT_PRIVATE_KEY | RSA-2048 private key for signing tokens (auto-generated) |
JWT_PUBLIC_KEY | RSA-2048 public key for token verification (exposed via JWKS) |
DEPLOY_SIGNING_PRIVATE_KEY | Ed25519 private key for signing deployment apply payloads (auto-generated) |
DEPLOY_SIGNING_PUBLIC_KEY | Ed25519 public key derived during deployment for deployment node verification (auto-generated) |
VAPID_PRIVATE_KEY | Base64url P-256 private scalar used to authenticate Web Push delivery (auto-generated) |
VAPID_PUBLIC_KEY | Uncompressed base64url P-256 public key returned to browsers at runtime (derived during deployment) |
VAPID_SUBJECT | RFC 8292 contact URI for Web Push, defaulting to the deployment app origin (generated during deployment) |
CF_API_TOKEN | Cloudflare API token for infrastructure, DNS, Origin CA certificate issuance, observability, AI Gateway, Containers, and admin logs. Requires Account → Containers → Edit and Account → SSL and Certificates → Edit. |
CF_AIG_TOKEN | Optional narrower Cloudflare AI Gateway Unified Billing token |
CF_ZONE_ID | Cloudflare zone ID for DNS record management |
CF_ACCOUNT_ID | Cloudflare account ID |
DEVCONTAINER_CACHE_CLOUDFLARE_API_TOKEN | Optional narrower Cloudflare token for managed devcontainer registry credentials |
DEVCONTAINER_CACHE_CLOUDFLARE_ACCOUNT_ID | Optional Cloudflare account override for managed devcontainer registry credentials |
GITHUB_CLIENT_ID | Optional fallback GitHub App client ID for OAuth; runtime admin config takes precedence |
GITHUB_CLIENT_SECRET | Optional fallback GitHub App client secret for OAuth; runtime admin config takes precedence |
GITHUB_APP_ID | Optional fallback GitHub App ID for installation tokens; runtime admin config takes precedence |
GITHUB_APP_PRIVATE_KEY | Optional fallback GitHub App private key (PEM or base64); runtime admin config takes precedence |
GITHUB_APP_SLUG | Optional fallback GitHub App URL slug; runtime admin config takes precedence |
GITHUB_WEBHOOK_SECRET | Optional fallback GitHub App webhook HMAC secret; runtime admin config takes precedence |
GITLAB_HOST | Optional fallback GitLab OAuth host, such as https://gitlab.com; runtime admin config takes precedence |
GITLAB_CLIENT_ID | Optional fallback GitLab OAuth application ID; runtime admin config takes precedence |
GITLAB_CLIENT_SECRET | Optional fallback GitLab OAuth secret; runtime admin config takes precedence |
TRIAL_CLAIM_TOKEN_SECRET | Trial onboarding HMAC secret (auto-generated) |
Worker Variables
Section titled “Worker Variables”Unless the deploy sets them itself (as it does BASE_DOMAIN), Worker variables come from [vars]
in apps/api/wrangler.toml, falling back to the default in the code. On a self-hosted instance, a
GitHub Environment variable of the same name replaces that value at deploy time, but only if the
deploy forwards it: scripts/deploy/sync-wrangler-config.ts must read it, and the Sync Wrangler
Config steps in .github/workflows/deploy-reusable.yml must pass it from vars. After a deploy,
the Worker’s Settings → Variables and Secrets page in the Cloudflare dashboard shows the value it
got. To change any other variable, edit
wrangler.toml in your fork; updates then need a manual merge (see
Updating an Existing Self-Hosted Instance).
| Variable | Default | Description |
|---|---|---|
BASE_DOMAIN | — | Root domain for the deployment (e.g., example.com) |
PREVIEW_BASE_DOMAIN | preview.BASE_DOMAIN | Full isolated hostname used for interactive HTML previews |
PREVIEW_URL_TTL_SECONDS | 300 | Lifetime of project/file/version-scoped interactive preview URLs in seconds |
PREVIEW_SIGNING_KEY | generated | Deployment-owned HMAC key generated and persisted by Pulumi; not a manual prerequisite |
VERSION | — | Deployment version string |
SETUP_TOKEN | — | Plaintext first-run setup token generated during deploy and readable in the Cloudflare dashboard while setup is incomplete |
SETUP_FORCE | (unset) | Set to true to reopen /setup for lockout recovery |
SETUP_RATE_LIMIT_MAX_ATTEMPTS | 10 | Max setup-token attempts per identifier/window |
SETUP_RATE_LIMIT_WINDOW_SECONDS | 900 | Setup-token attempt window in seconds |
D1_SESSION_MODE | first-primary | D1 Sessions API anchor for the Worker fetch handler. first-primary runs each request against one D1 session whose first query goes to the primary and whose later queries may be served by a caught-up read replica — same freshness, one wide-area round trip per request instead of one per query. disabled sends every query straight at the primary. An unrecognised value falls back to first-primary and is logged once per isolate. scheduled() and Durable Objects always use the unsessioned binding. |
PLATFORM_CONFIG_CACHE_MS | 60000 | Per-isolate cache TTL for the resolved platform integration config (the GITHUB_*/GITLAB_*/GOOGLE_LOGIN_* fallbacks above and their runtime admin overrides). Resolving costs 14 D1 queries and runs on the auth preamble of every authenticated request. After a config change, isolates that already hold a cached copy converge within this window. Set to 0 to disable caching and always re-read D1. |
GITHUB_INSTALLATION_TOKEN_CACHE_TTL_SECONDS | 3000 | KV cache TTL for GitHub App installation tokens. The default is shorter than GitHub’s one-hour token lifetime. Set to 0 to disable writes for new cache entries. |
GITHUB_INSTALLATION_TOKEN_REFRESH_MARGIN_SECONDS | 300 | Cached GitHub App installation tokens that expire within this many seconds are minted again instead of reused, so a long cache TTL can never hand out an expiring token. Capped at 1800, half the one-hour token lifetime. |
GITHUB_REPO_ACCESS_CACHE_TTL_SECONDS | 300 | KV cache TTL for per-user, per-installation, per-repository GitHub access checks used by the Files page. Set to 0 to disable writes for new cache entries. |
GITHUB_TREE_CACHE_TTL_SECONDS | 86400 | KV cache TTL for immutable Git tree responses keyed by commit SHA. Branch refs are resolved to a commit SHA before lookup. Set to 0 to disable writes for new cache entries. |
PROJECT_MULTIPLAYER_CACHE_TTL_MS | 10000 | Per-isolate cache TTL for project multiplayer state counts used by trigger-bearing pages. Set to 0 to disable the cache. |
CREDENTIAL_ATTRIBUTION_CACHE_TTL_MS | 10000 | Per-isolate cache TTL for project credential attribution health used by trigger-bearing pages. Set to 0 to disable the cache. |
GitHub Environment Variables
Section titled “GitHub Environment Variables”Set in GitHub Settings → Environments → production:
| Variable | Description | Example |
|---|---|---|
BASE_DOMAIN | Deployment domain | example.com |
RESOURCE_PREFIX | Domain-derived Cloudflare resource name prefix | sa379a6 |
PULUMI_STATE_BUCKET | R2 bucket for Pulumi state | sa379a6-pulumi-state |
CF_CONTAINER_ENABLED | Optional instant-session runtime toggle. Generated deploys default to true; set false to force VM runtime. | false |
WORKER_SECRET_BULK_MAX_OPS | Optional deploy-script limit for queued Worker secret create/update/delete operations in one wrangler secret bulk payload. Defaults to 100; range 1–100. | 75 |
D1_READ_REPLICATION_MODE | Optional D1 read-replication mode applied on every deploy. Defaults to auto; set disabled to remove replicas. | disabled |
D1_SESSION_MODE | Optional Worker D1 Sessions anchor. Defaults to first-primary; set disabled to send every query to the primary. An unrecognised value falls back to the default and is logged once per isolate. | disabled |
D1_RESTORE_RECOVERY_WINDOW_DAYS | Optional D1 restore window for accounts with narrower retention. Defaults to 30; range 1–30. | 7 |
D1_MIGRATION_CHURNING_TABLES | Optional comma-separated <binding>.<table> subset of the reviewed retention/expiry table list. May narrow the built-in list but cannot expand it. | OBSERVABILITY_DATABASE.platform_errors |
D1_MIGRATION_CHURNING_TABLE_MAX_DECREASE_PERCENT | Maximum allowed decrease for reviewed churning tables. Defaults to 50; range 0–100. A decrease exactly at the limit is accepted. | 25 |
The reviewed default churning selectors are DATABASE.deployment_releases, DATABASE.github_webhook_deliveries, DATABASE.project_files, DATABASE.registry_credential_rate_limits, DATABASE.session_snapshots, DATABASE.sessions, DATABASE.trial_waitlist, DATABASE.trigger_executions, DATABASE.verifications, DATABASE.webhook_deliveries, and OBSERVABILITY_DATABASE.platform_errors. All other application tables retain zero row-decrease tolerance. Leave D1_MIGRATION_CHURNING_TABLES unset to use the complete reviewed default list.
RESOURCE_PREFIX is generated from BASE_DOMAIN as s plus the first six hex
characters of the domain’s SHA-256 hash. The self-host onboarding flow fills it
in for you.
App deployment image-resolution safety
Section titled “App deployment image-resolution safety”These optional Worker variables bound the server-side OCI registry lookups used when a deployment release submits tag-based images. Digest-pinned images are stored without registry network resolution.
| Variable | Default | Description |
|---|---|---|
DEPLOYMENT_IMAGE_RESOLVE_REQUEST_TIMEOUT_MS | 10000 | Per-registry request timeout |
DEPLOYMENT_IMAGE_RESOLVE_TOTAL_TIMEOUT_MS | 60000 | Total tag-resolution wall-clock budget per release submission |
DEPLOYMENT_IMAGE_RESOLVE_MAX_FETCH_ATTEMPTS | 200 | Maximum outbound registry/token fetches per resolver instance |
DEPLOYMENT_IMAGE_RESOLVE_MAX_REDIRECTS | 2 | Maximum manually validated HTTPS redirects per outbound request |
DEPLOYMENT_IMAGE_RESOLVE_TOKEN_RESPONSE_MAX_BYTES | 65536 | Maximum bearer-token JSON response size |
DEPLOYMENT_IMAGE_RESOLVE_MAX_CONCURRENT_FETCHES | 4 | Maximum simultaneous outbound resolver fetches |
DEPLOYMENT_IMAGE_RESOLVE_MAX_SERVICES | 50 | Maximum tag-based image references resolved per release submission |
Required GitHub Actions secrets include CF_API_TOKEN, CF_ACCOUNT_ID, CF_ZONE_ID, R2_ACCESS_KEY_ID, R2_SECRET_ACCESS_KEY, and PULUMI_CONFIG_PASSPHRASE. GitHub App/OAuth secrets (GH_CLIENT_ID, GH_CLIENT_SECRET, GH_APP_ID, GH_APP_PRIVATE_KEY, GH_APP_SLUG, GH_WEBHOOK_SECRET) and Google login OAuth secrets (GOOGLE_LOGIN_CLIENT_ID, GOOGLE_LOGIN_CLIENT_SECRET) are optional environment fallbacks; fresh deployments can set them through /setup instead. The separate Google infra/GCP OAuth pair (GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET) is used only for WIF and can be configured by a superadmin at /admin/integrations; runtime values override the environment fallback. Service-account JSON users need no infrastructure OAuth client. Deploy signing keys are generated and persisted by Pulumi during deployment; GitHub Environment values are only needed for explicit key overrides.
CLI Environment Variables
Section titled “CLI Environment Variables”These variables affect the local sam CLI process only. They are not Worker runtime variables or GitHub Actions secrets.
| Variable | Default | Description |
|---|---|---|
SAM_CLI_MAX_API_RESPONSE_BYTES | 1048576 | Maximum API response body bytes the CLI reads before truncating/aborting. |
Feature Flags
Section titled “Feature Flags”Codex and Claude Code guided subscription login have no feature-on environment variable. They are
available by default when the deployment includes the SANDBOX,
CREDENTIAL_SETUP_SESSION, and SETUP_SESSION_POOL Worker bindings generated by
SAM’s deployment configuration. Omitting one of those bindings disables the
guided flow. SANDBOX_ENABLED continues to control separate administrative
Sandbox runtime surfaces and is not required for guided login.
| Variable | Default | Description |
|---|---|---|
MAX_CONCURRENT_SETUP_SESSIONS | 2 | Maximum concurrent guided credential-setup sessions. |
SETUP_SESSION_TTL_MS | 900000 | Guided session lifetime before automatic teardown. |
SETUP_SESSION_CAPTURE_POLL_MS | 3000 | Interval for checking device-login and credential-capture state. |
CODEX_DEVICE_AUTH_REQUEST_TIMEOUT_MS | 30000 | Timeout for each Codex app-server JSON-RPC request. |
CLAUDE_SETUP_ENTER_DELAY_MS | 1000 | Delay before sending Enter as a separate stdin write after pasting Claude’s browser-displayed code. |
CLAUDE_SETUP_EXCHANGE_TIMEOUT_MS | 120000 | Maximum wait for Claude’s CLI code exchange before a visible timeout. |
CLAUDE_SETUP_REJECTION_SETTLE_MS | 400 | Wait for Claude CLI Ink redraws to settle before classifying an OAuth error. |
CLAUDE_SETUP_VERIFICATION_POLL_MS | 500 | Interval for checking the sandbox handoff file for Claude’s browser-displayed code. |
CLAUDE_SETUP_TTY_COLUMNS | 512 | PTY width for claude setup-token, reducing opaque-token wrapping. |
CLAUDE_SETUP_OUTPUT_BUFFER_BYTES | 32768 | Maximum in-memory Claude PTY output retained for parsing. |
CLAUDE_VERIFICATION_CODE_MAX_LENGTH | 1024 | Maximum accepted length of Claude’s browser-displayed code#state value. |
CLAUDE_SETUP_ERROR_DETAIL_MAX_LENGTH | 160 | Maximum sanitized Claude CLI diagnostic length shown to the user. |
CLAUDE_OAUTH_TOKEN_MAX_LENGTH | 8192 | Maximum captured Claude OAuth token length. |
SETUP_SESSION_SWEEP_MAX_CANDIDATES | 50 | Maximum expired sessions cleaned up by one scheduled sweep. |
POOL_LEASE_BUFFER_MS | 300000 | Grace period after session TTL before a leaked capacity lease self-prunes. |
The variables below tune the Instant (Cloudflare Container) runtime — how long a session stays awake, how long a wake may take, and how many snapshot restores are attempted before a session is failed. See Instant Sessions for what each of these means to a user.
| Variable | Default | Description |
|---|---|---|
CF_CONTAINER_ENABLED | true | Enables Cloudflare Container instant sessions for matching profiles and zero-config runtime selection. Set false to force cloud VM runtime. |
CF_CONTAINER_SLEEP_AFTER | 1h | Normal inactivity window before an Instant container sleeps. Sleep remains recoverable through the runtime-neutral session snapshot. |
CF_CONTAINER_ACTIVE_WORK_MAX_MS | 7200000 | Defensive maximum lifetime for an active-work keepalive lease. |
CF_CONTAINER_KEEPALIVE_RENEW_INTERVAL_MS | 300000 | Interval used to renew the container activity timeout while prompt work is active. |
CF_CONTAINER_WAKE_TIMEOUT_MS | 120000 | Maximum time for a sleeping container to launch, restore its snapshot, and accept the triggering request. |
CF_CONTAINER_RECOVERY_MAX_ATTEMPTS | 2 | Maximum snapshot restore attempts before SAM reconciles the runtime, workspace, agent session, and active task to a visible terminal recovery failure. |
INSTANT_STALE_CALLBACK_MARGIN_MS | 60000 (60 sec) | Freshness margin used to reject destructive (error/failed) callbacks arriving from a superseded Instant container generation after the runtime row was reconciled by a completed recovery. |
CF_CONTAINER_CREATE_WORKSPACE_TIMEOUT_MS | 120000 | Budget for the synchronous instant-session create-workspace request, which includes the repository clone inside the container. |
CF_CONTAINER_CLONE_FILTER | blob:none | Git partial-clone filter forwarded to instant containers as STANDALONE_CLONE_FILTER. Set off to force full clones. |
CF_CONTAINER_RECOVERY_MAX_ATTEMPTS has a deployment-safety minimum of 2; smaller positive values resolve to 2.
Persistent session snapshots and sleep
Section titled “Persistent session snapshots and sleep”Sleeping and reclaimed Instant and VM sessions are restored from a snapshot of the agent’s home directory and the repository work in progress. An ordinary sleep requires a complete snapshot before SAM tears down VM compute. When snapshots keep failing, a bounded fallback can instead sleep an idle VM session on the Git recovery point an earlier snapshot saved, keeping the transcript but not every file (SAM could not save a complete snapshot). None of these limits are surfaced in the UI, so operators should set expectations deliberately — see What gets restored.
| Variable | Default | Description |
|---|---|---|
SESSION_SNAPSHOT_TTL_DAYS | 7 | Snapshot retention. A session sleeping longer than this cannot be fully restored. |
SESSION_SNAPSHOT_TOTAL_BUDGET_BYTES | 268435456 (256 MiB) | Max combined size of the home + work-in-progress snapshot. The higher default favors bounded retained R2 state over keeping a VM alive when a typical agent harness has accumulated substantial durable state. |
SESSION_SNAPSHOT_ENTRY_THRESHOLD_BYTES | 268435456 (256 MiB) | Largest single file the snapshot scanner will include. This matches the total budget so durable agent state databases are not skipped solely because they are larger than the former 50 MiB cap. |
SESSION_SNAPSHOT_TRANSFER_IDLE_TIMEOUT_MS | 30000 (30 sec) | No-progress timeout for each snapshot upload or download. |
SESSION_SNAPSHOT_UPLOAD_URL_TTL_SECONDS | 900 (15 min) | Lifetime of direct R2 upload URLs used so large snapshots do not traverse the Worker request-body boundary. Current agents bind exact length and SHA-256; busy legacy VM agents stream through a current same-user VM relay that independently authenticates both nodes and removes callback credentials before R2. When R2 S3 credentials are unavailable, SAM retains the Worker upload path. |
SESSION_SNAPSHOT_REQUEST_TIMEOUT_MS | 300000 (5 min) | Budget for the vm-agent to accept the final checkpoint request. Durable completion is governed by progress reporting rather than this fixed wall clock. |
SESSION_SNAPSHOT_PROGRESS_IDLE_TIMEOUT_MS | 120000 (2 min) | No-progress watchdog for an accepted final checkpoint. Current vm-agents periodically advance D1 progress while walking HOME or uploading artifacts; if progress stops, SAM records a degraded snapshot and that sleep attempt fails, counting against the sleep failure budget below. |
SESSION_SNAPSHOT_POLL_INTERVAL_MS | 1000 (1 sec) | Interval used while the Worker waits for a VM agent’s asynchronous final checkpoint to commit in D1. |
SESSION_SNAPSHOT_OPERATION_TIMEOUT | 15m | VM-agent checkpoint/restore deadline, using Go duration syntax. Before the first restore RPC, TaskRunner pins this duration plus SESSION_SNAPSHOT_REQUEST_TIMEOUT_MS as its retry window; retries and restarts cannot renew it. |
SESSION_SNAPSHOT_PROGRESS_REPORT_INTERVAL | 15s | VM-agent throttle for best-effort progress callbacks during data-scaled snapshot work. This uses Go duration syntax and is passed to newly provisioned VMs and Instant containers. |
SESSION_SNAPSHOT_PROGRESS_REPORT_TIMEOUT | 5s | VM-agent timeout for each best-effort snapshot progress callback. This uses Go duration syntax and is passed to newly provisioned VMs and Instant containers. |
SESSION_SNAPSHOT_JSON_BODY_MAX_BYTES | 262144 (256 KB) | Maximum snapshot coordination request size accepted by the Worker. |
SESSION_SNAPSHOT_R2_PREFIX | session-snapshots | Private object prefix. Session objects are deleted by the Worker from D1 lifecycle state, not by object age. |
SESSION_SNAPSHOT_RECOVERY_MAX_ATTEMPTS | 3 | Replacement-VM wake attempts allowed in a burst. Once spent, further wakes are refused (recovery_attempts_exhausted) until SESSION_SNAPSHOT_RECOVERY_ATTEMPT_DECAY_MS has passed since the last failed attempt; the snapshot’s own expiry remains the hard limit. |
SESSION_SNAPSHOT_RECOVERY_ATTEMPT_DECAY_MS | 900000 (15 min) | How long a spent wake-attempt burst stays spent before a new wake may be tried |
SESSION_RECOVERY_LINEAGE_MAX_DEPTH | 256 | How many wake-to-wake links a wake follows back to the conversation’s first run to decide whether that run explicitly pinned its region. A wake without such a request only prefers the region it slept in. |
SESSION_SLEEP_AFTER_MS | 900000 (15 min) | ProjectData-recorded idle interval before SAM automatically sleeps a VM session. Runtime heartbeats do not extend this clock. Completed and failed tasks queue sleep immediately. Their still-active final prompt becomes eligible after this interval from its later activity or the task’s end, so a working agent is not slept mid-turn. Ledger cleanup also uses it to protect the final response. |
SESSION_SLEEP_SWEEP_BATCH_SIZE | 10 | Maximum due session sleep candidates selected and individually claimed by one scheduled sweep. |
SESSION_SLEEP_SWEEP_WALL_BUDGET_MS | 20000 (20 sec) | Soft wall-clock budget for bounded D1/ProjectData eligibility and claim work. After a durable claim, final snapshot and teardown run through the scheduled event’s out-of-band lifetime. Remaining unclaimed rows stay due for the next sweep. |
SESSION_SLEEP_RETRY_DELAY_MS | 300000 (5 min) | Retry delay after a fail-closed automatic sleep attempt. |
SESSION_SLEEP_FAILURE_MAX_ATTEMPTS | 3 | Failed full-snapshot sleep attempts in one episode before SAM tries the fallback: release an idle VM session’s compute, keeping its transcript and the exact Git commit, branch and uncommitted changes an earlier snapshot saved. Instant sessions, and VM sessions without such a recovery point, end the episode blocked instead. An episode ends when the session sleeps or wakes, or a person sends a message. |
SESSION_SLEEP_FAILURE_MAX_ELAPSED_MS | 900000 (15 min) | Time since an episode’s first sleep attempt after which the fallback is tried whatever the attempt count, once at least one attempt has failed. With the default sweep and retry delay this is about three attempts. |
SESSION_SLEEP_MAX_ATTEMPTS | 9 | Ceiling on failed attempts in one sleep episode, counting full snapshots and fallback attempts. It is raised to at least SESSION_SLEEP_FAILURE_MAX_ATTEMPTS + 1 so the fallback always gets a try. At the ceiling the episode ends blocked: automatic sleep stops and the chat says so. A task failure starts a fresh episode; a failed task whose episode ends blocked has its runtime torn down, with a chat notice. Raising this value re-arms rows exhausted before bounded episodes existed that are still below the new limit. |
FAILED_TASK_PRESERVATION_MAX_WAIT_MS | 28800000 (8 hours) | Longest a failed task’s runtime stays awake waiting for its preservation sleep, measured from the latest of the failure, an in-place wake and the start of the agent’s current turn. A turn that never ends defers that sleep without spending an attempt; past this wait the sweep tears the runtime down and says so in the chat (releaseStalledFailedTaskPreservation()). |
SESSION_SLEEP_CLAIM_LEASE_MS | 600000 (10 min) | Time after which an interrupted automatic-sleep claim can be safely reclaimed. |
HARNESS_BACKGROUND_WORK_LEASE_MS | 300000 (5 min) | Finite sleep-protection lease renewed by normalized harness background-work lifecycle signals. Expiry fails open to ordinary idle-sleep eligibility so a missing terminal signal cannot pin compute forever. |
HARNESS_BACKGROUND_WORK_MAX_DURATION_MS | 1800000 (30 min) | Absolute ceiling, measured from the last harness lifecycle progress edge rather than the last heartbeat, on how long background work may defer sleep. The sliding lease above is refreshed by periodic re-reports, so an adapter faithfully re-reporting a stale task set (for example an abandoned run_in_background dev server) would otherwise pin compute awake indefinitely. |
ACP_ACTIVITY_ADMISSION_ENABLED | true | Enables Worker-side admission control for ACP activity callbacks. Redundant intermediate state is coalesced to protect ProjectData load; terminal/error and final idle transitions still bypass coalescing. |
ACP_ACTIVITY_COALESCE_WINDOW_MS | 2000 (2 sec) | Minimum interval between redundant intermediate ProjectData activity writes, and the first retry delay for a coalesced report whose flush hit a retryable ProjectData failure (including a CPU-limit reset or lost connection); each further retry doubles the delay, capped at ACP_ACTIVITY_COALESCE_TTL_MS. Activity transitions, new prompt epochs, and terminal/error reports bypass this window. |
ACP_ACTIVITY_COALESCE_TTL_MS | 60000 (1 min) | Maximum lifetime for a pending coalesced activity report before it is evicted and left to probe-backed session-activity reconciliation. |
ACP_ACTIVITY_COALESCE_MAX_PENDING | 512 | Maximum pending coalesced activity reports retained by one Worker isolate. Capacity evictions are logged as activity telemetry rather than silently dropped. |
ACP_ACTIVITY_BINDING_CACHE_TTL_MS | 30000 (30 sec) | Short-lived cache for already-authorized ACP session bindings used to avoid ProjectData reads during callback storms. Callback JWT authorization and D1 node/workspace liveness checks still run per request. |
ACP_ACTIVITY_BINDING_CACHE_MAX_ENTRIES | 2048 | Maximum cached ACP activity bindings retained by one Worker isolate. |
SESSION_SNAPSHOT_RECOVERY_CLAIM_LEASE_MS | 600000 (10 min) | Time after which an interrupted replacement-runtime wake claim can be reconciled or reclaimed. |
SESSION_LIFECYCLE_ERROR_MAX_LENGTH | 2048 | Maximum session lifecycle and agent activity failure diagnostic detail stored in lifecycle records. |
SESSION_SNAPSHOT_PURGE_ENABLED | true | Enables bounded expiry cleanup: terminalizes the sleeping chat, deletes its R2 objects, then removes D1 metadata. Chats slept by the fallback expire on the same seven-day schedule. |
SESSION_SNAPSHOT_PURGE_BATCH_SIZE | 250 | Maximum expired snapshot rows deleted per daily purge. |
SESSION_SLEEP_IN_FLIGHT_MAX_AGE_MS | 1800000 (30 minutes) | Absolute ceiling for preserving in-flight sleep lifecycle rows (scheduled, preparing, stopping, retry-eligible failed) from terminal session destroyers. Rows older than this are treated as wedged and must escape through bounded repair by runSessionSleepLifecycleRepair() in apps/api/src/scheduled/session-sleep-lifecycle-repair.ts instead of deferring forever. |
SESSION_SLEEP_IN_FLIGHT_REPAIR_BATCH_SIZE | 25 | Maximum stale post-capture in-flight sleep rows repaired per scheduled sweep by runSessionSleepLifecycleRepair() in apps/api/src/scheduled/session-sleep-lifecycle-repair.ts. The repair only completes restorable, unexpired preparing/stopping rows as sleeping; it does not wake or replay work. Values above 100 are capped. |
TERMINAL_SESSION_RECONCILE_PROJECT_BATCH_SIZE | 25 | Maximum projects inspected for stale active ProjectData session ledgers per scheduled sweep. Values above 200 are capped. |
TERMINAL_SESSION_RECONCILE_BATCH_SIZE | 25 | Maximum active ProjectData chat_sessions candidates reconciled per project per scheduled sweep. Values above 200 are capped. |
TERMINAL_SESSION_SUMMARY_RECONCILE_BATCH_SIZE | 25 | Maximum active D1 session_summaries candidates reconciled globally per scheduled sweep. Values above 200 are capped. |
TERMINAL_SESSION_RECONCILE_DEFER_MS | 3600000 (1 hour) | Retry delay for live-head, snapshot-protected, or temporarily ineligible terminal-session ledger candidates. Values above 86400000 (24 hours) are capped. |
TERMINAL_NODE_LIFECYCLE_REPAIR_BATCH_SIZE | 25 | Maximum active-looking workspace rows on terminal/deleted nodes repaired per scheduled sweep by runTerminalNodeLifecycleRepair() in apps/api/src/scheduled/terminal-node-lifecycle-repair.ts. The repair marks non-sleeping workspaces stopped, closes non-terminal agent sessions and open compute usage, and routes ProjectData cleanup through the sleeping-snapshot guard. Values above 100 are capped. |
TERMINAL_NODE_LIFECYCLE_REPAIR_WALL_BUDGET_MS | 10000 (10 seconds) | Wall-clock budget for runTerminalNodeLifecycleRepair() in apps/api/src/scheduled/terminal-node-lifecycle-repair.ts inside the scheduled sweep. Values above 30000 are capped so this repair cannot monopolize the cron event. |
REQUIRE_APPROVAL | (unset) | Default signup approval gate. Superadmins can override it at runtime in Admin → Users without redeploying; when no runtime override exists, this value is used. The first genuine human becomes superadmin regardless of this flag — see First Login & Admin Access. |
TRIAL_ANONYMOUS_USER_ID | system_anonymous_trials | Id of the internal anonymous-trial sentinel user, excluded from first-user superadmin checks. Override only if your deployment uses a different sentinel id. |
CAPACITY_POOL_BACKFILL_SCOPE_BATCH_SIZE | 25 | Maximum user scopes and maximum project scopes reconciled by one unscoped capacity-pool backfill pass. Values above 200 are capped; rerun the backfill to continue. |
CAPACITY_POOL_SCHEDULED_RECONCILIATION_INTERVAL_MS | 86400000 (24 hours) | Minimum interval between scheduled capacity-pool reconciliation runs. Set to 0 to let every operational cron sweep reconcile; explicit UI/API reconcile requests are not throttled by this setting. |
CAPACITY_POOL_LEGACY_WORKLOAD_MAPPING_JSON | built-in slices | Environment fallback for the versioned legacy small/medium/large to workload requirements adapter. Persisted platform_settings.capacityPools.legacyWorkloadMapping.v1 wins when present. Values are workload slices, not old whole-VM shapes. |
CAPACITY_POOL_PLATFORM_DEFAULTS_JSON | built-in defaults | Environment fallback for platform resource requirement defaults used when no task/trigger/skill/profile/project/user layer sets a field. Persisted platform_settings.capacityPools.platformDefaults.v1 wins when present. Values are validated by the shared ResourceRequirements validator and must provide every field. |
CAPACITY_POOL_SELECTION_SETTINGS_JSON | built-in scoring weights | Environment fallback for default capacity-pool selection weights and ranking rollout. Persisted platform_settings.capacityPools.selectionSettings.v1 wins when present. Candidate priority remains explicit pool policy; price comparisons are normalized by unit and currency, with unknown price sorted after known comparable prices. |
ORIGIN_CA_CERT_VALIDITY_DAYS | 7 | Validity for per-node Cloudflare Origin CA certificates issued from node-generated CSRs. Must be one of Cloudflare’s supported values: 7, 30, 90, 365, 730, 1095, or 5475. |
The rolloutCohortPercent field in capacity-pool selection settings accepts 0–100
(default 100). resolvePlacementRollout in services/placement-rollout.ts assigns
stable user/pool cohorts. Enabled cohorts use the pool’s configured ranking;
other cohorts use native balanced ranking while the reuse selector records
which host the configured strategy would select and why the selections differ.
The same eligible host set feeds both comparisons. Pool precedence, membership,
credential generation, workload role, aggregate reservations, and paid allocation
fences remain enforced at every percentage. Reducing rollout changes ranking;
it never restores legacy size labels as allocation authority. Settings and plan
columns remain additive, and readers accept plans without rollout diagnostics.
For the user-facing explanation of what these settings control — pool scopes and precedence, allowed offerings, strategies, exhaustion policies, and resource requirements — see the Compute Pools guide.
Upgrading existing compute pools
Section titled “Upgrading existing compute pools”Deploy the normal additive migrations before starting the updated Worker. Existing tasks, workspaces, credentials, and recorded hardware remain in place. Background reconciliation creates missing default pools from existing credentials and resumes in bounded batches; opening the settings page is not required. Larger installations may need several scheduled passes before every scope is ready.
In project Infrastructure settings, inspect the effective default pool before starting new work. A project default takes precedence over a personal default, which takes precedence over installation capacity. These states need different responses:
| Pool state | What to do |
|---|---|
| Migration pending | Allow reconciliation to finish; if it persists, check scheduled reconciliation errors and the affected credential. |
| Configured empty | Select a supported offering in that pool. An empty configured pool intentionally blocks new allocation. |
| Source disabled | Re-enable or replace the pool’s credential source. |
| Catalog unavailable | Check provider access and retry after inventory refresh. A failed refresh preserves the last valid inventory. |
| Configured ready | Start a small test workload and check its requested resources and provider-native hardware in the node details. |
An administrator can verify completion in D1 by inspecting capacity_pools:
migration_state must be complete for the affected pool. The durable user and
project backfill cursors are stored in platform_settings under
capacityPools.backfill.userCursor.v1 and capacityPools.backfill.projectCursor.v1.
Their presence indicates resumable progress, not an error. Do not delete pools or
reset cursors to resolve an unavailable credential.
Existing nodes without verified pool and provider identity may finish their current work but are not automatically treated as eligible pool capacity. New work must pass the current pool, credential, and resource checks. Previously recorded hardware remains visible even if an offering is later removed. Old browser, API, CLI, and MCP size fields remain accepted as compatibility inputs; new resource fields take precedence at the same configuration layer. Saved reservations survive retry rather than adopting changed defaults.
For a ranking rollback, reduce rolloutCohortPercent in the effective selection
settings. This uses balanced ranking for the excluded cohort while retaining
pool authorization and capacity checks. It does not roll back migrations, revive
removed offerings, or permit reuse of unverified nodes. Keep the additive schema
and saved plans; do not drop columns or recreate tables as a rollback step.
Activity coalescing and binding caches are per Worker isolate, so burst reduction scales with the number of active isolates for the same session. Delayed flushes carry their original observed event time, and ProjectData rejects stale writes so a delayed intermediate report cannot overwrite a newer idle/error state from another isolate.
Project file library cleanup
Section titled “Project file library cleanup”| Variable | Default | Description |
|---|---|---|
LIBRARY_PROJECT_DELETE_CLEANUP_BATCH_SIZE | 1000 | Maximum project-owned library objects listed and deleted per R2 page after project deletion. Values above R2’s 1,000-object page maximum are capped. |
Deployment release and compose artifact retention
Section titled “Deployment release and compose artifact retention”The scheduled Worker first reconciles provably stale non-terminal compose releases, then
prunes terminal deployment releases outside the protected window
(apps/api/src/scheduled/d1-retention.ts:runDeploymentReleaseRetention()). Terminal
retention always retains the newest releases per environment and the version reported in
deployment_environments.observed_applied_seq. The stale reconciler only marks a
created/applying compose-artifact release failed when D1 shows old release status
activity, stable authenticated deployment-node observed state, no recent release
fetch/apply events, a valid manifest, and a release version that is not the observed
applied version. Unknown statuses, malformed manifests, missing observed state, active
applying observations, and recent release events fail closed. Compose artifact cleanup
then re-derives references from the remaining manifests
(apps/api/src/scheduled/compose-image-artifact-cleanup.ts:runComposeImageArtifactCleanup()).
| Variable | Default | Description |
|---|---|---|
DEPLOYMENT_RELEASE_RETENTION_ENABLED | true | Enables bounded terminal release pruning. |
DEPLOYMENT_RELEASE_RETENTION_COUNT | 3 | Newest releases protected per environment, in addition to observed-applied and non-terminal releases. |
DEPLOYMENT_RELEASE_RETENTION_BATCH_SIZE | 250 | Maximum release rows deleted per run. |
DEPLOYMENT_RELEASE_RETENTION_INTERVAL_HOURS | 24 | Minimum interval between release retention runs. |
DEPLOYMENT_RELEASE_RETENTION_LAST_RUN_KV_KEY | cleanup:deployment-releases:last-run | KV interval marker. |
DEPLOYMENT_RELEASE_RECONCILIATION_ENABLED | true | Enables stale non-terminal compose release reconciliation before terminal retention. |
DEPLOYMENT_RELEASE_RECONCILIATION_BATCH_SIZE | 50 | Maximum stale non-terminal releases marked failed per retention run. |
DEPLOYMENT_RELEASE_RECONCILIATION_STALE_HOURS | 168 | Minimum release status age before reconciliation can terminalize a stale non-terminal release. |
DEPLOYMENT_RELEASE_RECONCILIATION_ACTIVITY_GRACE_HOURS | 6 | Recent release-event window that protects active fetch/apply work from reconciliation. |
COMPOSE_IMAGE_ARTIFACT_CLEANUP_BATCH_SIZE | 250 | Maximum abandoned compose archives deleted per daily run. |
R2 object lifecycle retention
Section titled “R2 object lifecycle retention”Pulumi updates the existing assets bucket lifecycle resource on upgrades and creates
the same rules on clean installs (infra/resources/storage.ts:r2BucketLifecycle).
temp-uploads/ is transient browser-upload staging; tts/ is a regenerable audio
cache. Durable library/ content is deleted only with its project, and reachable
compose-image-artifacts/ are governed by deployment release retention. Archived
ProjectData tool payloads stay private and retrievable through Worker/MCP access
paths while message rows retain their text in the Durable Object, so these durable
prefixes do not have age-only lifecycle rules.
| Pulumi option | Default | Object prefix | Description |
|---|---|---|---|
sessionSnapshotTtlDays | 7 | session-snapshots/ | Worker-owned retention from actual sleep; no age-only R2 lifecycle |
diagnosticIncidentTtlDays | 7 | configured private | Private diagnostic artifact retention |
tempUploadTtlDays | 1 | temp-uploads/ | Abandoned presigned browser upload retention |
ttsTtlDays | 30 | tts/ | Regenerable TTS audio-cache retention |
| n/a | n/a | project-data/tool-payloads/ | Private ProjectData archive; Worker-owned retention only |
| n/a | n/a | resource-history/ | Private workspace resource chunks; Worker-owned retention only |
All TTL options must be positive integers. Set overrides with pulumi config set
against the target stack before running its deployment workflow.
Google OAuth and GCP provisioning
Section titled “Google OAuth and GCP provisioning”Google login and Google infrastructure authorization are independent credential families:
| Variables | Purpose | Runtime precedence | Redirect URIs |
|---|---|---|---|
GOOGLE_LOGIN_CLIENT_ID, GOOGLE_LOGIN_CLIENT_SECRET | BetterAuth user login | /setup or superadmin runtime D1 → Worker env → unset | /api/auth/callback/google |
GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET | Keyless GCP/WIF setup only | Superadmin runtime D1 → Worker env → unset | /auth/google/callback and /api/deployment/gcp/callback |
Configuring one family never enables or modifies the other. Users who choose service-account JSON do not need either infrastructure OAuth variable.
| Variable | Default | Description |
|---|---|---|
GCP_SERVICE_ACCOUNT_JSON_MAX_BYTES | 65536 | Maximum UTF-8 byte size accepted by PUT /api/gcp/service-account |
GCP_DEFAULT_ZONE | us-central1-a | Default Compute zone |
GCP_IMAGE_FAMILY | ubuntu-2404-lts-amd64 | Compute image family. Native image overrides may be a family name or a Compute Engine image/family reference. |
GCP_IMAGE_PROJECT | ubuntu-os-cloud | Compute image project |
GCP_DISK_SIZE_GB | 50 | Default boot disk size for GCP legacy callers and native requests without bootDiskSizeGb. A native VM request with bootDiskSizeGb overrides this value before the Compute Engine insert call. |
GCP_TOKEN_CACHE_TTL_SECONDS | 3300 | Maximum derivative access-token cache TTL; actual TTL is capped by Google’s returned expiry |
GCP_IDENTITY_TOKEN_EXPIRY_SECONDS | 600 | SAM identity-token lifetime for WIF |
GCP_OPERATION_POLL_TIMEOUT_MS | 300000 | Maximum wait for GCP asynchronous operations |
GCP_API_TIMEOUT_MS | 30000 | GCP OAuth, IAM, and Compute request timeout |
GCP_STS_SCOPE | https://www.googleapis.com/auth/cloud-platform | WIF STS exchange scope |
GCP_SA_IMPERSONATION_SCOPES | https://www.googleapis.com/auth/compute | Comma-separated scopes for WIF service-account impersonation |
GCP_SA_TOKEN_LIFETIME_SECONDS | 3600 | WIF impersonated access-token lifetime |
GCP_STS_TOKEN_URL | https://sts.googleapis.com/v1/token | WIF STS endpoint override for controlled environments |
GCP_IAM_CREDENTIALS_BASE_URL | Google IAM Credentials API | WIF impersonation base URL override |
The service-account JWT bearer flow always uses https://oauth2.googleapis.com/token; it has no endpoint override, and uploaded token_uri values are ignored. Source credentials are encrypted in D1. Only derivative short-lived tokens are cached.
AI Idea Title Generation
Section titled “AI Idea Title Generation”| Variable | Default | Description |
|---|---|---|
TASK_TITLE_MODEL | @cf/google/gemma-4-26b-a4b-it | Workers AI model for title generation |
TASK_TITLE_MAX_LENGTH | 100 | Max characters in generated title |
TASK_TITLE_TIMEOUT_MS | 5000 | Timeout before falling back to truncation |
TASK_TITLE_GENERATION_ENABLED | true | Set false to disable AI generation |
TASK_TITLE_SHORT_MESSAGE_THRESHOLD | 100 | Messages at or below this length bypass AI |
TASK_TITLE_MAX_RETRIES | 2 | Max retry attempts on failure |
TASK_TITLE_RETRY_DELAY_MS | 1000 | Base delay between retries (exponential backoff) |
TASK_TITLE_RETRY_MAX_DELAY_MS | 4000 | Max delay cap for backoff |
TASK_TITLE_ERROR_DIAGNOSTIC_MAX_LENGTH | 512 | Max sanitized provider-error diagnostic length |
Task Output Branches
Section titled “Task Output Branches”| Variable | Default | Description |
|---|---|---|
BRANCH_NAME_PREFIX | sam/ | Prefix for generated task output branches. Include the trailing separator (for example agent/). |
Task workspaces are checked out on the generated output branch, and SAM refuses to auto-push a completed task while the workspace is still on the project’s default branch. See Where the work lands.
Deployment Debugging Agent
Section titled “Deployment Debugging Agent”| Variable | Default | Description |
|---|---|---|
DEBUG_AGENT_MODEL | @cf/zai-org/glm-5.2 | Workers AI model for superadmin deployment diagnosis |
DEBUG_AGENT_MAX_TURNS | 6 | Maximum model/tool turns per diagnosis |
DEBUG_AGENT_RUN_TOKEN_LIMIT | 96000 | Combined token ceiling per diagnosis |
DEBUG_AGENT_MODEL_OUTPUT_TOKENS | 4096 | Maximum output tokens requested per model turn |
DEBUG_AGENT_DAILY_TOKEN_LIMIT | 480000 | Daily diagnosis token budget, counted per feature |
DEBUG_AGENT_TOOL_RESULT_LIMIT | 50 | Maximum rows returned by a diagnosis tool |
DEBUG_AGENT_TOOL_RESULT_BYTES | 32768 | Maximum serialized bytes per model-visible tool result |
DEBUG_AGENT_MAX_WINDOW_HOURS | 24 | Maximum selectable diagnosis window |
DEBUG_AGENT_TIMEOUT_MS | 120000 | Timeout for each diagnosis model request |
DEBUG_AGENT_HARD_DEADLINE_MS | 900000 | Hard deadline for an active diagnosis |
DEBUG_AGENT_STALE_HEARTBEAT_MS | 120000 | Orphan reconciler heartbeat threshold |
DEBUG_AGENT_RETRY_BASE_DELAY_MS | 2000 | Initial transient step retry delay |
DEBUG_AGENT_RETRY_MAX_DELAY_MS | 60000 | Maximum transient step retry delay |
DEBUG_AGENT_STEP_MAX_RETRIES | 3 | Maximum classified transient retries per step |
The /admin/errors view remains superadmin-only and may show local user IDs, IP addresses, and user-agent strings. Before any tool result enters model context, SAM recursively removes those fields plus credential-shaped values such as API tokens, JWTs, authorization headers, private keys, and long secret-like strings. Cloudflare credentials stay server-side and are never included in model messages or saved diagnosis text.
Same-installation VM diagnostic evidence
Section titled “Same-installation VM diagnostic evidence”VM failures use a durable local SQLite outbox and a private R2 artifact. Generated deployments set the R2 prefix and object lifecycle from Pulumi; the remaining Worker bounds can be overridden through deployment environment variables.
| Worker variable | Default | Description |
|---|---|---|
MAX_VM_AGENT_ERROR_BODY_BYTES | 32768 | Maximum VM error batch body |
MAX_VM_AGENT_ERROR_BATCH_SIZE | 10 | Maximum errors per VM batch |
MAX_VM_AGENT_ERROR_SOURCE_LENGTH | 256 | Maximum redacted VM error source length |
OBSERVABILITY_ERROR_MESSAGE_MAX_LENGTH | 2048 | Maximum persisted observability error message length |
OBSERVABILITY_ERROR_STACK_MAX_LENGTH | 4096 | Maximum persisted observability stack length |
OBSERVABILITY_ERROR_USER_AGENT_MAX_LENGTH | 512 | Maximum persisted observability user-agent length |
VM_INCIDENT_R2_PREFIX | diagnostic-incidents | Private object prefix; generated from the Pulumi output |
VM_INCIDENT_ARTIFACT_MAX_BYTES | 2097152 | Maximum compressed artifact size |
VM_INCIDENT_REGISTRATION_MAX_BYTES | 262144 | Maximum registration JSON body |
VM_INCIDENT_MANIFEST_MAX_BYTES | 131072 | Maximum redacted manifest |
VM_INCIDENT_PREVIEW_MAX_BYTES | 131072 | Maximum redacted model/UI preview |
VM_INCIDENT_MAX_ARTIFACTS_PER_NODE | 50 | Active artifact quota per node |
VM_INCIDENT_MAX_BYTES_PER_NODE | 104857600 | Active expected-byte quota per node |
VM_INCIDENT_RETENTION_DAYS | 7 | Private object and active metadata retention |
VM_INCIDENT_METADATA_RETENTION_DAYS | 30 | Expired metadata retention after object deletion |
VM_INCIDENT_PENDING_TIMEOUT_MINUTES | 30 | Incomplete-upload timeout and upload-lease duration |
VM_INCIDENT_RECONCILE_BATCH_SIZE | 50 | Maximum artifacts/incidents repaired per scheduled pass (minimum: 6) |
The VM Agent process accepts the corresponding ERROR_REPORT_* overrides for flush interval, batch size/bytes, outbox size and path, SQLite busy timeout, HTTP timeout, retry bounds, attempts, spool path/bytes, artifact bytes, retention, collector timeout/count/concurrency, document bytes, recursive value depth/items, string bytes, structured event limit, response-read bytes, and persisted-error bytes. Generated deployments pass these validated values through cloud-init into the VM Agent systemd service, so overrides apply to newly provisioned nodes. Defaults are listed in apps/api/.env.example; the common defaults are a 32 KiB error batch, 1,000-row outbox, 2 MiB artifact, 20 MiB spool, and 24-hour local retention.
Pulumi options diagnosticIncidentPrefix (default diagnostic-incidents) and diagnosticIncidentTtlDays (default 7, any positive integer) configure the private prefix and an independent R2 lifecycle rule. They do not require a separate bucket or manually managed Worker variable. The prefix cannot begin with the application-owned namespaces agents, cli, compose-image-artifacts, library, resource-history, session-snapshots, temp-uploads, or tts, because the lifecycle would otherwise expire unrelated objects.
Platform Feedback Triage
Section titled “Platform Feedback Triage”| Variable | Default | Description |
|---|---|---|
PLATFORM_FEEDBACK_PROJECT_ID | unset | Bootstrap/environment fallback for the project that receives user issue reports and automated triage draft Ideas. The Admin → Integrations runtime setting is preferred and overrides it. |
PLATFORM_FEEDBACK_TRIAGE_WINDOW_MINUTES | 60 | Lookback window for grouping recent platform errors |
PLATFORM_FEEDBACK_TRIAGE_ERROR_LIMIT | 100 | Maximum platform error rows scanned per triage sweep |
PLATFORM_FEEDBACK_TRIAGE_GROUP_LIMIT | 5 | Maximum grouped feedback candidates processed per triage sweep |
PLATFORM_FEEDBACK_TRIAGE_EVIDENCE_LIMIT | 10 | Maximum bounded error references retained per grouped feedback record |
PLATFORM_FEEDBACK_TRIAGE_CLAIM_TTL_MS | 600000 | Claim lease duration before a later sweep can reclaim the group |
PLATFORM_FEEDBACK_TRIAGE_MAX_FAILURES | 3 | Maximum failed attempts before a group is rejected from auto-triage |
PLATFORM_FEEDBACK_TRIAGE_FAILURE_REASON_MAX_LENGTH | 240 | Maximum characters stored or returned for sanitized failure reasons |
PLATFORM_FEEDBACK_TRIAGE_BUDGET_DEFER_MS | 86400000 | Retry delay for per-run budget deferrals |
PLATFORM_FEEDBACK_INCIDENT_DISPATCH_LEASE_TTL_MS | 7200000 | Dispatch lease before a failed incident trigger handoff can be reclaimed |
PLATFORM_FEEDBACK_INCIDENT_AGENT_LEASE_TTL_MS | 3600000 | Agent claim lease before another task can reclaim a private incident |
PLATFORM_FEEDBACK_INCIDENT_MAX_DISPATCH_ATTEMPTS | 3 | Agent-reported failed dispatch attempts before an incident is rejected |
PLATFORM_FEEDBACK_INCIDENT_REOPEN_COOLDOWN_MS | 1800000 | Minimum elapsed time after terminal resolution/expiry before a newer occurrence can reopen the same signature; set 0 to disable cooldown-only suppression |
PLATFORM_FEEDBACK_INCIDENT_RECLAIM_LIMIT | 25 | Maximum expired dispatch leases reclaimed by one incident sweep |
PLATFORM_FEEDBACK_INCIDENT_MAX_AGE_MS | 2592000000 | Maximum active incident age before expiry |
PLATFORM_FEEDBACK_INCIDENT_STALE_SINGLETON_MAX_AGE_MS | 259200000 | Maximum age for one-off pending incidents with no recurrence |
PLATFORM_FEEDBACK_INCIDENT_STALE_SINGLETON_EXPIRY_BATCH_SIZE | 25 | Maximum stale singleton incidents expired per sweep |
PLATFORM_FEEDBACK_INCIDENT_MIN_DISPATCH_SEVERITY | error | Minimum severity admitted to automatic VM incident dispatch |
PLATFORM_FEEDBACK_INCIDENT_MIN_DISPATCH_BATCH_SIZE | 2 | Dispatch immediately once this many eligible incidents are ready |
PLATFORM_FEEDBACK_INCIDENT_MIN_PENDING_AGE_MS | 1800000 | Dispatch a smaller eligible batch after this pending age |
PLATFORM_FEEDBACK_INCIDENT_DISPATCH_RATE_WINDOW_MS | 3600000 | Rate-cap window for each incident trigger |
PLATFORM_FEEDBACK_INCIDENT_MAX_DISPATCHES_PER_TRIGGER_WINDOW | 1 | Maximum dispatches one incident trigger may submit per rate window |
PLATFORM_FEEDBACK_INCIDENT_AUTO_TRIGGER_ENABLED | true | Auto-create one private incident trigger when pending incidents exist and no incident trigger exists |
PLATFORM_FEEDBACK_INCIDENT_TRIGGER_LIMIT | 5 | Maximum active incident triggers inspected per sweep |
PLATFORM_FEEDBACK_INCIDENT_TRIGGER_NAME | built-in | Name for the auto-created private incident trigger |
PLATFORM_FEEDBACK_INCIDENT_TRIGGER_TEMPLATE | built-in | Prompt template for the auto-created private incident trigger |
PLATFORM_FEEDBACK_INCIDENT_SUMMARY_LIMIT | 10 | Maximum grouped incidents included in one incident-trigger backlog summary |
PLATFORM_FEEDBACK_INCIDENT_EVIDENCE_REF_LIMIT | 10 | Maximum bounded evidence references retained per incident |
PLATFORM_FEEDBACK_INCIDENT_EVIDENCE_MAX_BYTES | 32768 | Maximum serialized evidence bytes retained per incident |
PLATFORM_FEEDBACK_INCIDENT_RESOLUTION_NOTE_MAX_LENGTH | 2000 | Maximum private incident resolution-note length |
Resolved and expired incident signatures reopen only when a newer occurrence arrives after PLATFORM_FEEDBACK_INCIDENT_REOPEN_COOLDOWN_MS; older lookback-window occurrences remain closed. VM-agent incidents resolved with fix evidence also wait for occurrences from nodes reporting the current VM_AGENT_REQUIRED_VERSION, because Worker deploys do not update already-running VM binaries. Dispatch attempts are consumed only when the incident task reports its own failure; platform-side handoff/session failures release the dispatch without incrementing dispatch_attempts.
Automated triage and superadmin-initiated diagnosis read the same DEBUG_AGENT_DAILY_TOKEN_LIMIT value but count against independent per-feature counters, so worst-case daily spend across both is twice this value. Automated triage treats budget exhaustion as a retryable deferral: daily exhaustion retries after the next UTC day starts, and per-run exhaustion uses PLATFORM_FEEDBACK_TRIAGE_BUDGET_DEFER_MS. Incident trigger agents run from the private grouped backlog and dispatch one agent for a bounded backlog summary, not one agent per occurrence. Automatic incident dispatch ignores pending signatures already linked to open tracked work and warning-only signatures below the configured severity floor, then applies the batch/age gate and per-trigger rate cap before reserving incidents.
Report an Issue
Section titled “Report an Issue”The in-app Report an Issue flow files user-submitted reports as draft Ideas in the effective private feedback project. Configure it from Admin → Integrations when possible; PLATFORM_FEEDBACK_PROJECT_ID remains the environment fallback when no runtime setting is saved. The feature is hidden entirely — both UI entry points disappear and GET /api/report-issue/config returns enabled: false — when no effective project exists or the effective project does not exist in this deployment’s database.
| Variable | Default | Description |
|---|---|---|
REPORT_ISSUE_TITLE_MAX_LENGTH | 200 | Truncation ceiling for the stored title (lowers only — see below) |
REPORT_ISSUE_DESCRIPTION_MAX_LENGTH | 5000 | Truncation ceiling for the stored description (lowers only) |
REPORT_ISSUE_CONTENT_MAX_LENGTH | 65536 | Maximum stored Idea body, including attached technical references |
RATE_LIMIT_REPORT_ISSUE_POST | 20 | Report submissions allowed per clock hour, per authenticated user |
The two length variables apply after request validation, so they can only lower the stored length — the request schema and the report dialog both enforce the built-in 200 / 5,000 caps regardless of what you set here.
See Reporting Issues for the user-facing flow and the untrusted-evidence Idea format.
Agent Model Catalog
Section titled “Agent Model Catalog”SAM loads OpenCode Zen and OpenCode Go model choices through the authenticated model-catalog API, backed by Models.dev and cached in KV. If the upstream catalog or cache is unavailable, SAM falls back to the static catalog shipped with the app.
| Variable | Default | Description |
|---|---|---|
MODEL_CATALOG_SOURCE_URL | https://models.dev/api.json | Source URL for the dynamic model catalog |
MODEL_CATALOG_CACHE_TTL_SECONDS | 3600 | KV cache TTL for normalized dynamic model catalog payloads |
MODEL_CATALOG_FETCH_TIMEOUT_MS | 5000 | Timeout for the upstream catalog fetch before static fallback |
Dashboard
Section titled “Dashboard”The dashboard’s Active Tasks list (GET /api/dashboard/active-tasks, apps/api/src/routes/dashboard.ts)
reads up to DASHBOARD_ACTIVE_TASK_CANDIDATE_LIMIT of the user’s most recently started active tasks,
ranks them by newest message (or by when they started, if they have none yet), and only then applies
the display limit. The open dashboard refreshes the list every 15 seconds while its tab is visible.
| Variable | Default | Description |
|---|---|---|
DASHBOARD_ACTIVE_TASK_LIMIT | 6 | Most recently active tasks the list shows |
DASHBOARD_ACTIVE_TASK_CANDIDATE_LIMIT | 100 | Active tasks read and ranked before the display limit applies (also the maximum; each project’s candidates share one SQL statement’s binds) |
DASHBOARD_INACTIVE_THRESHOLD_MS | 900000 (15 min) | A working task whose last message is newer than this shows Active; older shows Working |
HTTP Response Caching
Section titled “HTTP Response Caching”Conservative Cache-Control budgets for stable and semi-stable API GETs, letting the browser
serve a cached body instantly while it revalidates in the background. All values are seconds and
are clamped to [0, 86400]; an unparseable or negative value falls back to the default rather than
caching for longer.
Authenticated responses are always emitted as private with Vary: Cookie, so neither a shared
cache nor a second account in the same browser can be served another user’s body. Only the
unauthenticated /api/config/* endpoints are marked public. Endpoints returning real-time data
(chat messages, task status, session and workspace state) are deliberately excluded.
| Variable | Default | Description |
|---|---|---|
PUBLIC_CONFIG_CACHE_MAX_AGE_SECONDS | 60 | max-age for the unauthenticated /api/config/* endpoints |
PUBLIC_CONFIG_CACHE_SWR_SECONDS | 300 | stale-while-revalidate for /api/config/* |
MODEL_CATALOG_CACHE_MAX_AGE_SECONDS | 60 | max-age for GET /api/model-catalog/:agentType |
MODEL_CATALOG_CACHE_SWR_SECONDS | 300 | stale-while-revalidate for the model catalog response |
PROJECT_REFERENCE_CACHE_MAX_AGE_SECONDS | 0 | max-age for project agent-profile and skill lists (0 = always revalidate) |
PROJECT_REFERENCE_CACHE_SWR_SECONDS | 30 | stale-while-revalidate for project agent-profile and skill lists |
Durable Direct Provisioning
Section titled “Durable Direct Provisioning”Direct node allocation and workspace creation use isolated NodeLifecycle Durable Object instances. These optional Worker runtime overrides accept positive integers; invalid or unset values use the defaults.
| Variable | Default | Description |
|---|---|---|
NODE_PROVISIONING_REQUEST_TIMEOUT_MS | 5000 (5 s) | Allocation/reconciliation request budget, also used for background readiness and workspace dispatch. |
NODE_PROVISIONING_RETRY_INTERVAL_MS | 30000 (30 s) | Delay between durable provisioning attempts. |
NODE_PROVISIONING_MAX_AGE_MS | 900000 (15 min) | Maximum provisioning intent age before retries stop. |
NODE_PROVISIONING_MAX_ATTEMPTS | 30 | Maximum provisioning attempts before retries stop. |
Reaching either the age or attempt limit stops allocation retries. Diagnostic publication then uses a separate retry budget with the same interval and attempt limit; the age limit applies only to allocation and reconciliation. The unresolved intent is retained for inspection. Empty provider inventory does not authorize another provider create request or establish cleanup proof.
Warm Node Pooling
Section titled “Warm Node Pooling”| Variable | Default | Description |
|---|---|---|
NODE_WARM_TIMEOUT_MS | 1800000 (30 min) | Time a managed auto-provisioned node stays warm after its last active workspace leaves |
MAX_AUTO_NODE_LIFETIME_MS | 14400000 (4 hr) | Max lifetime for an auto-provisioned node holding no active workspaces |
NODE_WARM_GRACE_PERIOD_MS | 2100000 (35 min) | Cron sweep grace period (must be > warm timeout) |
NODE_LIFECYCLE_ALARM_RETRY_MS | 60000 (1 min) | Retry delay for DO alarm failures |
NODE_LIFECYCLE_MAX_DESTROYING_AGE_MS | 86400000 (24 hr) | Backstop after which a destroying-state alarm self-cleans; infrastructure teardown remains owned by cron/provider reconciliation |
DEFAULT_TASK_AGENT_TYPE | opencode | Default agent for autonomous idea execution |
Idle & Orphan Node Reaping
Section titled “Idle & Orphan Node Reaping”The cleanup sweep measures idleness from a node’s last workspace activity
(COALESCE(MAX(workspaces.updated_at), nodes.created_at)), never from
nodes.updated_at — heartbeats rewrite updated_at on every beat, so it tracks
liveness rather than idleness. The eligibility check is implemented by
claimNodeForCleanup() in apps/api/src/scheduled/node-cleanup/shared.ts.
Reaping only ever applies to nodes with node_role = 'workspace' and
node_class != 'user-owned'. Deployment nodes host long-running user applications
and legitimately hold zero workspaces forever, so they are never reaped by these
timers; they are released when their last deployment environment is deleted.
Stopped managed VM nodes created directly through a canonical pool can also be reaped using their server-recorded pool, credential, and native-offering identity. They retain the same workspace-activity and active-claim guards. A short project warm timeout does not bypass the workspace idle window. Runtime teardown preserves the saved snapshot and conversation for recovery on a fresh node.
| Variable | Default | Description |
|---|---|---|
NODE_WORKSPACE_IDLE_TIMEOUT_MS | 1800000 (30 min) | Last-workspace-activity window before an auto-provisioned node_role = 'workspace' node with no active workspaces can be destroyed. Uses COALESCE(MAX(workspaces.updated_at), nodes.created_at), never heartbeat-updated nodes.updated_at. |
NODE_ORPHAN_IDLE_TIMEOUT_MS | legacy alias | Backward-compatible alias used only when NODE_WORKSPACE_IDLE_TIMEOUT_MS is unset. |
NODE_ABSOLUTE_MAX_LIFETIME_MS | 86400000 (24 hr) | Hard ceiling on auto-provisioned workspace node age. Applies even when a workspace row still reports running, provided no workspace has reported activity within the idle window — this is what stops a stuck workspace row from making a node immortal. |
NODE_CLEANUP_SWEEP_LIMIT | 25 | Max node candidates processed per cleanup phase per cron run. |
NODE_CLEANUP_FAILURE_BACKOFF_MS | 3600000 (1 hr) | Expiring exclusion applied to failed cleanup candidates so a permanent provider error cannot monopolize the bounded page. |
NODE_UNHEALTHY_DRAIN_AFTER_MS | 600000 (10 min) | Heartbeat-loss window before SAM posts a session notice and requests sleep on a managed workspace VM. |
NODE_UNHEALTHY_RELEASE_AFTER_MS | 1800000 (30 min) | Heartbeat-loss window before SAM attempts strict provider deletion, even if the node still has active workspace rows. The value is clamped above the drain threshold. |
NODE_UNHEALTHY_FLEET_MAX_FRACTION | 0.5 | Holds destructive cleanup when this fraction of at least three managed workspace VMs lose heartbeat together, indicating a possible heartbeat-intake incident. The hold escalates after the configured release and drain windows combined; it does not delete busy nodes on a timer. |
NODE_UNHEALTHY_FLEET_MIN_NODES | 3 | Minimum number of managed workspace VMs required before the fleet-wide heartbeat-loss guard applies. |
NODE_UNHEALTHY_PRESERVATION_TIMEOUT_MS | 5000 (5 sec) | Per-node budget for attempting chat notices and sleep requests before cleanup proceeds. |
NODE_UNHEALTHY_RETRY_MS | 60000 (1 min) | Retry delay after provider deletion of an unhealthy node fails. |
NODE_STOPPED_HANDOFF_SWEEP_BUDGET_MS | 20000 (20 sec) | Wall-clock budget for stopped-node handoff. Candidates not started within the budget remain eligible for the next sweep. |
NODE_STOPPED_HANDOFF_REQUEST_TIMEOUT_MS | 5000 (5 sec) | Per-candidate provider/DNS deadline during stopped-node handoff, capped by remaining sweep time. Provider failures enter cleanup backoff. |
WORKSPACE_CLEANUP_SWEEP_LIMIT | 50 | Max workspace candidates processed per cleanup phase per cron run. |
NODE_AGENT_BACKGROUND_REQUEST_TIMEOUT_MS | 5000 (5 s) | VM-agent request timeout for background sweeps. Deliberately far below the interactive NODE_AGENT_REQUEST_TIMEOUT_MS (30 s) so a sweep over unreachable nodes cannot exhaust the Worker’s wall-clock budget. |
WORKSPACE_DELETION_RETRY_BASE_MS | 60000 (1 min) | Initial retry delay after a VM workspace deletion remains unconfirmed. |
WORKSPACE_DELETION_RETRY_MAX_MS | 3600000 (1 hr) | Maximum exponential backoff for an unconfirmed workspace deletion. |
WORKSPACE_DELETION_MAX_RESIDENCE_MS | 86400000 (24 hr) | Maximum hot-retry residence before an unconfirmed deletion enters durable operator quarantine; its workspace remains stopping and replacement-fenced. |
WORKSPACE_DELETION_ALARM_BATCH_SIZE | 3 | Maximum due workspace-deletion entries processed by one NodeLifecycle alarm, sized to stay within the Cloudflare Free-plan D1 query budget on the worst successful linked-workspace path. |
WORKSPACE_DELETION_CALLBACK_SIGNAL_TTL_SECONDS | 300 (5 min) | Per-workspace and callback-kind throttle for payload-free workspace.deletion_unconfirmed_callback activity evidence. |
WORKSPACE_DELETION_CALLBACK_SIGNAL_CLEANUP_LIMIT | 25 | Maximum expired callback telemetry throttle claims removed by one signal attempt. |
WORKSPACE_DELETION_DIAGNOSTIC_MAX_LENGTH | 500 characters | Maximum sanitized deletion-attempt diagnostic stored in workspaces.error_message. |
Operational Control-Loop Safety
Section titled “Operational Control-Loop Safety”The cron and Durable Object switches are availability brakes: an absent key or
KV read error means enabled (fail-open). This differs deliberately from the
fail-closed trials entitlement switch. Superadmins can inspect and update both
brakes through /api/admin/runtime-controls; emergency operators can use the
KV procedure in .claude/rules/55-runaway-cost-emergency-ops.md.
| Variable | Default | Description |
|---|---|---|
CRON_SWEEPS_ENABLED_KV_KEY | control-loops:cron-enabled | KV key gating the five-minute operational sweep block |
DO_ALARMS_ENABLED_KV_KEY | control-loops:alarms-enabled | Shared KV key gating alarm-bearing Durable Objects |
CONTROL_LOOP_KILL_SWITCH_CACHE_MS | 30000 | In-memory switch cache; runtime clamps it to at most 30 seconds |
CONTROL_LOOP_DISABLED_ALARM_RETRY_MS | 300000 (5 min) | Safe alarm recheck interval while DO work is disabled; values below 60 seconds are clamped |
CRON_FAILURE_NOTIFICATION_THROTTLE_MS | 3600000 (1 hr) | Per-sweep throttle enforced by a KV cache plus an atomic per-user Notification DO claim; also how long an archive-breaker alert claim is held |
CRON_FAILURE_NOTIFICATION_KV_PREFIX | cron-failure-notification | KV prefix for notification throttle markers |
DIAGNOSIS_COMPLETED_STEP_MIN_DELAY_MS | 1000 | Minimum delayed re-arm for an already-completed diagnosis step |
ORCHESTRATOR_ZERO_TASK_GRACE_MS | 600000 (10 min) | Grace period before an active mission with no tasks terminalizes |
ORCHESTRATOR_MAX_MISSION_LIFETIME_MS | 86400000 (24 hr) | Backstop that force-completes active/completing missions |
The scheduled Durable Object billing monitor reads these non-secret variables from the selected GitHub Environment, not from the API Worker runtime:
| Variable | Default/fallback | Description |
|---|---|---|
DO_WALL_TIME_SCRIPT_NAMES | none | Optional comma-separated API Worker filter for wall-time and invocation-rate analysis |
DO_INVOCATION_RATE_REGRESSION_RATIO | 2 | Recent-versus-seven-day-baseline request-rate failure ratio |
DO_CRON_LIVENESS_MAX_AGE_HOURS | 3 | Maximum age of the most recent targeted cron.completed event |
DO_CRON_LIVENESS_SCRIPT_NAMES | DO_WALL_TIME_SCRIPT_NAMES | Explicit API Worker service target for cron liveness; the GitHub workflow derives both from RESOURCE_PREFIX and the selected stack when unset |
DO_CRON_LIVENESS_ENDPOINT | Cloudflare Workers Observability query endpoint | Optional endpoint override for compatible/private telemetry gateways |
The selected GitHub Environment’s CF_API_TOKEN secret must include the
Cloudflare Workers Observability Write permission. Cloudflare requires that
permission for the telemetry query endpoint even though this monitor only reads
aggregated liveness telemetry.
Provider-Side Orphan Reconciliation
Section titled “Provider-Side Orphan Reconciliation”Reclaims cloud servers that exist at the provider but which no live database row claims — for example when a server was created but the control plane failed before recording its instance ID.
Because this is the only path that destroys infrastructure on the basis of absent
evidence, it fails closed at every step. A server must carry both the current
control-plane env value and the exact Pulumi-generated installation marker before
SAM consults D1. SAM then re-reads and revalidates the same provider resource immediately
before it calls the provider delete API. Provider-account membership, server
names, resource prefixes, and absence from this installation’s D1 are not ownership
proof.
Pulumi generates the non-secret installation identity automatically on first deploy,
persists it in the stack state, and injects it into the Worker as
SAM_INSTALLATION_ID; there is no manual GitHub Environment setting. An upgrade does
not relabel existing servers. Legacy servers without the marker remain usable and are
preserved indefinitely, while servers provisioned after the upgrade participate in
normal orphan cleanup. If the Pulumi state is lost or recreated, the new identity
safely leaves the old fleet unattributable instead of adopting it destructively. Any
missing/malformed identity, ambiguous provider metadata, or failed/malformed D1 lookup
skips deletion. Resources surfaced to reconciliation with non-owning metadata emit
aggregate operator-visible counters.
| Variable | Default | Description |
|---|---|---|
PROVIDER_ORPHAN_RECONCILIATION_ENABLED | true | Set to false to disable provider-side reconciliation entirely. |
PROVIDER_ORPHAN_MIN_AGE_MS | 3600000 (1 hr) | Minimum server age before it can be treated as an orphan. Must comfortably exceed provisioning time, since a server’s instance ID is recorded only after the provider returns it. |
PROVIDER_ORPHAN_DESTROY_LIMIT | 5 | Max servers destroyed per reconciliation run. |
PROVIDER_ORPHAN_RECONCILE_INTERVAL_MS | 3600000 (1 hr) | Minimum interval between runs. Invoked by the 5-minute cron but self-throttled to this interval via KV. |
Project Invites
Section titled “Project Invites”| Variable | Default | Description |
|---|---|---|
PROJECT_INVITE_TOKEN_BYTES | 32 | Random bytes used for generated project invite link tokens |
PROJECT_INVITE_DEFAULT_EXPIRY_DAYS | 7 | Default lifetime for invite links created without an explicit expiry |
PROJECT_INVITE_MAX_EXPIRY_DAYS | 30 | Maximum allowed invite link lifetime, including explicit expiry-date input |
PROJECT_OFFBOARDING_PLAN_TTL_SECONDS | 900 | Lifetime for project member offboarding preview plans before recomputation |
Notification System
Section titled “Notification System”| Variable | Default | Description |
|---|---|---|
NOTIFICATION_PROGRESS_BATCH_WINDOW_MS | 300000 (5 min) | Min interval between progress notifications per idea |
NOTIFICATION_DEDUP_WINDOW_MS | 60000 (60s) | Dedup window for task_complete notifications |
NOTIFICATION_AUTO_DELETE_AGE_MS | 7776000000 (90 days) | Auto-delete old notifications |
MAX_NOTIFICATIONS_PER_USER | 500 | Max stored notifications per user |
NOTIFICATION_PAGE_SIZE | 50 | Default page size for notification list |
MAX_NOTIFICATION_PAGE_SIZE | 100 | Max allowed page size |
HUMAN_INPUT_TIMEOUT_MS | 7200000 (2 hr) | Initial needs-input response window |
HUMAN_INPUT_ESCALATION_FRACTIONS | 0.25,0.75 | Reminder points within the initial response window |
HUMAN_INPUT_UNDELIVERED_GRACE_MS | 7200000 (2 hr) | Extension without confirmed push delivery |
HUMAN_INPUT_MAX_WAIT_MS | 86400000 (24 hr) | Hard maximum needs-input marker lifetime |
WEB_PUSH_TTL_SECONDS | 86400 | Push-service message TTL |
WEB_PUSH_VAPID_TTL_SECONDS | 43200 | VAPID authorization-token lifetime |
WEB_PUSH_DELIVERY_TIMEOUT_MS | 10000 | Per-attempt push-service timeout |
WEB_PUSH_DELIVERY_BUDGET_MS | 25000 | Total fan-out budget, hard-capped at 25s below Worker background limit |
WEB_PUSH_FANOUT_CONCURRENCY | 8 | Maximum concurrent endpoint deliveries |
WEB_PUSH_MAX_ATTEMPTS | 3 | Bounded transient delivery attempts |
WEB_PUSH_MAX_RETRY_AFTER_SECONDS | 30 | Maximum honored Retry-After delay |
WEB_PUSH_MAX_PAYLOAD_BYTES | 3500 | Maximum unencrypted payload size |
WEB_PUSH_FAILURE_THRESHOLD | 5 | Consecutive failures before disabling a subscription |
WEB_PUSH_MAX_SUBSCRIPTIONS_PER_USER | 8 | Maximum retained browser endpoints per user |
WEB_PUSH_USER_AGENT_MAX_LENGTH | 512 | Maximum stored browser description length |
RATE_LIMIT_PUSH_SUBSCRIPTION | 30 | Subscription mutations per user per hour |
Event Trigger Cleanup
Section titled “Event Trigger Cleanup”| Variable | Default | Description |
|---|---|---|
TRIGGER_STALE_EXECUTION_TIMEOUT_MS | 1800000 (30 min) | Age before running executions are checked against linked task liveness |
TRIGGER_STALE_QUEUED_TIMEOUT_MS | 300000 (5 min) | Age before queued executions are checked against linked task liveness |
TRIGGER_EXECUTION_HARD_MAX_RESIDENCE_HOURS | 48 | Hard maximum execution residence backstop; live linked tasks still control concurrency and incident dispatch use |
TRIGGER_EXECUTION_LOG_RETENTION_DAYS | 90 | Completed/failed/skipped execution log retention |
TRIGGER_EXECUTION_CLEANUP_ENABLED | enabled | Set to false to disable the cleanup sweep |
TRIGGER_STALE_RECOVERY_BATCH_SIZE | 100 | Maximum stale execution candidates processed per sweep |
MAX_TRIGGERS_PER_PROJECT | 20 | Platform default trigger cap per project. A project owner can raise or lower it per project via UI (Project → Settings → Scaling & Scheduling → Task Limits). |
Generic Webhook Triggers
Section titled “Generic Webhook Triggers”| Variable | Default | Description |
|---|---|---|
WEBHOOK_TRIGGERS_ENABLED | true | Public generic webhook ingress kill switch |
WEBHOOK_CREDENTIAL_CLAIM_TTL_SECONDS | 600 | Authenticated one-time MCP webhook credential claim lifetime |
WEBHOOK_TRIGGER_MAX_BODY_BYTES | 65536 | Maximum JSON request body size |
WEBHOOK_TRIGGER_MAX_FILTERS | 10 | Maximum deterministic filters per trigger |
WEBHOOK_TRIGGER_MAX_FILTER_PATH_LENGTH | 200 | Maximum configured filter dot-path length |
WEBHOOK_TRIGGER_MAX_FILTER_PATH_DEPTH | 8 | Maximum filter nesting depth at evaluation time |
WEBHOOK_TRIGGER_MAX_INCLUDED_HEADERS | 10 | Maximum safe request headers copied into template context |
WEBHOOK_TRIGGER_MAX_HEADER_NAME_LENGTH | 100 | Maximum configured included-header name length |
WEBHOOK_TRIGGER_MAX_SOURCE_LABEL_LENGTH | 100 | Maximum optional source label length |
WEBHOOK_TRIGGER_MAX_IDEMPOTENCY_KEY_LENGTH | 200 | Maximum accepted Idempotency-Key length |
WEBHOOK_INGRESS_RATE_LIMIT_PER_MINUTE | 120 | Best-effort pre-auth request damping per client IP/window |
WEBHOOK_TRIGGER_RATE_LIMIT_PER_MINUTE | 60 | Best-effort request damping per trigger/window |
WEBHOOK_INVALID_TOKEN_RATE_LIMIT_PER_MINUTE | 30 | Best-effort invalid-token damping per client IP/window |
WEBHOOK_RATE_LIMIT_WINDOW_SECONDS | 60 | Fixed rate-limit window length |
WEBHOOK_DELIVERY_RETENTION_DAYS | 7 | Retention for redacted delivery audit metadata |
WEBHOOK_DELIVERY_CLEANUP_BATCH_SIZE | 500 | Maximum expired audit rows deleted per cleanup pass |
WEBHOOK_DELIVERY_DEFAULT_PAGE_SIZE | 25 | Default delivery-history page size |
WEBHOOK_DELIVERY_MAX_PAGE_SIZE | 100 | Maximum delivery-history page size |
WEBHOOK_DELIVERY_PROCESSING_LEASE_SECONDS | 300 | Lease before an unsubmitted processing delivery can recover |
Webhook tokens use the existing ENCRYPTION_KEY as keyed-hash material and do not require a separate deployment secret. See Webhook Triggers for request, credential, filtering, and audit behavior.
Webhook damping uses Cloudflare KV’s eventually consistent read-update-write behavior. It reduces accidental bursts and abuse but is not a strict distributed quota.
Agent Requests in Chat
Section titled “Agent Requests in Chat”These control whether agents can stop and ask the person who started a chat for permission, ask
them a question, or send a link to open — see
When the Agent Needs You. All three switches
are false in the checked-in configuration; set them as GitHub Environment variables to turn them
on (see Let agents ask in chat). Turning a
switch on applies to agent sessions started afterwards; sleep and wake preserve their recorded
interaction settings;
turning one off refuses new requests at once, even in running sessions.
A permission request in a Chat session, and every question, waits up to
ACP_INTERACTION_PERMISSION_CONVERSATION_DEADLINE_MS; a permission request in a Task waits up to
ACP_INTERACTION_PERMISSION_TASK_DEADLINE_MS. Both are cut short a minute
(ACP_INTERACTION_DEADLINE_MARGIN_MS) before the agent’s own turn would time out.
| Variable | Default | Description |
|---|---|---|
ACP_INTERACTIONS_ENABLED | false | Turns on permission requests: an agent can ask the person who started a chat before it acts, and waits for the answer. While false, SAM refuses every request at once. The other two switches need this one too. Turning it off stops new requests; ones already waiting can still be answered until they expire. |
ACP_INTERACTION_FORMS_ENABLED | false | Turns on questions: a short form (choices, short text, numbers, yes/no) the agent asks in a Chat (conversation-mode) session. Agents decide when to ask; this switch only lets them. A form SAM can’t display is cancelled. Questions and answers are stored encrypted, and only the person who started the chat can see them. |
ACP_INTERACTION_URLS_ENABLED | false | Turns on links to open: a tool, usually an MCP server, asks the person who started a Chat session to open an https:// page to sign in or approve something. SAM never opens the link itself and refuses local addresses and localhost callbacks. |
ACP_INTERACTION_URL_DEADLINE_MS | 600000 (10 min) | Maximum URL request window, further limited by the live prompt deadline. Existing requests remain answerable when the URL switch is turned off. |
ACP_INTERACTION_URL_MAX_CHARS | 8192 | Maximum HTTPS URL length. Overrides can lower this bound. |
ACP_INTERACTION_URL_ELICITATION_ID_MAX_CHARS | 256 | Maximum wrapper URL request ID length. Overrides can lower this bound. |
ACP_INTERACTION_URL_REDIRECT_DEPTH | 2 | Maximum nesting of explicit redirect/callback query URLs. Overrides can lower this bound; SAM does not fetch redirects. |
ACP_INTERACTION_FORM_SCHEMA_MAX_BYTES | 16384 (16 KiB) | Maximum supported form schema size, including descriptions and previews. |
ACP_INTERACTION_FORM_SCHEMA_MAX_PROPERTIES | 20 | Maximum fields in one supported form. |
ACP_INTERACTION_FORM_SCHEMA_MAX_ENUM | 50 | Maximum choices per select or multi-select field. |
ACP_INTERACTION_ANSWER_MAX_BYTES | 16384 (16 KiB) | Maximum accepted form answer payload size. |
ACP_INTERACTION_ANSWER_STRING_MAX_BYTES | 4096 (4 KiB) | Maximum size of a string answer or selected choice. |
ACP_INTERACTION_PERMISSION_TASK_DEADLINE_MS | 1800000 (30 min) | Runtime permission deadline for task sessions. |
ACP_INTERACTION_PERMISSION_CONVERSATION_DEADLINE_MS | 7200000 (2 hours) | Runtime permission deadline for conversation sessions. |
ACP_INTERACTION_DEADLINE_MARGIN_MS | 60000 (1 min) | Safety margin before the enclosing prompt deadline when computing a permission deadline. |
ACP_INTERACTION_MAX_DEADLINE_MS | 14400000 (4 hours) | Longest any request may wait. Keep every deadline setting at or below it: while requests are on, a deadline above it makes the VM agent reject the session start, so new agent sessions fail to start until you fix it. |
ACP_INTERACTION_MAX_PENDING_PER_SESSION | 8 | Most requests that can wait in one chat at once; a further request is refused. |
ACP_INTERACTION_OPTIONS_MAX_COUNT | 16 | Most buttons a permission request can offer. |
ACP_INTERACTION_REQUEST_MAX_BYTES | 32768 (32 KiB) | Largest request detail SAM accepts from an agent. |
ACP_INTERACTION_DELIVERY_WINDOW_MS | 900000 (15 min) | How long SAM keeps trying to hand an answer to the agent (never past the request’s deadline) before the card says delivery is unconfirmed. |
ACP_INTERACTION_SENSITIVE_PURGE_MS | 3600000 (1 hour) | How long a request’s question and answer are kept after it is settled; after that only its outcome remains. |
ACP_INTERACTION_SUMMARY_RETENTION_MS | 2592000000 (30 days) | How long settled requests’ outcomes are kept. The newest 100 per chat are kept regardless. |
ACP_INTERACTION_OPTION_ID_MAX_CHARS | 128 | Maximum permission option or form field identifier length accepted by the runtime bridge. |
ACP_INTERACTION_OPTION_NAME_MAX_CHARS | 200 | Maximum permission option or form field label length accepted by the runtime bridge. |
ACP_INTERACTION_RUNTIME_RECEIPT_LIMIT | 256 | Maximum in-memory idempotency receipts and distinct URL request IDs retained per live SessionHost generation. Exhausting the URL ID limit fails closed until a new runtime generation starts; retained IDs prevent late duplicate completion from binding to a new request. |
ACP_INTERACTION_RUNTIME_RESPONSE_MAX_BYTES | 65536 (64 KiB) | Maximum response body read by the runtime for interaction creation and settlement. |
The remaining ACP_INTERACTION_* settings (*_BATCH_SIZE, RETRY_*, ALARM_*, *_LAST_SETTLED)
tune how SAM stores and delivers requests internally and rarely need changing; their defaults are in
packages/shared/src/acp-interactions.ts.
ACP Session Lifecycle
Section titled “ACP Session Lifecycle”| Variable | Default | Description |
|---|---|---|
ACP_SESSION_DETECTION_WINDOW_MS | 300000 (5 min) | Stale ProjectData ACP heartbeat detection window. VM sessions are not interrupted solely from stale/missing ProjectData heartbeat rows; the timeout must be paired with conclusive runtime/workspace evidence. |
ACP_SESSION_HEARTBEAT_INTERVAL_MS | 60000 (60s) | How often VM agent sends heartbeats |
ACP_SESSION_RECONCILIATION_TIMEOUT_MS | 30000 (30s) | VM agent startup reconciliation timeout |
ACP_SESSION_MAX_FORK_DEPTH | 10 | Maximum session fork chain depth |
ACP_SESSION_FORK_CONTEXT_MESSAGES | 20 | Context messages included when forking |
ACP Protocol (VM Agent)
Section titled “ACP Protocol (VM Agent)”Task-managed ACP prompts use control-plane inactivity classification and the task absolute ceiling. ACP_TASK_PROMPT_TIMEOUT has been removed; legacy values no longer impose a duration-only failure. ACP_PROMPT_TIMEOUT applies only to unmanaged workspace sessions.
| Variable | Default | Description |
|---|---|---|
ACP_MESSAGE_BUFFER_SIZE | 5000 | Buffer size for ACP messages |
ACP_STDERR_BUFFER_BYTES | 4096 | Agent stderr bytes retained for crash reports |
ACP_PING_INTERVAL | 30s | WebSocket keepalive ping interval |
ACP_PONG_TIMEOUT | 10s | Pong response timeout |
ACP_PROMPT_RETRY_MAX_RETRIES | 2 | Max transient provider prompt retries after the initial attempt |
ACP_PROMPT_RETRY_INITIAL_BACKOFF | 15s | Initial backoff before retrying transient provider prompt errors |
ACP_PROMPT_RETRY_MAX_BACKOFF | 2m | Max exponential backoff for transient provider prompt retries |
ACTIVITY_REREPORT_INTERVAL | 60s | Re-send prompting activity while a prompt is active |
ACP_HARNESS_ACTIVITY_REPORT_DEBOUNCE | 750ms | Debounce ACP harness/tool-call activity reports before callbacks |
ACP_CHECKPOINT_PREEMPT_GRACE | 30s | Graceful ACP cancel/close wait before harness force-stop |
ACP_CHECKPOINT_PREEMPT_MAX_GRACE | 2m | Maximum caller-selected checkpoint rollover grace |
ACP_CHECKPOINT_ROLLOVER_TIMEOUT | 2m | Full checkpoint restart and strict LoadSession deadline |
ACTIVITY_TERMINAL_REPORT_ATTEMPTS | 5 | Retry attempts for terminal activity reports |
ACTIVITY_TERMINAL_REPORT_BACKOFF | 1s | Backoff between terminal activity report retries |
ACP_IDLE_SUSPEND_TIMEOUT | 30m | Idle session auto-suspend timeout |
ACP_NOTIF_SERIALIZE_TIMEOUT | 5s | Notification serialization timeout |
MCP (Agent Tools)
Section titled “MCP (Agent Tools)”| Variable | Default | Description |
|---|---|---|
MCP_TOKEN_TTL_SECONDS | 28800 (8 hours) | Sliding inactivity timeout for agent MCP access |
MCP_RATE_LIMIT | 120 | Max MCP requests per window |
MCP_RATE_LIMIT_WINDOW_SECONDS | 60 | Rate limit window |
MCP_DISPATCH_MAX_DEPTH | 3 | Max recursion depth for dispatch_task |
MCP_DISPATCH_MAX_PER_TASK | 5 | Max dispatched tasks per parent task |
MCP_DISPATCH_MAX_ACTIVE_PER_PROJECT | 10 | Max active dispatched tasks per project |
ORCHESTRATOR_STOP_CAS_MAX_ATTEMPTS | 2 | Task-status CAS attempts after a hard stop |
Voice & Text-to-Speech
Section titled “Voice & Text-to-Speech”| Variable | Default | Description |
|---|---|---|
WHISPER_MODEL_ID | @cf/openai/whisper-large-v3-turbo | Transcription model |
MAX_AUDIO_SIZE_BYTES | 10485760 (10 MB) | Max upload audio size |
MAX_AUDIO_DURATION_SECONDS | 60 | Max recording duration |
RATE_LIMIT_TRANSCRIBE | 30 | Max transcriptions per user per window |
RATE_LIMIT_TRANSCRIBE_WINDOW_SECONDS | 60 | Transcription rate-limit window (seconds) |
TTS_ENABLED | true | Enable/disable text-to-speech |
TTS_MODEL | @cf/deepgram/aura-2-en | TTS model |
TTS_SPEAKER | luna | TTS voice selection |
TTS_ENCODING | mp3 | Audio output format |
TTS_MAX_TEXT_LENGTH | 100000 | Max characters per TTS synthesis |
TTS_TIMEOUT_MS | 60000 | TTS synthesis timeout |
Idea Execution Timeouts
Section titled “Idea Execution Timeouts”| Variable | Default | Description |
|---|---|---|
TASK_RUN_MAX_EXECUTION_MS | 14400000 (4 hr) | Age from which each stuck-task sweep checks an in_progress task’s task-scoped runtime liveness. A conclusively dead runtime is failed; a live one is kept, bounded only by TASK_RUN_ABSOLUTE_CEILING_MS |
TASK_STUCK_QUEUED_TIMEOUT_MS | 1200000 (20 min) | Timeout for tasks stuck in queued state |
TASK_STUCK_DELEGATED_TIMEOUT_MS | 1860000 (31 min) | Timeout for tasks stuck in delegated state |
TASK_DO_MISMATCH_GRACE_MS | 300000 (5 min) | Minimum age before reconciling completed TaskRunner state with task-scoped liveness |
STUCK_TASK_MAX_CANDIDATES_PER_SWEEP | 100 | Maximum active tasks inspected by each recovery sweep |
STUCK_TASK_SCAN_CURSOR_KV_KEY | scheduled:stuck-tasks:scan-cursor:v1 | KV key used to resume bounded recovery scans fairly across active tasks |
TASK_LIVENESS_MAX_ACP_SESSIONS | 5 | Maximum task-scoped ACP sessions inspected per liveness probe |
TASK_LIVENESS_PROBE_TIMEOUT_MS | 5000 (5 sec) | Per-candidate timeout for ACP and Instant lifecycle probes used by ProjectData heartbeat deferral, idle cleanup, and stuck-task reconciliation; a timeout is inconclusive and preserves the task and workspace |
TASK_LIVENESS_NODE_HEALTH_PROBE_TIMEOUT_MS | 5000 (5 sec) | Per-candidate timeout for stale-VM-node health probes used by ProjectData idle cleanup and stuck-task reconciliation; a timeout is inconclusive and preserves the task and workspace |
IDLE_CLEANUP_MAX_CANDIDATES_PER_SWEEP | 5 | Maximum exact-session task candidates inspected by a ProjectData idle-cleanup pass; workspace deletion is deferred when this bound cannot prove every reporter-scoped runtime conclusively dead |
IDLE_CLEANUP_MAX_RESIDENCE_MS | 7200000 (2 hr) | Maximum residence for a ProjectData idle-cleanup schedule before repeated preserved/error outcomes stop re-arming, preserve the workspace, and surface an attention marker |
WORKSPACE_IDLE_TIMEOUT_MS | 7200000 (2 hr) | Installation default for how long an active chat session’s workspace can go without messages or terminal activity before ProjectData retires it, once its runtime is conclusively dead; a project’s Workspace Idle Timeout (30 min to 24 hr) overrides it |
WORKSPACE_IDLE_BACKOFF_BASE_MS | 600000 (10 min) | First retry delay after a ProjectData workspace-idle check finds an idle workspace it cannot retire yet: inconclusive task candidates, a live or unprovable runtime, a missing project identity, or a failed check |
WORKSPACE_IDLE_BACKOFF_MAX_MS | 21600000 (6 hr) | Maximum retry delay for repeated workspace-idle checks that cannot retire an idle workspace; the delay doubles from the base and resets on new activity or when the session wakes |
TASK_RUN_ABSOLUTE_CEILING_MS | 86400000 (24 hr) | Absolute runaway-cost ceiling; fails even a task with a demonstrably live runtime. Measured from the age of the currently allocated runtime generation (workspaces.created_at), not from tasks.started_at, and applied only while a runtime generation exists — a conversation whose workspace has been released holds no compute for it to bound |
TASK_RUN_ABSOLUTE_CEILING_SLEEP_GRACE_MS | 3600000 (1 hr) | Longest the absolute ceiling waits for a sleep that is still only in flight (scheduled, capturing, stopping or retrying) once the ceiling has passed. Measured on runtime-generation age, so sleep retries cannot renew it; a restorable sleep record always defers the ceiling. Keep it above SESSION_SLEEP_IN_FLIGHT_MAX_AGE_MS |
STALLED_TASK_CLASSIFIER_ENABLED | true | Once a task is past TASK_RUN_MAX_EXECUTION_MS and kept alive only by an open agent turn, ask a Workers AI classifier whether the turn is stalled; a confident stalled verdict fails the task with “SAM detected a stalled agent turn after N minutes”. Set false to keep such turns running until the absolute ceiling |
STALLED_TASK_CLASSIFIER_MIN_ACTIVITY_AGE_MS | 3600000 (1 hr) | How long the agent’s current turn (or its running work) must have gone on, and the transcript stayed quiet, before the classifier is asked |
STALLED_TASK_CLASSIFIER_CONFIDENCE_THRESHOLD | 0.8 | Stalled probability required to fail the task; uncertain or failed classifications keep the task running |
STALLED_TASK_CLASSIFIER_MODEL | @cf/cloudflare/clef | Workers AI model used for the stall classification |
STALLED_TASK_CLASSIFIER_SELECTOR | clef | Clef selector passed with each classification |
STALLED_TASK_CLASSIFIER_TIMEOUT_MS | 10000 (10 sec) | Per-classification timeout; a timeout counts as uncertain |
STALLED_TASK_CLASSIFIER_MESSAGE_LIMIT | 200 | Most recent transcript rows read for the classification |
STALLED_TASK_CLASSIFIER_TRANSCRIPT_MAX_CHARS | 24000 | Maximum transcript characters, lightly redacted, sent to the classifier |
CLAUDE_CODE_COMPACTION_LOOP_DETECTOR_ENABLED | true | Enable Claude Code compaction-loop shutdown from recent message evidence |
CLAUDE_CODE_COMPACTION_LOOP_RECENT_MESSAGE_LIMIT | 40 | Recent task-session messages to inspect for compaction-loop evidence |
CLAUDE_CODE_COMPACTION_LOOP_WINDOW_MESSAGES | 20 | Rolling recent-message window used for compaction-loop detection |
CLAUDE_CODE_COMPACTION_LOOP_MIN_PAIRS | 3 | Minimum Compacting... / Compacting completed marker pairs before failing a task |
TASK_CALLBACK_TIMEOUT_MS | 10000 | Callback response timeout |
TASK_CALLBACK_RETRY_MAX_ATTEMPTS | 3 | Max callback retry attempts |
TASK_RUN_CLEANUP_DELAY_MS | 5000 | Delay before task cleanup |
TASK_RECONCILIATION_IDLE_MS | 300000 (5 min) | Idle threshold before SAM sends a visible task check-in |
TASK_RECONCILIATION_RESPONSE_DEADLINE_MS | 60000 (1 min) | Response deadline after a visible task check-in |
TASK_RECONCILIATION_PROMPT_SOFT_STALL_MS | 1800000 (30 min) | In-flight prompt observation threshold before a non-interrupting reconciliation event |
TASK_RECONCILIATION_PROMPT_HARD_STALL_MS | 7200000 (2 hr) | In-flight prompt hard-stall threshold before SAM requests prompt cancellation |
TASK_RECONCILIATION_ACTIVE_WORK_HARD_STALL_MS | 7200000 (2 hr) | Hard ceiling for deferring an expired task check-in because prompt/tool work is still active |
TASK_RECONCILIATION_MIN_ALARM_DELAY_MS | 10000 (10 sec) | Minimum delay before the next reconciliation alarm can fire |
TASK_RECONCILIATION_MAX_CANDIDATES_PER_SWEEP | 5 | Maximum reconciliation candidates assessed per ProjectData alarm pass |
TASK_RECONCILIATION_NODE_CALL_TIMEOUT_MS | 5000 (5 sec) | Bounded timeout for reconciliation check-in delivery and prompt cancellation |
TASK_RECONCILIATION_CANDIDATE_LEASE_MS | 30000 (30 sec) | Durable claim floor preventing overlapping alarms from repeating reconciliation; the effective lease is clamped to cover the configured liveness-probe, node-call, and minimum-alarm-delay budgets |
TASK_RECONCILIATION_MAX_CHECKINS | 3 | Maximum automatic check-ins without confirmed tool progress. At the limit, pause nudges and ask Clef once; human input or new completed tool work resets the budget. Classifier failure never grants more retries. |
TASK_RECONCILIATION_PROBE_MAX_ATTEMPTS | 3 | Consecutive inconclusive task reconciliation attempts before quarantine |
TASK_RECONCILIATION_QUARANTINE_MS | 300000 (5 min) | Cooldown after task reconciliation exhausts its inconclusive-attempt budget |
INSTANT_START_STALE_TIMEOUT_MS | 600000 (10 min) | How long an Instant session may sit mid-launch (execution step instant_persistence) before the recovery sweep treats its start as stuck and fails it. Instant starts are accepted and then finished in the background, so this bounds a launch that never completes. |
Durable prompt delivery and checkpoint storage
Section titled “Durable prompt delivery and checkpoint storage”Durable prompt delivery is enabled by default so a follow-up can remain queued while a sleeping VM is replaced and restored. Legacy VM compatibility remains disabled: targets must advertise stable delivery receipts, and receipt ambiguity fails visibly rather than being guessed or replayed.
| Variable | Default | Description |
|---|---|---|
DURABLE_PROMPT_DELIVERY_ENABLED | true | Persist prompts and deliver them from ProjectData alarms, including sleeping-session wake. |
PROMPT_DELIVERY_LEGACY_VM_COMPAT_ENABLED | false | Explicit old-VM compatibility switch; receipt ambiguity still fails visibly and is never guessed or replayed. |
PROMPT_DELIVERY_MAX_CANDIDATES_PER_ALARM | 5 | Maximum delivery claims started by one alarm pass. |
PROMPT_DELIVERY_MAX_ATTEMPTS | 5 | Counted delivery attempts before retryable busy/not-ready waits use capped backoff; TTL remains the hard bound. |
PROMPT_DELIVERY_RETRY_BASE_MS | 5000 | Initial retry delay. |
PROMPT_DELIVERY_RETRY_MAX_MS | 300000 | Maximum exponential retry delay. |
PROMPT_DELIVERY_TTL_MS | 3600000 | Maximum unresolved delivery lifetime. |
PROMPT_DELIVERY_RECEIPT_TIMEOUT_MS | 30000 | Age at which an unconfirmed claim enters receipt reconciliation. |
PROMPT_DELIVERY_BACKGROUND_TIMEOUT_MS | 5000 | Deadline for pre-send target/recovery preparation; also bounds background VM submit and receipt calls. Preparation timeouts remain retryable without sending a prompt. |
PROMPT_DELIVERY_MIN_ALARM_DELAY_MS | 1000 | Minimum delay before the next delivery alarm. |
ACP_LONG_TURN_SUPERVISOR_ENABLED | false | Reserved long-turn candidate/preemption engine switch; this release leaves it inert. |
ACP_LONG_TURN_CHECKPOINT_MS | 18000000 (5 hr) | Reserved checkpoint eligibility threshold. |
ACP_CHECKPOINT_PREEMPT_GRACE_MS | 30000 | Reserved graceful preemption window. |
ORCHESTRATOR_WAIT_RECONCILE_INTERVAL_MS | 30000 | D1 reconciliation backstop interval for active parent waits. |
ORCHESTRATOR_WAIT_MAX_CHILDREN | 20 | Maximum same-project task IDs selected by one durable wait (hard ceiling: 90, preserving D1 bind headroom). |
ORCHESTRATOR_WAIT_MAX_ACTIVE_PER_PROJECT | 100 | Maximum active durable parent waits per project. |
ORCHESTRATOR_WAIT_MAX_DURATION_MS | 86400000 | Maximum finite wait deadline. |
ORCHESTRATOR_WAIT_MAX_CANDIDATES_PER_ALARM | 10 | Maximum wait subscriptions reconciled by one ProjectData alarm. |
ProjectData stores a single prompt-delivery queue and checkpoint episodes keyed by ACP session and prompt epoch. Sleeping-session prompts stay in that queue until strict restore succeeds, then use stable receipts for exactly-once acceptance. Task agents can register wait_for_subtasks for same-project tasks; terminal hooks provide low-latency nudges, bounded alarms reconcile missed writers, and one stable delivery ID wakes the caller exactly once. Automatic checkpoint preemption remains disabled.
Liveness-gated recovery. Stuck-task recovery for
in_progresstasks (including task-mode work paused at theawaiting_followupexecution step) is gated on task-scoped liveness — a live workspace, node reachability, and an active task-scoped ACP session. A healthy, recent D1 node mirror satisfies the reachability gate directly; a stale or unhealthy mirror is only suspect and must be contradicted by a successful bounded node-health probe before task-scoped ACP classification continues. A shared-node heartbeat or health response alone is never sufficient. Consequently,TASK_RUN_MAX_EXECUTION_MSis the point from which a task with no proven live runtime is failed; a task with a demonstrably live runtime is preserved past it, bounded only byTASK_RUN_ABSOLUTE_CEILING_MSas a runaway-cost backstop (there is no separate hard timeout). When liveness cannot be determined (including a probe timeout, transport error, or failed health response), the task is left untouched (fail-safe). A live verdict does not mean work is happening: the task’s own agent may simply be alive and idle, waiting for its user. Each preserve is logged asstuck_task.skipped_active_heartbeatwith the liveness basis (livenessReason), the agent’s work state (idle,prompt_turn_active,runtime_work_active, orprompt_turn_unprovenfor a turn that has reported nothing pastSESSION_ACTIVITY_STALE_THRESHOLD_MS, logged as a warning) with its ages, and, for an idle agent, whether a sleep is in flight to release it. Onestuck_task_heartbeat_skiprow per task and basis is kept in the observability database. Idle runtimes are released by automatic sleep, not by stuck-task recovery (apps/api/src/scheduled/stuck-task-live-runtime.ts).
A sleeping conversation is never failed. Sleeping is what deletes a workspace and destroys its node, so a slept session is indistinguishable from a dead one by those signals alone. Before any terminal verdict, recovery reads the chat session’s own
session_snapshotsrow through the same predicate the wake path uses to authorize a restore (restorableOrInFlightSleepSnapshotPredicateSqlinapps/api/src/services/session-snapshot-sleep-predicate.ts, vialoadTaskSleepPreservation). A session that is asleep and restorable, or with a sleep capture still in flight, is preserved with no status change and no error message — a restorable one including pastTASK_RUN_ABSOLUTE_CEILING_MS, because a released runtime holds no compute for the cost ceiling to bound. The preserve is bounded on both arms and cannot strand a task: the restorable arm expires with the snapshot’s ownexpires_at, and the in-flight arm withSESSION_SLEEP_IN_FLIGHT_MAX_AGE_MSafter the sleep’s latest attempt. Because every retry renews that age, a sleep that is only in flight on still-allocated compute defers the ceiling for at mostTASK_RUN_ABSOLUTE_CEILING_SLEEP_GRACE_MS, measured on runtime-generation age (evaluateRunawayCostCeilinginapps/api/src/scheduled/stuck-task-ceiling.ts); past it the ceiling terminalizes the task and the terminal gate does not re-defer to that sleep. Once either lapses, the ordinary liveness path terminalizes the task with its accurate reason. A failed snapshot read withholds the verdict until the next sweep.
Node & Workspace Readiness
Section titled “Node & Workspace Readiness”| Variable | Default | Description |
|---|---|---|
NODE_AGENT_READY_TIMEOUT_MS | 900000 (15 min) | Wait for VM agent to report ready |
NODE_AGENT_READY_POLL_INTERVAL_MS | 5000 | Poll interval for agent readiness |
VM_AGENT_REQUIRED_VERSION | (deploy-generated) | Required vm-agent build for reusable VM nodes. Official deploys derive this from the last commit that changed a vm-agent build input, not from the deployment commit, so a Worker-only deploy does not make every running node ineligible for reuse. Build inputs are packages/vm-agent/** minus exactly four pathspecs — .claude/, the top-level AGENTS.md, **/*_test.go and *_test.go — so a commit touching only those deliberately leaves the release, and the node pool, untouched. Any other file under that directory rotates the release even if it reads as documentation; the resolver excludes rather than includes so that an unrecognised new file fails safe. Binaries are published under that immutable release key first, and cloud-init requests that exact release. Leave unset only for local/manual development or skip-agent deploys. |
TASK_RUNNER_STEP_MAX_RETRIES | 3 | Max retries per TaskRunner step before failing the task |
TASK_RUNNER_RETRY_BASE_DELAY_MS | 5000 | Base delay for TaskRunner retry backoff |
TASK_RUNNER_RETRY_MAX_DELAY_MS | 60000 | Maximum delay for TaskRunner retry backoff |
TASK_RUNNER_AGENT_POLL_INTERVAL_MS | 5000 | TaskRunner D1 poll interval while waiting for a fresh provisioned VM agent |
TASK_RUNNER_AGENT_READY_TIMEOUT_MS | 900000 (15 min) | TaskRunner max wait for a provisioned VM agent before failing the task |
TASK_RUNNER_AGENT_READY_FRESHNESS_SKEW_MS | 30000 | Timestamp skew tolerated between TaskRunner wait start, node heartbeat, and /ready signals |
TASK_RUNNER_WORKSPACE_DISPATCH_TIMEOUT_MS | 600000 (10 min) | Max wait for VM-agent workspace dispatch acknowledgement |
TASK_RUNNER_WORKSPACE_DISPATCH_BASE_DELAY_MS | 30000 | Base delay for workspace dispatch retry backoff |
TASK_RUNNER_WORKSPACE_DISPATCH_MAX_DELAY_MS | 120000 (2 min) | Maximum delay for workspace dispatch retry backoff |
TASK_RUNNER_WORKSPACE_READY_TIMEOUT_MS | 1800000 (30 min) | Max wait for workspace-ready callback |
TASK_RUNNER_WORKSPACE_READY_POLL_INTERVAL_MS | 30000 | D1 poll interval during the TaskRunner workspace-ready step |
TASK_RUNNER_PROVISION_POLL_INTERVAL_MS | 10000 | TaskRunner provision status poll interval |
TASK_RUNNER_PROVISION_TIMEOUT_MS | 900000 (15 min) | Max time for TaskRunner node provisioning before permanent failure |
PROVISIONING_TIMEOUT_MS | 1800000 (30 min) | Cron marks stuck workspaces as error |
NODE_HEARTBEAT_STALE_SECONDS | 180 | Seconds without a heartbeat before a node is treated as stale |
Host resource protection
Section titled “Host resource protection”Each VM runs the SAM agent and the workspace containers side by side. systemd slices keep the agent from being starved by the work it is supervising: sam-infra.slice holds the agent, sam-workload.slice is Docker’s cgroup parent, and both sit under sam.slice.
Memory and CPU are protected differently on purpose. Memory exhaustion kills the agent, so it gets a hard reservation. CPU contention only slows it, so it gets a proportional share — which Linux applies only when something is actually competing, and therefore costs nothing on an idle machine.
| Variable | Default | Description |
|---|---|---|
SAM_INFRA_SLICE_MEMORY_MIN_MB | 256 | systemd MemoryMin reserved for the agent slice |
VM_AGENT_MEMORY_RESERVE_MB | (set by deployment) | Memory held back from the workload slice’s MemoryMax, leaving headroom for the agent |
SAM_INFRA_SLICE_CPU_WEIGHT | 1000 | systemd CPUWeight for the agent slice (cgroup v2 range 1–10000). Higher than the workload weight so a saturated node cannot delay the heartbeat and get the node declared dead |
SAM_WORKLOAD_SLICE_CPU_WEIGHT | 100 | systemd CPUWeight for the Docker workload slice — the cgroup v2 default |
An out-of-range weight makes the unit fail to load, which would take the slice hierarchy and its memory reservation down with it, so cloud-init generation rejects anything outside 1–10000 rather than emitting it.
App Deployment Routing
Section titled “App Deployment Routing”| Variable | Default | Description |
|---|---|---|
DEPLOY_PAYLOAD_EXPIRY_SECONDS | 3600 | Signed deployment apply payload lifetime |
DEPLOYMENT_ROUTE_PORT_BASE | 35000 | First node-local loopback port reserved for app routes |
DEPLOYMENT_ROUTE_PORT_SPAN | 100 | Number of loopback ports reserved per deployment environment |
AGENT_DEPLOYMENT_RESERVED_ENVIRONMENT_NAMES | prod,production | Comma-separated environment names agents cannot create through MCP |
MAX_ENVIRONMENTS_PER_DEPLOYMENT_NODE | 5 | Maximum deployment environments to place on one deployment node |
DEPLOYMENT_DEFAULT_VM_SIZE | small | Default VM size for deployment nodes |
DEPLOYMENT_MODEL_RUNNER_VM_SIZE | medium | VM size for deployment nodes that need Docker Model Runner |
DEPLOYMENT_DEFAULT_CPU_LIMIT_MILLIS | 250 | Per-service CPU reservation when the whole resource block is omitted |
DEPLOYMENT_DEFAULT_MEMORY_LIMIT_MB | 256 | Per-service memory limit/reservation when the whole resource block is omitted |
DEPLOYMENT_DEFAULT_ROOT_DISK_MB | 1024 | Per-service root-disk reservation used for deployment placement |
DEPLOYMENT_LOG_MAX_SIZE | 10m | Default json-file log max-size for compose-publish releases |
DEPLOYMENT_LOG_MAX_FILE | 3 | Default json-file log max-file for compose-publish releases |
MCP_DEPLOYMENT_COMPOSE_PREVIEW_MAX_BYTES | 128000 | Max Compose YAML size accepted by deployment route preview MCP tool |
BUILD_PUBLISH_TOOL_TIMEOUT_MS | 1260000 | Worker-to-VM proxy timeout for build_and_publish |
DEPLOY_ACME_EMAIL | (unset) | Optional ACME contact email emitted into deployment-node Caddy config |
DEPLOY_ACME_CA | (unset) | Optional ACME CA directory override, useful for Let’s Encrypt staging |
DOH_RESOLVER_URL | https://cloudflare-dns.com/dns-query | DNS-over-HTTPS resolver used to verify deployment custom domains |
DOH_TIMEOUT_MS | 10000 | Timeout for deployment custom-domain DNS verification lookups |
DEPLOY_COMPOSE_CMD | docker compose | Docker Compose command used by the deployment engine |
DEPLOY_HEALTH_TIMEOUT | 5m | Deployment health-check timeout used by the VM agent |
DEPLOY_RUNTIME_TIMEOUT | 15m | VM-agent max time for deployment-node host dependency setup |
GRACEFUL_SHUTDOWN_TIMEOUT | 30s | VM-agent max time for graceful HTTP server shutdown after SIGTERM |
SYSTEM_PROVISIONING_TIMEOUT | 15m | VM-agent max time for workspace host provisioning before bootstrap |
CF_IP_FETCH_TIMEOUT | 10s | VM-agent timeout for Cloudflare IP range fetches during provisioning |
BOOT_LOG_HTTP_TIMEOUT | 10s | VM-agent timeout for boot-log callbacks to the control plane |
MCP_SHORT_COMMAND_TIMEOUT | 10s | VM-agent timeout for short MCP workspace command probes |
MCP_DIFF_COMMAND_TIMEOUT | 30s | VM-agent timeout for MCP diff-summary git commands |
MCP_BUILD_PREPARE_TIMEOUT | 30s | VM-agent timeout for MCP build/publish preparation probes |
JWKS_FETCH_TIMEOUT | 10s | VM-agent startup JWKS fetch timeout |
ACP_CREDENTIAL_SYNC_TIMEOUT | 10s | VM-agent ACP auth-file sync-back timeout during shutdown |
ACP_RESTART_ATTEMPT_TIMEOUT | 5m | VM-agent bound on one automatic ACP agent restart attempt |
ACP_ACTIVITY_REPORT_TIMEOUT | 10s | VM-agent timeout for each ACP activity callback attempt |
ACP_USAGE_PROBE_TIMEOUT | 10s | VM-agent bound on one post-turn provider usage probe (Codex rollout read or OpenCode Go usage request) |
OPENCODE_GO_USAGE_URL | https://opencode.ai/zen/go/v1/usage | VM-agent OpenCode Go usage endpoint probed with the session’s key after each completed opencode-go turn Must be https unless the host is loopback (localhost, 127.0.0.1, ::1), because the probe sends the key as a bearer token; the probe never follows redirects. |
DEVCONTAINER_CACHE_PUSH_TIMEOUT | 10m | VM-agent best-effort devcontainer cache image push timeout |
DEPLOY_PREFLIGHT_COMMAND_TIMEOUT | 15s | VM-agent deployment preflight diagnostic command timeout |
LOG_STREAM_PING_WRITE_TIMEOUT | 10s | VM-agent log-stream WebSocket ping write deadline |
DEPLOY_TEARDOWN_TIMEOUT | 2m | VM-agent max time for deployment environment teardown (stop/start) |
DEPLOY_APPLY_IDLE_TIMEOUT | 15m | VM-agent idle watchdog for deployment apply (no-progress only) |
DEPLOY_BUILD_PUBLISH_TIMEOUT | 20m | VM-agent max time for host build + push + release publish |
DEPLOY_ARTIFACT_DIAL_TIMEOUT | 30s | VM-agent TCP dial timeout for artifact downloads |
DEPLOY_ARTIFACT_TLS_HANDSHAKE_TIMEOUT | 15s | VM-agent TLS handshake timeout for artifact downloads |
DEPLOY_ARTIFACT_RESPONSE_HEADER_TIMEOUT | 60s | VM-agent first-response-header timeout for artifact downloads |
DEPLOY_ARTIFACT_IDLE_TIMEOUT | 2m | VM-agent idle watchdog for artifact body-read progress |
DEFAULT_RESOURCE_EVENT_BUFFER_SIZE | 64 | Capacity of each bounded pressure/Docker event queue; positive integer |
DEFAULT_PSI_POLL_INTERVAL_SECONDS | 10 | Linux memory PSI sampling interval, in seconds |
DEFAULT_CONTAINER_STATS_INTERVAL_SECONDS | 30 | Docker resource statistics sampling interval, in seconds |
DEFAULT_PSI_MEMORY_SOME_WARNING_THRESHOLD | 25 | Warning threshold for the maximum PSI some-memory avg10/avg60 percentage |
DEFAULT_PSI_MEMORY_SOME_CRITICAL_THRESHOLD | 50 | Critical threshold for the maximum PSI some-memory avg10/avg60 percentage |
DEFAULT_PSI_MEMORY_FULL_WARNING_THRESHOLD | 10 | Warning threshold for the maximum PSI full-memory avg10/avg60 percentage |
DEFAULT_PSI_MEMORY_FULL_CRITICAL_THRESHOLD | 25 | Critical threshold for the maximum PSI full-memory avg10/avg60 percentage |
DEFAULT_EVICTION_DEBOUNCE_SECONDS | 30 | Minimum cooldown between ResourceGuard eviction attempts |
DEFAULT_EVICTION_SNAPSHOT_TIMEOUT_SECONDS | 120 | VM-agent pre-stop ResourceGuard eviction snapshot deadline |
DEFAULT_EVICTION_DOCKER_STOP_TIMEOUT_SECONDS | 10 | Grace period passed to docker stop --time during ResourceGuard eviction |
DEFAULT_EVICTION_CALLBACK_RETRY_MAX_SECONDS | 300 | Backoff cap for durable eviction callback retries, in seconds; the operation lease is a lower bound and can exceed this cap. Delivery starts on a later heartbeat |
DEFAULT_EVICTION_RESOLVE_TIMEOUT_SECONDS | 5 | Docker label resolution deadline before ResourceGuard eviction |
RESOURCE_HISTORY_SAMPLE_INTERVAL | 5s | VM-agent retained resource-history cgroup sampling cadence |
RESOURCE_HISTORY_CHUNK_INTERVAL | 15m | VM-agent retained resource-history chunk duration before upload |
RESOURCE_HISTORY_SPOOL_DIR | /var/lib/vm-agent/resource-history | Node-local retry spool for resource-history chunks |
RESOURCE_HISTORY_SPOOL_MAX_BYTES | 20971520 | Max node-local resource-history retry spool bytes |
RESOURCE_HISTORY_UPLOAD_TIMEOUT | 10s | VM-agent deadline for one resource-history upload callback |
RESOURCE_HISTORY_MAX_SAMPLES | 4096 | Max resource samples packed into one uploaded chunk |
Platform Limits
Section titled “Platform Limits”| Variable | Default | Description |
|---|---|---|
MAX_NODES_PER_USER | 10 | Max nodes per user |
TASK_RUN_NODE_CPU_SHARE_BUDGET_PERCENT | 100 | CPU millicore share budget per node for aggregate workspace reservations |
TASK_RUN_NODE_HOST_MEMORY_RESERVE_MB | 512 | Memory reserved for the host/VM agent before admitting occupied-node packing |
TASK_RUN_NODE_DISK_PRESSURE_THRESHOLD_PERCENT | 90 | Fresh node disk telemetry at or above this percent vetoes VM workspace reuse |
TASK_RUN_NODE_METRICS_TTL_MS | 180000 | Freshness window for occupied-node resource telemetry |
TASK_RUN_NODE_CPU_SCORE_WEIGHT_PERCENT | 40 | CPU weight in existing-node load scoring after load average is normalized by vCPU |
TASK_RUN_NODE_MEMORY_SCORE_WEIGHT_PERCENT | 60 | Memory weight in existing-node load scoring |
VM_ADMISSION_CONTROL_MODE | enforce | VM task/session admission mode: off, shadow, or enforce |
VM_ADMISSION_LEASE_TTL_MS | 1200000 (20 min) | Fenced provisioning-claim lease duration |
VM_ADMISSION_RETRY_MIN_MS | 15000 | Minimum retry delay for tasks waiting on VM capacity |
VM_ADMISSION_RETRY_MAX_MS | 60000 | Maximum retry delay for tasks waiting on VM capacity |
VM_ADMISSION_WAIT_TIMEOUT_MS | 7200000 (2 h) | Maximum visible wait for VM capacity before failing the task |
VM_ADMISSION_PROVIDER_COOLDOWN_MS | 600000 (10 min) | Cooldown after provider/account capacity errors such as Hetzner server limits |
VM_ADMISSION_WAKE_BATCH_SIZE | 25 | Maximum waiting TaskRunner DOs nudged by one capacity event |
VM_ADMISSION_DIAGNOSTIC_MESSAGE_MAX_LENGTH | 500 | Maximum provider diagnostic message length stored on admission records |
MAX_AGENT_SESSIONS_PER_WORKSPACE | 10 | Max concurrent agent sessions |
MAX_PROJECTS_PER_USER | 100 | Max projects per user |
MAX_TASKS_PER_PROJECT | 10000 | Max ideas per project |
MAX_TASK_MESSAGE_LENGTH | 16000 | Max task description and reserved prompt length |
RESERVED_TASK_BRANCH_NAME_SEED_MAX_LENGTH | 512 | Max branch-name seed characters accepted by reserved task submissions |
RESERVED_TASK_SOURCE_DISPLAY_NAME_MAX_LENGTH | 512 | Max source display-name characters accepted by reserved task submissions |
RESERVED_TASK_REPOSITORY_ACCESS_FLOW_MAX_LENGTH | 512 | Max repository-access audit flow characters accepted by reserved submissions |
RESERVED_TASK_INITIAL_STATUS_REASON_MAX_LENGTH | 1024 | Max initial status reason characters accepted by reserved task submissions |
RATE_LIMIT_CALLBACK_TOKEN_RENEWAL | 12 | Workspace callback-token renewals accepted per workspace in each window; past it the renewal route answers 429 with Retry-After |
RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS | 3600 | Window for RATE_LIMIT_CALLBACK_TOKEN_RENEWAL |
Durable Object Limits
Section titled “Durable Object Limits”| Variable | Default | Description |
|---|---|---|
MAX_SESSIONS_PER_PROJECT | 10000 | Max chat sessions per project |
MAX_MESSAGES_PER_SESSION | 100000 | Max messages per chat session |
COMMENT_BODY_MAX_LENGTH | 8000 | Max characters per message-anchored comment or reply body |
COMMENT_QUOTE_MAX_LENGTH | 2000 | Max characters preserved from quoted message text |
COMMENT_IDEMPOTENCY_KEY_MAX_LENGTH | 200 | Max clientMutationId length for message-anchored comment writes |
COMMENT_LIST_LIMIT_DEFAULT | 100 | Default page size for comment thread lists |
COMMENT_LIST_LIMIT_MAX | 500 | Max page size for comment thread lists |
COMMENT_THREADS_PER_SESSION_MAX | 1000 | Max message-anchored comment threads per chat session |
COMMENT_REPLIES_PER_THREAD_MAX | 200 | Max replies per message-anchored comment thread |
PROJECT_COMMENT_LIST_LIMIT | 100 | Page size for the project-wide comment inbox |
PROJECT_COMMENT_LIST_MAX | 300 | Max page size for the project-wide comment inbox |
PROJECT_COMMENT_LIST_MAX_BYTES | 4000000 | Byte budget for one project-wide comment inbox response, so a few very long threads cannot exhaust the Durable Object RPC limit |
DOCUMENT_CARD_RAW_OUTPUT_MAX_BYTES | 16384 | Max compact metadata bytes preserved for library document cards |
PROJECT_DATA_TOOL_METADATA_MAX_BYTES | 131072 | Max stored tool_metadata bytes per message before oversized tool content is stripped into bounded metadata |
PROJECT_DATA_STORAGE_TELEMETRY_ENABLED | true | Enables ProjectData databaseSize alarm measurement and D1 telemetry writes |
PROJECT_DATA_STORAGE_LIMIT_BYTES | 10000000000 | Cloudflare SQLite-backed Durable Object storage limit used for ProjectData usage classification |
PROJECT_DATA_STORAGE_MEASURE_INTERVAL_MS | 3600000 | Interval between full per-object ProjectData storage measurements. Each one appends a telemetry history row and evaluates the storage alerts; cleanup passes publish the latest size but never postpone it (measureAndPersistProjectDataStorage) |
PROJECT_DATA_STORAGE_ALERT_INTERVAL_MS | 21600000 | Minimum interval between repeated ProjectData storage observability alerts of the same kind and status. Threshold (warning/critical/degraded) and cleanup-target-unreachable alerts are throttled separately (maybePersistProjectDataStorageAlert) |
PROJECT_DATA_STORAGE_NOTICE_RATIO | 0.6 | ProjectData storage usage ratio classified as notice |
PROJECT_DATA_STORAGE_WARNING_RATIO | 0.8 | ProjectData storage usage ratio classified as warning |
PROJECT_DATA_STORAGE_CRITICAL_RATIO | 0.9 | ProjectData storage usage ratio classified as critical |
PROJECT_DATA_STORAGE_DEGRADED_RATIO | 0.95 | ProjectData storage usage ratio classified as degraded |
PROJECT_DATA_STORAGE_EMERGENCY_TARGET_RATIO | 0.9 | Target usage ratio for explicit superadmin ProjectData emergency purge calls |
PROJECT_DATA_STORAGE_EMERGENCY_BATCH_ROWS | 500 | Oldest activity_events and acp_session_events rows deleted per table per emergency purge batch |
PROJECT_DATA_STORAGE_EMERGENCY_MAX_BATCHES | 4 | Maximum emergency purge batches per explicit call |
PROJECT_DATA_STORAGE_GROWTH_LOOKBACK_DAYS | 7 | Lookback window used to estimate ProjectData bytes/day growth and days to storage limit |
PROJECT_DATA_STORAGE_TELEMETRY_LIST_LIMIT_DEFAULT | 50 | Default row count for admin ProjectData storage telemetry and history lists |
PROJECT_DATA_STORAGE_TELEMETRY_LIST_LIMIT_MAX | 200 | Max accepted row count for admin ProjectData storage telemetry and history lists |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_ENABLED | true | Enables automatic ProjectData cleanup that archives expandable tool_metadata.content payloads to private R2 before stripping them from old message rows. NOT sufficient on its own: setting PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_CUTOFF_CREATED_AT arms the stricter approved-manifest plan, which additionally requires _PLAN_ID, _MANIFEST_KEY, _MANIFEST_SHA256 and all four _MAX_TOTAL_* ceilings. With the cutoff set and any of those blank, cleanup is enabled but builds no plan; the refusal is logged as project_data.tool_payload_cleanup.config_refused with the unmet field names. |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_PLAN_ID | (empty) | Immutable operator plan identifier; required with a fixed cleanup cutoff and persisted with continuation state to reject configuration drift |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MANIFEST_KEY | (empty) | Immutable R2 root key for the verified row-level target manifest; required with a fixed cleanup cutoff |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MANIFEST_SHA256 | (empty) | Expected SHA-256 of the approved target-manifest root; required with a fixed cleanup cutoff |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_BATCH_MANIFEST_MAX_BYTES | 2000000 | Verified approved-plan batch-manifest read/write ceiling; included in the immutable exact-plan fingerprint |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_ROOT_MANIFEST_MAX_BYTES | 1000000 | Verified approved-plan root-manifest read/write ceiling; included in the immutable exact-plan fingerprint |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_ROWS | (empty) | Hard cumulative approved target-row ceiling; required with a fixed cleanup cutoff |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_BYTES | (empty) | Hard cumulative approved projected-reclaim ceiling; required with a fixed cleanup cutoff |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_R2_OPERATIONS | (empty) | Hard cumulative R2 operation ceiling including manifest verification; exact execution transactionally charges the full per-pass allowance before external work and never refunds it |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_TOTAL_WALL_TIME_MS | (empty) | Hard cumulative cleanup wall-time ceiling; exact execution transactionally charges the full per-pass allowance before external work and never refunds it |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_PROJECT_IDS | (empty) | Optional comma-separated project allowlist for automatic cleanup; empty preserves cleanup eligibility for every project, while an emergency rollout can scope elevated budgets to exact projects. One variable, two meanings — see the note below |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_CUTOFF_CREATED_AT | (empty) | Optional fixed exclusive tool-message cutoff in epoch milliseconds; when set, the immutable verified manifest, cumulative ceilings, plan ID, and exactly one allowlisted project are required; invalid values fail closed |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_TRIGGER_RATIO | 0.8 | ProjectData storage usage ratio that starts automatic tool payload archival cleanup even before the retention cadence is due |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_TARGET_RATIO | 0.75 | ProjectData storage usage ratio below which automatic tool payload cleanup stops |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_BATCH_ROWS | 500 | Maximum eligible tool-message candidates processed by one automatic cleanup alarm batch; ordinary selection may examine a larger physical row window bounded by PROJECT_DATA_STORAGE_RELIEF_MEASURE_MAX_BATCH_ROWS |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_BATCH_BYTES | 2097152 | Legacy tool_metadata read budget per automatic pass; an ordinary retention pass may admit its first oversized candidate up to the archive metadata ceiling, while a fixed exact plan requires its row ceiling to fit this budget |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_ROW_BYTES | 1048576 | Maximum single legacy tool_metadata row bytes read into JS by archival cleanup; larger rows fail closed unless the operator deliberately raises this limit |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MIN_SESSION_AGE_DAYS | 7 | Legacy terminal-session age guard retained for storage telemetry compatibility; tool payload archival uses PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_RETENTION_DAYS |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_RECHECK_MS | 86400000 | Delay before the next automatic cleanup alarm batch when more candidates remain; daily by default |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_MAX_SESSIONS_PER_ALARM | 25 | Deprecated legacy terminal-session cleanup knob retained for env compatibility; archival cleanup scans tool-message rows directly and is bounded by rows, bytes, and wall time instead |
PROJECT_DATA_TOOL_PAYLOAD_CLEANUP_WALL_TIME_MS | 20000 | Absolute wall-clock deadline shared by candidate processing and every R2 write/read-back operation in one cleanup pass |
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_MAX_BATCH_ROWS | 500 | Hard row cap for one explicit superadmin manual ProjectData tool payload archival cleanup pass |
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_MAX_BATCH_BYTES | 2097152 | Maximum configurable legacy tool_metadata read budget for one explicit manual pass; the ordinary-path first-row exception applies unless a fixed exact plan binds the row ceiling within this budget |
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_MAX_WALL_TIME_MS | 20000 | Hard wall-clock cap for one explicit manual cleanup pass |
PROJECT_DATA_TOOL_PAYLOAD_MANUAL_CLEANUP_RECHECK_MS | 86400000 | Persisted project-scoped cooldown after an explicit manual cleanup pass; daily by default |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_RETENTION_DAYS | 5 | Message age before expandable tool payload JSON may be archived to private R2 and stripped from the ProjectData DO |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_INTERVAL_MS | 86400000 | Cadence for the retention-driven ProjectData tool payload archival scan |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_R2_PREFIX | project-data/tool-payloads | Private R2 prefix used for archived ProjectData tool payload JSON objects |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_WRITE_TIMEOUT_MS | 5000 | Per-R2 write/read-back verification timeout for automatic archival cleanup; timeout leaves the original payload in ProjectData and defers retry |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_MAX_OPERATIONS | 1500 | Maximum R2 PUT, GET, and body-read operations reserved by one cleanup pass |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_RETRY_DELAY_MS | 300000 | Row-level retry deferral after retryable archive/write failures |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_CHUNK_BYTES | 524288 | R2 chunk size used when archiving legacy tool payload metadata larger than one archive object slice |
PROJECT_DATA_TOOL_PAYLOAD_ARCHIVE_MAX_METADATA_BYTES | 1900000 | Absolute bounded read cap for legacy oversized tool payload metadata; larger rows fail closed and remain in ProjectData |
PROJECT_DATA_STORAGE_RELIEF_MEASURE_BATCH_ROWS | 500 | Default row budget for superadmin ProjectData relief measurement slices |
PROJECT_DATA_STORAGE_RELIEF_MEASURE_MAX_BATCH_ROWS | 5000 | Maximum physical row window accepted for superadmin/preflight relief measurements and examined by one ordinary cleanup candidate-selection slice |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_ENABLED | false | Enables the scheduled, read-only, fixed-cutoff ProjectData tool-payload relief preflight; requires the exact plan, project, and cutoff variables below |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_PLAN_ID | (empty) | Immutable operator plan identifier used to resume one preflight without mixing evidence from another run |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_PROJECT_ID | (empty) | Exact ProjectData project targeted by the enabled preflight |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_CUTOFF_CREATED_AT | (empty) | Fixed exclusive tool-message creation cutoff in epoch milliseconds; required while preflight is enabled |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_BATCH_ROWS | 5000 | Maximum physical chat_messages rowid window examined by one preflight slice before eligibility filtering |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_INTERVAL_MS | 300000 | Persisted minimum interval between preflight slices |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_BATCHES | 100 | Overall claimed-attempt ceiling; an incomplete successful scan becomes truncated at the ceiling, while the final failed attempt becomes failed |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_ROWS | 500000 | Overall physical chat_messages rows-examined ceiling for one preflight plan before eligibility filtering |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_BYTES | 2000000000 | Overall projected net reclaimable-byte evidence ceiling for one preflight plan |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_LEASE_MS | 60000 | D1 claim lease that prevents overlapping slices; must exceed the wall-time budget by the configured lease margin |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_WALL_TIME_MS | 20000 | Absolute per-slice deadline shared by bounded ProjectData measurement and verified R2 manifest writes/read-backs; failures retain the prior cursor and fail closed for retry |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_SLICES_PER_RUN | 1 | Maximum sequential, separately leased slices in one scheduled invocation; later slices bypass only that invocation’s persisted cadence |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_RUN_WALL_TIME_MS | 25000 | Aggregate invocation admission budget used to decide whether another full sequential slice can start; must exceed the per-slice wall-time budget by the return margin |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_LEASE_MARGIN_MS | 5000 | Required lease headroom above the per-slice wall-time budget |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_RETURN_MARGIN_MS | 500 | Required return headroom inside measurement, slice, and aggregate run budgets |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MEASUREMENT_WALL_TIME_MS | 10000 | Explicit ProjectData measurement budget inside one slice; plus the return margin it must not exceed the slice wall budget |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_MAX_STATE_BYTES | 1750000 | Combined D1 JSON byte ceiling for accumulated session and target-batch proof state in one preflight row |
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_ERROR_MAX_LENGTH | 1000 | Character ceiling for a persisted preflight failure diagnostic |
PROJECT_DATA_ARCHIVE_SHARDING_ENABLED | false | Feature switch for exact archive read routing. It is disabled when unset; enabling it alone does not run the unscoped scheduled sweep |
PROJECT_DATA_ARCHIVE_GLOBAL_SWEEP_ENABLED | false | Separate feature switch for the unscoped scheduled archive-sharding sweep. It is disabled when unset; scheduled global migration also requires exact archive routing to be enabled |
PROJECT_DATA_ARCHIVE_GLOBAL_SWEEP_INTERVAL_MS | 86400000 | Persisted cadence gate for unscoped scheduled archive-sharding sweeps; the five-minute handler claims the cadence row only when due. Code fallback is daily; the checked-in wrangler.toml ships 1080000 (18 min), which lands a claim on the fourth five-minute tick (occasionally the fifth, when the archive step starts over two minutes earlier than at the previous claim). A claim can migrate zero or more sessions (usually one, because a real candidate outlasts the PROJECT_DATA_ARCHIVE_WALL_TIME_MS gate checked between candidates), subject to the per-sweep session and message budgets and PROJECT_DATA_ARCHIVE_DAILY_WRITE_BUDGET. The interval allows roughly 72 installation-wide claims/day under the observed cron timing, so cadence alone does not set the drain rate; refused, failed or budget-limited claims publish nothing |
PROJECT_DATA_ARCHIVE_SHARD_COUNT | 128 | Deterministic archive-shard fanout used when assigning terminal sessions to ProjectData archive Durable Objects |
PROJECT_DATA_ARCHIVE_SWEEP_PROJECTS | 1 | Maximum projects selected by one archive-sharding cron pass |
PROJECT_DATA_ARCHIVE_SWEEP_SESSIONS | 10 | Hard ceiling on sessions (in-flight plus new candidates) one archive-sharding pass may process; the checked-in wrangler.toml ships 8 |
PROJECT_DATA_ARCHIVE_SWEEP_MESSAGE_BUDGET | 20000 | Cumulative session_summaries.message_count of new candidates one pass may journal; candidates are selected largest first, and a single session above the budget is still selected alone; the checked-in wrangler.toml ships 10000 |
PROJECT_DATA_ARCHIVE_SESSION_GRACE_MS | 604800000 | Minimum terminal-session age before archive-sharding may consider a session |
PROJECT_DATA_ARCHIVE_PRECOPY_REFUSAL_RETRY_MS | 604800000 | How long a session the root object refused at prepare (pre-copy invariant, e.g. active_session_state) stays out of unscoped sweep selection; journal frozen/precopy_refused, location back at root; a named-session canary bypasses it |
PROJECT_DATA_ARCHIVE_FAILED_RETRY_DELAY_MS | 3600000 | Minimum age of a failed archive journal before an unscoped sweep reclaims it, so PROJECT_DATA_ARCHIVE_POISON_AFTER_ATTEMPTS spans hours rather than consecutive cadence ticks; 0 disables it and a named-session canary bypasses it |
PROJECT_DATA_ARCHIVE_CHUNK_ROWS | 500 | Maximum rows exported in one idempotent transcript archive chunk |
PROJECT_DATA_ARCHIVE_CHUNK_BYTES | 16777216 | Maximum bytes exported in one archive chunk; clamped below Cloudflare’s 32MiB RPC ceiling |
PROJECT_DATA_ARCHIVE_HASH_PAGE_ROWS | 500 | Rows read per statement while streaming a session’s terminal-version hash inside the ProjectData Durable Object; bounds object memory by page size instead of session size |
PROJECT_DATA_ARCHIVE_LEASE_MS | 300000 | D1 CAS journal lease duration for archive-sharding work |
PROJECT_DATA_ARCHIVE_WALL_TIME_MS | 5000 | Soft wall-clock budget for one archive-sharding cron pass, checked between candidates (one large session may run past it). Code fallback 5 s; the checked-in wrangler.toml ships 10000 after the 2026-09-08 billing firebreak |
PROJECT_DATA_ARCHIVE_ROLLOUT_LIST_LIMIT_DEFAULT | 25 | Default row limit for superadmin archive-sharding rollout inspection and failed/poisoned/frozen migration list endpoints |
PROJECT_DATA_ARCHIVE_ROLLOUT_LIST_LIMIT_MAX | 100 | Maximum accepted row limit for D1-only superadmin archive-sharding rollout inspection endpoints. Frozen-intent DO inspection is capped separately. |
PROJECT_DATA_ARCHIVE_FROZEN_INTENT_INSPECTION_LIMIT_DEFAULT | 5 | Default row limit for GET /archive-sharding/frozen-intents; each row may perform source and target ProjectData DO RPCs, so this endpoint intentionally has a smaller hot-object fan-out budget |
PROJECT_DATA_ARCHIVE_FROZEN_INTENT_INSPECTION_LIMIT_MAX | 10 | Maximum accepted frozen-intent detail inspection limit; configured values above 10 are clamped to the reviewed hot-DO fan-out ceiling and request limits above the configured max are rejected |
PROJECT_DATA_ARCHIVE_MANUAL_CANARY_MAX_SESSIONS | 5 | Maximum sessions a superadmin scoped manual archive-sharding canary request may select; default requests select one session and dry-run by default |
PROJECT_DATA_ARCHIVE_MANUAL_CANARY_MAX_WALL_TIME_MS | 15000 | Maximum wall-clock budget accepted by the superadmin scoped manual archive-sharding canary endpoint |
PROJECT_DATA_ARCHIVE_ROLLOUT_WARNING_EXAMPLES_MAX | 5 | Maximum malformed-row warning examples returned by archive rollout read/list endpoints |
PROJECT_DATA_ARCHIVE_ROLLOUT_WARNING_REASON_MAX_LENGTH | 300 | Maximum characters per malformed-row warning reason returned by archive rollout read/list endpoints |
PROJECT_DATA_ARCHIVE_POISON_AFTER_ATTEMPTS | 3 | Failed archive-sharding attempts before the migration is poisoned and the project circuit breaker opens. Each opening sends every active superadmin one Operational Failure notification and records one /admin/errors entry (alertProjectDataArchiveBreakerOpened) |
PROJECT_DATA_ARCHIVE_R2_PREFIX | project-data/session-archives | Private R2 prefix for terminal-session archive recovery chunks and manifests |
PROJECT_DATA_ARCHIVE_SEARCH_MAX_OWNERS | 4 | Archive-shard owner batch size for one project-wide message-search continuation (maximum 64; invalid/out-of-range values use 4); callers continue until coverage is complete |
PROJECT_DATA_ARCHIVE_SEARCH_CONCURRENCY | 4 | Concurrent archive-owner queries in one batch (maximum 16; invalid/out-of-range values use 4) |
PROJECT_DATA_ARCHIVE_SEARCH_REPAIR_SESSIONS | 1 | Incomplete published archive sessions advanced per owner query (maximum 16); remaining repair stays explicit and resumable |
PROJECT_DATA_ARCHIVE_SEARCH_REPAIR_CHUNKS | 1 | Immutable compact R2 chunks consumed per session repair step (maximum 64) |
PROJECT_DATA_ARCHIVE_SEARCH_CONTINUATION_TTL_MS | 900000 | Lifetime of a signed, query-bound complete-history continuation (maximum 24 hours) |
PROJECT_DATA_ARCHIVE_SEARCH_CURSOR_MAX_BYTES | 1048576 | Maximum continuation bytes accepted before decoding or emitted after signing (maximum 4 MiB) |
PROJECT_DATA_ARCHIVE_SEARCH_ERROR_LIMIT | 20 | Maximum entries retained independently in each public execution- and index-error collection (maximum 100); full details remain in server logs |
PROJECT_DATA_SEARCH_FTS_CANDIDATE_LIMIT | 2000 | Newest full-text matches ranked per ProjectData message search; when more match, results report rootSearch.ftsCandidatesTruncated |
PROJECT_DATA_SEARCH_FTS_SCAN_LIMIT | 20000 | Full-text index entries inside a session’s rowid span that a session-scoped search examines, newest first, to collect that session’s newest matches; never smaller than PROJECT_DATA_SEARCH_FTS_CANDIDATE_LIMIT |
PROJECT_DATA_SEARCH_KEYWORD_SCAN_ROW_LIMIT | 50000 | Newest raw messages the keyword fallback scans for not-yet-indexed text; older rows are reported as rootSearch.keywordScanTruncated |
PROJECT_DATA_ALARM_SECTION_GATING_ENABLED | true | Run only the ProjectData alarm sections that are due each tick; false runs every section every tick |
PROJECT_DATA_ALARM_FULL_RUN_INTERVAL_MS | 900000 | Maximum interval between ProjectData alarm ticks that run every section regardless of due times |
PROJECT_DATA_ALARM_DUE_TOLERANCE_MS | 2000 | A ProjectData alarm section due within this many ms of the tick runs in that tick |
PROJECT_DATA_ALARM_SLOW_SECTION_MS | 1000 | ProjectData alarm sections at or above this wall time log project_data.alarm.section_slow |
WORKSPACE_RESOURCE_RAW_RETENTION_DAYS | 90 | Retention for immutable raw resource-history gzip chunks in the private archive R2 binding |
WORKSPACE_RESOURCE_SUMMARY_RETENTION_DAYS | 180 | Retention for bounded D1 workspace resource summary rows |
WORKSPACE_RESOURCE_UNCOMPRESSED_MAX_BYTES | 8388608 | Max decoded resource chunk JSON bytes accepted/read before rejecting detail payloads |
WORKSPACE_RESOURCE_METADATA_MAX_BYTES | 8192 | Max summary or completeness JSON bytes stored in D1 for one resource chunk/summary |
WORKSPACE_RESOURCE_TOOL_NAME_MAX_BYTES | 256 | Max UTF-8 bytes retained and returned for one resource-history tool name; longer metadata names are truncated at a complete Unicode character |
WORKSPACE_RESOURCE_UPLOAD_MAX_BYTES | 2097152 | Max compressed resource-history chunk bytes accepted by the callback upload route |
WORKSPACE_RESOURCE_DETAIL_MAX_POINTS | 720 | Max samples returned by one raw detail read after spike-preserving downsampling |
WORKSPACE_RESOURCE_LIST_LIMIT | 24 | Max resource-history chunk index rows returned for one contextual read |
WORKSPACE_RESOURCE_TIMELINE_MAX_CHUNKS | 1000 | Max chunks the session resource timeline lists; the oldest beyond it are left out and disclosed as omittedChunkCount |
WORKSPACE_RESOURCE_ROLLUP_BUCKET_MS | 60000 | Width of each per-chunk rollup bucket computed on upload |
WORKSPACE_RESOURCE_ROLLUP_MAX_BUCKETS | 60 | Max rollup buckets stored per chunk; the bucket width widens in whole multiples to fit |
WORKSPACE_RESOURCE_CLEANUP_BATCH_SIZE | 50 | Max expired resource-history chunks and summaries processed per scheduled cleanup sweep |
WORKSPACE_RESOURCE_OBJECT_CLEANUP_LIMIT | 5000 | Max resource-history R2 objects deleted when a project or workspace is deleted; deletion paginates until the prefix is empty or this safety budget is reached |
A completed fixed tool-payload plan writes an explicit terminal marker rather than an indefinite timestamp. Later alarm or manual entry points no-op after validating the same immutable fingerprint. A different fingerprint also fails closed; there is deliberately no automatic reset or supersession, so an emergency runbook must include a reviewed follow-up reset before ordinary cleanup is expected to resume.
Archive-sharding rollout is deliberately staged. The scheduled coordinator in
apps/api/src/scheduled/project-data-archive-sharding.ts still does no work unless
PROJECT_DATA_ARCHIVE_GLOBAL_SWEEP_ENABLED=true and exact archive routing is active through
PROJECT_DATA_ARCHIVE_SHARDING_ENABLED=true. Both switches are ordinary Worker [vars]: a GitHub
Environment variable of the same name replaces the checked-in wrangler.toml value at deploy time,
and the deploy log prints every such override that differs from wrangler.toml. When the sweep is
enabled, each cron.completed log carries projectDataArchiveShardingSkipReason (disabled,
exact_routing_disabled, missing_r2_binding, cadence_not_due, cadence_running,
cadence_unavailable) or the tick’s
selected/migrated/refused/failed counts. A session the root object refuses before any copy
(for example one that still holds an active session_state row) is returned to root in the same
tick with a frozen/precopy_refused journal and is skipped by the sweep for
PROJECT_DATA_ARCHIVE_PRECOPY_REFUSAL_RETRY_MS.
The shipped cadence (1080000, 18 minutes) and wall budget (10000) were sized against the SAM root object
(~3,350 backlog sessions plus 4-233 newly terminal sessions per day): one tick moves one
non-trivial session because the wall-time break fires between candidates, so a daily tick can
never keep up. Idle ticks cost two D1 statements. Per-tick copy work is bounded by the session ceiling and the
message budget; the wall budget is a soft gate checked between candidates, so one archive can run
past it. The sweep runs after every lifecycle sweep
in the cron chain so that budget cannot delay them. failed journals wait
PROJECT_DATA_ARCHIVE_FAILED_RETRY_DELAY_MS before a sweep retries them, so the three-attempt poison
budget still spans hours at the shorter cadence. The frozen-intent inspection route skips
precopy_refused rows (nothing exists on either object to inspect) and the problem-migrations list
sorts them after rows that need a human. Before changing the global sweep switch, operators
use the superadmin routes in apps/api/src/routes/admin/project-data-storage.ts to inspect one
project’s D1 journal/location/breaker state and run a scoped dry-run canary for a specific project
and optional session. Dry-runs work while both rollout switches are false and never call source deletion RPCs.
Non-dry scoped canaries fail closed unless exact archive routing is active, because publishing or
deleting source rows while exact reads still resolve to root can render conversations empty. Failed,
poisoned, or frozen rows are inspected through the frozen-intent route (and listed under Admin →
Storage → Problem migrations). A migration that never reached source deletion is cleared with
Abandon (the page’s button, or POST .../migrations/:migrationId/abandon); one past source deletion
is recovered through the copy-back helper with exact archive routing enabled. Both take an explicit
operator reason, and neither closes the project’s circuit breaker.
Safe operator sequence for ProjectData storage relief:
- Inspect current telemetry with
GET /api/admin/project-data/storage?projectId=<projectId>. Before building an approval manifest, freeze competing source mutations: disable tool-payload, event-log, and grouped/FTS cleanup, and leave both archive-sharding switches off. - For an emergency exact plan, configure the immutable
PROJECT_DATA_STORAGE_RELIEF_PREFLIGHT_{PLAN_ID,PROJECT_ID,CUTOFF_CREATED_AT}scope plus row, byte, attempt, slice, run-admission, state, and manifest limits. Enable the scheduled preflight only for that scope. - Wait for a terminal
project_data_storage_relief_preflightsrow. Accept only a complete or deliberately bounded truncated result with exact totals and a non-nulltarget_manifest_key, byte count, and SHA-256. The preflight reads every batch and root manifest back from R2 and verifies bytes and SHA-256 before persisting these proofs; independently inspect the root and referenced batch proofs before approval. - Disable preflight and keep cleanup frozen. Present the exact project, cutoff, row targets, manifest keys/hashes, cumulative ceilings, switches, failure behavior, expected relief, observation window, and stop conditions for explicit human approval. Broad rollout approval does not substitute for this plan approval.
- Only after approval, configure cleanup with the same plan/project/cutoff, exact root key/SHA-256, and cumulative row/byte/R2-operation/wall-time ceilings. Execute only approved bounded passes, using
POST /api/admin/project-data/storage/:projectId/tool-payload-cleanupwith a non-emptyreasonand uniqueidempotencyKeywhen an explicit manual slice is required. - Each payload archive object/chunk is written to an immutable content-addressed R2 key and read back for byte/SHA-256 verification before a guarded transaction changes only
chat_messages.tool_metadata;chat_messages.contentis never changed. Stop on any manifest, archive, source-hash, scope, cap, timeout, or circuit-breaker uncertainty. - After every pass, inspect
terminationReason, before/afterdatabaseSize, reclaimed bytes, row counts, cursor/recheck state, and cooldown, then re-measure withPOST /api/admin/project-data/storage/:projectId/measure. Stop at the approved target or any approved stop condition. - A completed fixed plan is terminal. Before ordinary retention cleanup resumes, use a separately reviewed follow-up reset; do not reuse or silently supersede the emergency plan fingerprint.
| Variable | Default | Description |
|---|---|---|
PROJECT_DATA_MATERIALIZATION_PAGE_ROWS | 500 | Tokens read into memory by one chat-search materialization SELECT, bounding Durable Object isolate memory |
PROJECT_DATA_MATERIALIZATION_MAX_ROWS_PER_PASS | 5000 | Tokens one chat-search materialization pass indexes before leaving the rest to the next pass |
PROJECT_DATA_MATERIALIZATION_MAX_GROUP_CHARS | 65536 | Grouped-row size past which a continuing run starts a new row instead of rewriting the accumulated content |
PROJECT_DATA_MATERIALIZATION_SWEEP_LIMIT | 50 | Sessions indexed by one chat-search materialization backfill call |
PROJECT_DATA_MATERIALIZATION_SWEEP_SCAN_LIMIT | 500 | Sessions examined by one chat-search materialization backfill call |
PROJECT_DATA_GROUPED_FTS_CLEANUP_ENABLED | false | Enables cleanup of old terminal-session grouped message rows and their FTS entries under storage pressure. wrangler.toml ships true, so this is ON; a cleaned session falls back to keyword search and is never re-indexed |
PROJECT_DATA_GROUPED_FTS_CLEANUP_TRIGGER_RATIO | 0.9 | ProjectData usage ratio that starts grouped/FTS derived-data cleanup |
PROJECT_DATA_GROUPED_FTS_CLEANUP_TARGET_RATIO | 0.85 | ProjectData usage ratio below which grouped/FTS cleanup stops |
PROJECT_DATA_GROUPED_FTS_CLEANUP_BATCH_SESSIONS | 2 | Maximum terminal sessions examined by one grouped/FTS slice; candidate selection pages IDs first and reads at most one extra session for continuation |
PROJECT_DATA_GROUPED_FTS_CLEANUP_BATCH_ROWS | 1000 | Maximum grouped message rows deleted by one grouped/FTS canary slice |
PROJECT_DATA_GROUPED_FTS_CLEANUP_BATCH_BYTES | 4194304 | Maximum grouped message content bytes deleted by one grouped/FTS canary slice |
PROJECT_DATA_GROUPED_FTS_CLEANUP_MIN_SESSION_AGE_DAYS | 7 | Minimum terminal-session age before grouped/FTS derived rows may be cleaned |
PROJECT_DATA_GROUPED_FTS_CLEANUP_RECHECK_MS | 300000 | Delay before the next grouped/FTS cleanup slice when more candidates remain. Also the back-off after an overload/reset error, measured from when that error was recorded |
PROJECT_DATA_GROUPED_FTS_CLEANUP_WALL_TIME_MS | 5000 | Soft wall-clock budget for one grouped/FTS cleanup slice |
PROJECT_DATA_GROUPED_FTS_CLEANUP_WALL_UNSAFE_RATIO | 0.98 | Refuses grouped/FTS cleanup writes when the object is too close to the configured storage limit |
PROJECT_DATA_GROUPED_FTS_CLEANUP_WEAK_RECLAIM_BYTES | 1 | Stops grouped/FTS cleanup if a slice deletes rows but databaseSize does not drop by at least this many bytes |
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_MAX_ROWS | 10000 | Ceiling on grouped rows one superadmin grouped/FTS wall-recovery call may prune |
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_MAX_BYTES | 33554432 | Ceiling on grouped content bytes one grouped/FTS wall-recovery call may prune |
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_MAX_SESSIONS | 500 | Ceiling on sessions one grouped/FTS wall-recovery call may consider, and on its skipSessionIds length |
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_TRANSACTION_ROWS | 500 | Grouped rows pruned per wall-recovery transaction |
PROJECT_DATA_GROUPED_FTS_WALL_RECOVERY_TRANSACTION_BYTES | 8388608 | Grouped content bytes held in memory per wall-recovery transaction |
PROJECT_DATA_EVENT_LOG_CLEANUP_ENABLED | true | Enables automatic deletion of old low-value terminal-session activity_events and terminal ACP event history when storage remains above the cleanup target |
PROJECT_DATA_EVENT_LOG_CLEANUP_BATCH_ROWS | 500 | Maximum terminal activity_events rows and terminal acp_session_events rows deleted per automatic cleanup alarm batch |
PROJECT_DATA_EVENT_LOG_CLEANUP_MIN_SESSION_AGE_DAYS | 7 | Minimum terminal-session age before automatic event-log cleanup may delete its activity/ACP event history |
PROJECT_DATA_EVENT_LOG_CLEANUP_RECHECK_MS | 86400000 | Delay before the next terminal event-log cleanup alarm batch when more candidates remain; daily by default |
PROJECT_EVENT_MAX_ACTIVE_SUBSCRIPTIONS_PER_PROJECT | 200 | Maximum active durable event subscriptions in one ProjectData object |
PROJECT_EVENT_FILTER_MAX_VALUES_PER_FIELD | 20 | Maximum exact/set values accepted for one v1 event filter field |
PROJECT_EVENT_FILTER_MAX_MATCH_KEYS | 100 | Maximum deterministic match keys compiled for one event subscription |
PROJECT_EVENT_FILTER_MAX_STRING_BYTES | 160 | Maximum bytes for event source/type/subject, filter values, delivery keys, fingerprints, and event idempotency strings |
PROJECT_EVENT_METADATA_MAX_BYTES | 8192 | Maximum normalized event metadata JSON bytes stored in ProjectData |
PROJECT_EVENT_METADATA_MAX_DEPTH | 4 | Maximum nesting depth for normalized event metadata |
PROJECT_EVENT_METADATA_MAX_KEYS | 64 | Maximum total object keys in normalized event metadata |
PROJECT_EVENT_METADATA_MAX_ARRAY_ITEMS | 64 | Maximum array items in normalized event metadata |
PROJECT_EVENT_DISPLAY_MAX_BYTES | 4096 | Maximum deterministic untrusted event display JSON bytes |
PROJECT_EVENT_DISPLAY_MAX_LABELS | 12 | Maximum labels accepted in event display data |
PROJECT_EVENT_RAW_PAYLOAD_REF_MAX_BYTES | 512 | Maximum optional raw-payload reference JSON bytes; raw webhook bodies are not persisted in ProjectData |
PROJECT_EVENT_REASON_MAX_BYTES | 1024 | Maximum diagnostic reason bytes on subscriptions, batches, and attempts |
PROJECT_EVENT_MAX_MATCHES_PER_EVENT | 100 | Maximum durable subscription matches written for one admitted normalized event |
PROJECT_EVENT_DELIVERY_BATCH_MAX_EVENTS | 50 | Maximum event matches in one durable delivery batch |
PROJECT_EVENT_DELIVERY_ATTEMPT_MAX_PER_BATCH | 10 | Maximum recorded delivery attempts for one event delivery batch |
PROJECT_EVENT_SCHEDULE_MAX_SCHEDULES | 128 | Maximum active schedules per project |
PROJECT_EVENT_SCHEDULE_MAX_WATCHES | 64 | Maximum active or paused standing watches per project |
PROJECT_EVENT_SCHEDULE_MAX_RETAINED_SCHEDULES | 4096 | Maximum retained schedule records of all states per project; cancellation does not free retained capacity |
PROJECT_EVENT_SCHEDULE_MAX_RETAINED_WATCHES | 256 | Maximum retained watch records of all states per project; revocation does not free retained capacity |
PROJECT_EVENT_SCHEDULE_PROMPT_MAX_BYTES | 32768 | Maximum scheduled action prompt size in UTF-8 bytes; start_session admission also validates MAX_TASK_MESSAGE_LENGTH |
PROJECT_EVENT_SCHEDULE_MAX_HORIZON_MS | 2592000000 | Maximum future scheduling horizon (30 days) |
PROJECT_EVENT_SCHEDULE_LATE_GRACE_MS | 86400000 | Maximum late admission grace after due time |
PROJECT_EVENT_SCHEDULE_DELIVERY_TTL_MS | 86400000 | Maximum durable message lifetime, also bounded by schedule expiry |
PROJECT_EVENT_SCHEDULE_SWEEP_BATCH_SIZE | 20 | Maximum schedule or watch candidates processed per alarm pass |
PROJECT_EVENT_SCHEDULE_CLAIM_LEASE_MS | 60000 | Finite admission/submission claim lease |
PROJECT_EVENT_SCHEDULE_RETRY_BASE_MS | 30000 | Retry and accepted execution reconciliation interval |
PROJECT_EVENT_SCHEDULE_MAX_ATTEMPTS | 8 | Maximum uncertain task submission attempts |
PROJECT_EVENT_SCHEDULE_MAX_DEFERRAL_MS | 86400000 | Maximum deferral before a new task may initially start |
PROJECT_EVENT_WATCH_COOLDOWN_MIN_MS | 60000 | Minimum time between standing watch executions |
PROJECT_EVENT_WATCH_MAX_EXECUTIONS | 100 | Maximum configurable executions per standing watch |
PROJECT_EVENT_WATCH_MAX_CONCURRENT | 3 | Maximum configurable concurrent actions per standing watch |
PROJECT_EVENT_CHANNEL_MAX_CHANNELS | 128 | Maximum catalog generations per project |
PROJECT_EVENT_CHANNEL_MESSAGE_MAX_BYTES | 4096 | Maximum channel message UTF-8 bytes, also subject to canonical metadata limits |
PROJECT_EVENT_CHANNEL_NAME_MAX_BYTES | 64 | Maximum channel name bytes |
PROJECT_EVENT_CHANNEL_PUBLISH_WINDOW_MS | 60000 | Per-project fixed publish window in milliseconds |
PROJECT_EVENT_CHANNEL_PUBLISH_MAX_PER_WINDOW | 120 | Maximum newly committed channel publishes per project window; retained replays do not consume quota |
PROJECT_EVENT_CHANNEL_CURSOR_TTL_MS | 3600000 | History cursor and unfinished catch-up lifetime in milliseconds; continuation never extends it |
PROJECT_EVENT_CHANNEL_CATALOG_IDLE_TTL_MS | 2592000000 | Minimum idle time before reclaiming an empty catalog generation without live catch-up |
AGENT_MESSAGE_CHANNELS_ENABLED | false | Code fallback is off; managed deployment sets true. Send notify/deliver agent messages over SAM-managed agent-dm.* pair channels. Effective only while PROJECT_EVENT_WAKE_ENABLED and durable prompt delivery are on |
AGENT_MESSAGE_CHANNEL_MAX_CHANNELS | 1024 | Maximum agent-dm.* pair channels per project, separate from PROJECT_EVENT_CHANNEL_MAX_CHANNELS |
AGENT_MESSAGE_SUBSCRIPTION_ROTATION_GRACE_MS | 300000 | Replace a managed pair subscription this close to the end of its wake lifetime when it owes no pending wake |
AGENT_MESSAGE_MAX_ACTIVE_SUBSCRIPTIONS | 100 | Share of PROJECT_EVENT_MAX_ACTIVE_SUBSCRIPTIONS_PER_PROJECT that SAM-managed agent-message subscriptions may hold; idle ones on other pairs are released first |
PROJECT_EVENT_LIST_LIMIT | 50 | Default ProjectData event-subscription list/status page size |
PROJECT_EVENT_LIST_MAX | 200 | Maximum ProjectData event-subscription list/status page size |
PROJECT_EVENT_SUBSCRIPTION_EVENT_CURSOR_MAX_LENGTH | 512 | Maximum opaque list_subscription_events cursor length accepted by ProjectData pull delivery |
PROJECT_EVENT_RECENT_STATUS_LIMIT | 50 | Maximum rows returned per section in recent event-subscription status inspection |
PROJECT_EVENT_RETENTION_DAYS | 30 | Retention window for terminal ProjectData event-subscription records |
PROJECT_EVENT_RETENTION_BATCH_ROWS | 500 | Maximum old rows pruned per event-subscription retention category per pass |
PROJECT_EVENT_SOURCE_OUTBOX_BATCH_ROWS | 25 | Maximum producer-side outbox row mutations per trigger cleanup pass. Must be an integer of at least 2 for claim plus settlement; invalid values reject before any writes. Per-call budget overrides have the same minimum. |
PROJECT_EVENT_SOURCE_OUTBOX_MAX_ATTEMPTS | 8 | Maximum ProjectData admission attempts before a source intent is marked permanent_failed |
PROJECT_EVENT_SOURCE_OUTBOX_TTL_MS | 86400000 | Maximum lifetime for a producer-side event admission intent |
PROJECT_EVENT_SOURCE_OUTBOX_RETRY_BASE_MS | 30000 | Initial retry delay for failed source event admission |
PROJECT_EVENT_SOURCE_OUTBOX_RETRY_MAX_MS | 1800000 | Maximum retry delay for failed source event admission |
PROJECT_EVENT_SOURCE_OUTBOX_PROCESSING_LEASE_MS | 60000 | Lease before an in-flight source event admission intent can be claimed by reconciliation |
PROJECT_EVENT_RETENTION_INTERVAL_MS | 86400000 | Default delay before the next ProjectData event retention alarm after a complete pass; daily by default |
PROJECT_EVENT_RETENTION_MIN_ALARM_DELAY_MS | 60000 | Minimum delay before a follow-up event retention alarm when more bounded work remains or a due checkpoint is re-armed |
PROJECT_EVENT_WAKE_ENABLED | false | Enables same-chat ProjectData event wake materialization for v2 existing_session_prompt and runtime_interrupt subscriptions; opt in with true; unset or false leaves event delivery pull-only |
PROJECT_EVENT_WAKE_MATERIALIZATION_MIN_ALARM_DELAY_MS | 1000 | Minimum delay before a ProjectData alarm retries wake materialization after due matches are present |
PROJECT_EVENT_WAKE_MATERIALIZATION_BACKOFF_BASE_MS | 5000 | Initial persisted backoff for a failed event-wake materialization or retention alarm phase |
PROJECT_EVENT_WAKE_MATERIALIZATION_BACKOFF_MAX_MS | 300000 | Maximum persisted backoff for repeated event-wake materialization or retention alarm failures |
PROJECT_EVENT_WAKE_PROMPT_TTL_MS | 86400000 | Hard lifetime for queued same-chat event wake prompts before physical delivery is rejected; also bounds how long an undelivered wake holds its chat |
PROJECT_EVENT_WAKE_READ_GRACE_MS | 86400000 | Read/ack grace retained on accepted event-wake batches after natural subscription expiry |
PROJECT_EVENT_WAKE_TARGET_COOLDOWN_MS | 30000 | Delay before retrying an event wake for a chat whose transcript or project mailbox is at capacity. A chat with an undelivered wake is skipped until that wake is accepted, without this cooldown |
PROJECT_EVENT_WAKE_SUBSCRIPTION_COOLDOWN_MS | 30000 | Per-subscription cooldown after a wake batch is queued |
PROJECT_EVENT_WAKE_SUBSCRIPTION_LIFETIME_MS | 86400000 | Maximum operational lifetime for same-chat wake delivery on a v2 event subscription when the requested subscription expiry is absent or farther out |
PROJECT_EVENT_WAKE_MAX_PER_SUBSCRIPTION | 50 | Maximum same-chat wake prompt batches materialized for one v2 event subscription |
PROJECT_EVENT_SOURCE_OUTBOX_SWEEP_WALL_MS | 2500 | Per-pass wall-clock budget for producer-side outbox reconciliation before the sweep yields |
PROJECT_EVENT_SOURCE_OUTBOX_ADMISSION_TIMEOUT_MS | 5000 | Per-intent timeout for the ProjectData admission call; timed-out calls are treated as ambiguous and retried through the fenced outbox |
PROJECT_EVENT_SOURCE_OUTBOX_TERMINAL_RETENTION_MS | 604800000 | Retention window for terminal outbox rows before the bounded cleanup removes admitted, expired, and permanent_failed history |
MESSAGE_SIZE_THRESHOLD | 102400 | Max message size in bytes |
ACTIVITY_RETENTION_DAYS | 90 | Days to retain activity events |
SESSION_IDLE_TIMEOUT_MINUTES | 60 | Idle session timeout |
SESSION_ACTIVITY_STALE_THRESHOLD_MS | 300000 (5 min) | Threshold before stale working activity is checked against authoritative SessionHost inventory |
SESSION_ACTIVITY_PROBE_TIMEOUT_MS | 5000 (5 s) | Timeout for the vm-agent session-activity probe. Background control-loop budget — deliberately far below the interactive node-agent timeout |
SESSION_ACTIVITY_PROBE_MAX_ATTEMPTS | 3 | Consecutive unreachable probes after which a stale working state is quarantined until an authoritative report refreshes it |
SESSION_ACTIVITY_PROBE_MAX_CANDIDATES | 10 | Stale-activity candidates probed per ProjectData alarm pass |
DO_SUMMARY_SYNC_DEBOUNCE_MS | 5000 | Debounce for DO-to-D1 summary sync |
SESSION_INDEX_MAX_ROWS | 1000 | Sessions mirrored into the D1 session_summaries index per project. A project holding more is recorded as incomplete and its chat sidebar reads fall back to the Durable Object |
SESSION_INDEX_MAX_STALENESS_MS | 900000 (15 min) | How stale the session index may be before the per-project sidebar list stops trusting it and falls back to the Durable Object |
CREDENTIAL_LIMIT_WARNING_PERCENT | 75 | Advisory warning threshold for credential quota utilization samples |
CREDENTIAL_LIMIT_CRITICAL_PERCENT | 90 | Advisory critical threshold for credential quota utilization samples |
CREDENTIAL_LIMIT_MAX_OBSERVATIONS_PER_REPORT | 16 | Maximum credential-limit observations accepted from one VM usage callback or proxy report |
CREDENTIAL_LIMIT_TRANSITION_RECOMPUTE_ATTEMPTS | 4 | Maximum predecessor-CAS recomputes for one credential-limit observation when concurrent samples update the same window |
CREDENTIAL_LIMIT_ADMISSION_MAX_ACTIVE_PER_PROJECT | 1000 | Maximum active credential-limit event admissions per project (range: 1–100000) |
CREDENTIAL_LIMIT_READ_MAX_ROWS | 200 | Maximum credential usage-limit window rows returned by one read request (GET /api/projects/:id/credential-limits, GET /api/credentials/limits, MCP get_credential_limits) |
CREDENTIAL_LIMIT_ADMISSION_RETRY_BATCH_SIZE | 25 | Maximum pending credential-limit event admissions retried per batch (range: 1–500) |
CREDENTIAL_LIMIT_ADMISSION_RETENTION_DAYS | 30 | Retention of inactive credential-limit window observations in days (range: 1–365) |
CREDENTIAL_LIMIT_USAGE_CALLBACK_MAX_BODY_BYTES | 32768 | Maximum raw JSON bytes accepted for one authenticated VM usage callback before schema validation |
CREDENTIAL_LIMIT_USAGE_CALLBACK_RATE_LIMIT_RPM | 120 | Authenticated VM usage callbacks accepted per session in each rate-limit window |
CREDENTIAL_LIMIT_USAGE_CALLBACK_RATE_LIMIT_WINDOW_SECONDS | 60 | Usage callback rate-limit window in seconds |
CREDENTIAL_LIMIT_OBSERVATION_MAX_AGE_MS | 86400000 | Oldest provider observation timestamp accepted for credential-limit telemetry |
CREDENTIAL_LIMIT_OBSERVATION_FUTURE_SKEW_MS | 300000 | Future clock skew accepted for credential-limit observation timestamps |
CREDENTIAL_LIMIT_RESET_MAX_FUTURE_MS | 691200000 | Maximum future provider reset timestamp accepted for credential-limit telemetry |
CREDENTIAL_LIMIT_SUPPORTED_PROVIDERS | anthropic,openai,opencode | Comma-separated allowlist of providers accepted by credential-limit telemetry |
CREDENTIAL_LIMIT_SUPPORTED_SOURCES | built-in sources | Comma-separated allowlist of VM-agent and AI-proxy telemetry source identifiers accepted by credential-limit telemetry |
CREDENTIAL_LIMIT_SUPPORTED_WINDOW_TYPES | built-in windows | Comma-separated allowlist of provider quota window identifiers accepted by credential-limit telemetry |
AI_PROXY_REQUEST_BODY_MAX_BYTES | 1048576 | Maximum raw JSON bytes accepted by OpenAI-compatible, Anthropic-native, and passthrough AI proxy request endpoints before request validation |
AI_PROXY_ALLOWED_MODELS | platform catalog | Comma-separated models the AI proxy serves on every route that spends platform credentials, the native Anthropic endpoint included; any other model returns 400 |
Ordinary ProjectData storage alarms record O(1) databaseSize telemetry and
bounded cleanup row/byte counters. Category breakdown scans are reserved for
explicit/admin measurement paths so hot ProjectData alarms do not delay
lifecycle bookkeeping.
Durable Object Retry
Section titled “Durable Object Retry”| Variable | Default | Description |
|---|---|---|
DO_RETRY_MAX_ATTEMPTS | 8 | Max attempts for transient Durable Object RPC reset/overload errors |
DO_RETRY_BASE_DELAY_MS | 100 | Base retry delay in milliseconds for transient Durable Object RPC failures |
DO_RETRY_MAX_DELAY_MS | 250 | Max per-attempt retry delay for transient Durable Object RPC failures |
DO_RETRY_CONNECTION_LOST_MAX_ATTEMPTS | 3 | Attempts (capped at DO_RETRY_MAX_ATTEMPTS) for an idempotent ProjectData read whose connection to the object was lost or that encounters SQLITE_NOMEM. Mutations never retry SQLITE_NOMEM. An idempotent read that exhausts its attempts on either error or a CPU-limit reset returns 503 PROJECT_DATA_UNAVAILABLE |
PROJECT_DATA_ENSURE_MEMO_MAX_ENTRIES | 2000 | Max ProjectData Durable Objects one Worker isolate remembers as already having a persisted projectId, so ensureProjectId costs one RPC per isolate instead of one before every DO call |
Runtime Config Limits
Section titled “Runtime Config Limits”| Variable | Default | Description |
|---|---|---|
MAX_PROJECT_RUNTIME_ENV_VARS_PER_PROJECT | 150 | Max env vars per project |
MAX_PROJECT_RUNTIME_FILES_PER_PROJECT | 50 | Max files per project |
MAX_PROJECT_RUNTIME_ENV_VALUE_BYTES | 8192 | Max bytes per env var value |
MAX_PROJECT_RUNTIME_FILE_CONTENT_BYTES | 131072 | Max bytes per file content |
MAX_PROJECT_RUNTIME_FILE_PATH_LENGTH | 256 | Max file path length |
MAX_DEPLOYMENT_ENV_VARS_PER_ENVIRONMENT | 100 | Max deployment config vars per environment |
MAX_DEPLOYMENT_ENV_VALUE_BYTES | 65536 | Max bytes per deployment config value |
MAX_DEPLOYMENT_ENV_TOTAL_BYTES | 262144 | Max aggregate deployment config env size |
MAX_MCP_CONNECTIONS_PER_SCOPE | 25 | Max bring-your-own MCP servers per scope |
MCP_CONNECTION_URL_MAX_BYTES | 2048 | Max MCP endpoint URL size |
MCP_CONNECTION_TOKEN_MAX_BYTES | 8192 | Max MCP bearer token size |
MAX_MCP_CONNECTION_HEADERS | 10 | Max custom headers per MCP server |
MCP_CONNECTION_HEADER_VALUE_MAX_BYTES | 8192 | Max bytes per MCP custom header value |
External API Timeouts
Section titled “External API Timeouts”| Variable | Default | Description |
|---|---|---|
HETZNER_API_TIMEOUT_MS | 30000 | Hetzner API request timeout |
HETZNER_CAPACITY_RETRY_INITIAL_DELAY_MS | 15000 | Initial delay for transient Hetzner capacity retry backoff |
HETZNER_CAPACITY_RETRY_MAX_DELAY_MS | 120000 | Maximum delay per transient Hetzner capacity retry wait |
HETZNER_CAPACITY_RETRY_MAX_ATTEMPTS | 10 | Maximum transient Hetzner capacity retry attempts |
HETZNER_CAPACITY_RETRY_BUDGET_MS | 300000 | Total transient Hetzner capacity retry budget |
HETZNER_MAX_LIST_PAGES | 100 | Maximum pages per Hetzner list request |
CF_API_TIMEOUT_MS | 30000 | Cloudflare API request timeout |
GCP_API_TIMEOUT_MS | 30000 | GCP OAuth, IAM, and Compute request timeout |
NODE_AGENT_REQUEST_TIMEOUT_MS | 30000 | VM Agent request timeout |
DIGITALOCEAN_API_TIMEOUT_MS | 30000 | DigitalOcean API request timeout |
DIGITALOCEAN_IP_POLL_TIMEOUT_MS | 20000 | Bounded best-effort public IPv4 poll budget |
DIGITALOCEAN_IP_POLL_INTERVAL_MS | 3000 | Public IPv4 poll interval |
DIGITALOCEAN_ACTION_POLL_TIMEOUT_MS | 60000 | Block Storage action completion budget |
DIGITALOCEAN_ACTION_POLL_INTERVAL_MS | 1000 | Block Storage action poll interval |
DIGITALOCEAN_MAX_LIST_PAGES | 20 | Maximum pages per DigitalOcean list request |
DIGITALOCEAN_REGION | fra1 | Default DigitalOcean region |
DIGITALOCEAN_IMAGE | ubuntu-24-04-x64 | Default Droplet image slug |
CF_CONTAINER_CREATE_WORKSPACE_TIMEOUT_MS | 120000 | Instant-session create-workspace budget (includes in-container clone) |
Admin Observability
Section titled “Admin Observability”| Variable | Default | Description |
|---|---|---|
OBSERVABILITY_ERROR_RETENTION_DAYS | 30 | Error log retention |
OBSERVABILITY_ERROR_MAX_ROWS | 100000 | Max stored error rows |
OBSERVABILITY_ERROR_BATCH_SIZE | 25 | Error ingestion batch size |
OBSERVABILITY_ERROR_MESSAGE_MAX_LENGTH | 2048 | Maximum persisted message length |
OBSERVABILITY_ERROR_STACK_MAX_LENGTH | 4096 | Maximum persisted stack length |
OBSERVABILITY_ERROR_USER_AGENT_MAX_LENGTH | 512 | Maximum persisted user-agent length |
OBSERVABILITY_LOG_QUERY_RATE_LIMIT | 30 | Log queries per minute per admin |
VM TLS
Section titled “VM TLS”| Variable | Default | Description |
|---|---|---|
VM_AGENT_PROTOCOL | https | Protocol for VM agent communication |
VM_AGENT_PORT | 8443 | VM agent listening port |
VM_AGENT_MEMORY_RESERVE_MB | 512 | Optional Docker workload-slice memory reserve for VM-agent reachability headroom |
SAM_INFRA_SLICE_MEMORY_MIN_MB | 256 | systemd MemoryMin for the VM-agent/system-services slice |
DOCKER_MEMORY_MIN_MB | 512 | Minimum Docker MemoryMax retained when VM_AGENT_MEMORY_RESERVE_MB is enabled |
SAM_INFRA_SLICE_CPU_WEIGHT | 1000 | systemd CPUWeight for the VM-agent slice (cgroup v2 range 1–10000) |
SAM_WORKLOAD_SLICE_CPU_WEIGHT | 100 | systemd CPUWeight for the Docker workload slice (cgroup v2 default) |
HEARTBEAT_DOCKER_STATS_TIMEOUT | 2s | VM-agent timeout for heartbeat Docker stats used by workspace memory telemetry |
HEARTBEAT_WORKSPACE_METRICS_MAX_CONTAINERS | 8 | Maximum workspace containers measured by one heartbeat |
HEARTBEAT_WORKSPACE_METRICS_MAX_OUTPUT_BYTES | 65536 | Maximum bytes read from each heartbeat Docker metric command |
ORIGIN_CA_CERT_VALIDITY_DAYS | 7 | Validity for per-node Origin CA certificates signed by the API Worker |
New nodes generate /etc/sam/tls/origin-ca-key.pem locally in cloud-init and fetch only the signed certificate from POST /api/nodes/:id/origin-ca-certificate (packages/cloud-init/src/template.ts, apps/api/src/routes/node-lifecycle.ts). Legacy ORIGIN_CA_CERT and ORIGIN_CA_KEY Worker secrets are not required for new node provisioning.
VM workspace admission uses persisted workspaces.resolved_reservation_json snapshots and provider capacity fields in a final single-statement D1 reservation (apps/api/src/services/workspace-placement.ts). Aggregate CPU, memory, and disk reservations plus exclusiveNode control packing. Fresh telemetry remains mandatory on occupied nodes; disk pressure and CPU saturation veto reuse, while live memory percentage is a scoring signal rather than a second capacity gate (apps/api/src/services/workspace-resource-capacity.ts, apps/api/src/durable-objects/task-runner/node-selection.ts). There is no workspace-count or co-tenant cap: MAX_WORKSPACES_PER_NODE and maxCoTenants were removed, and legacy reservation rows that still carry maxCoTenants are admitted on their resource reservations alone. Live memory-threshold settings remain scoring inputs, not gates. VM agents report optional per-workspace memory telemetry in heartbeat metrics when Docker stats can be collected within the configured bounds (packages/vm-agent/internal/server/health.go, packages/vm-agent/internal/sysinfo/docker_metrics.go).
For managed VM nodes, the effective host-memory reserve contract is shared by admission and cloud-init: project scaling overrides win, then TASK_RUN_NODE_HOST_MEMORY_RESERVE_MB, then VM_AGENT_MEMORY_RESERVE_MB, then the 512 MB default. Cloud-init applies the cgroup hierarchy only on newly provisioned nodes. Existing nodes need a drain/recreate, a VM-agent/bootstrap upgrade flow, or a manual in-place systemd/Docker reconfiguration before they can be treated as protected by the workload-slice cap; SAM does not destructively evict existing workspaces to retrofit this.
Journald Configuration (VM)
Section titled “Journald Configuration (VM)”Applied via cloud-init on each node:
| Setting | Default | Description |
|---|---|---|
SystemMaxUse | 500M | Max disk space for journal |
SystemKeepFree | 1G | Minimum free disk to maintain |
MaxRetentionSec | 7day | Max log retention period |
Storage | persistent | Persist logs across reboots |
Compress | yes | Compress stored entries |
File Upload & Download
Section titled “File Upload & Download”| Variable | Default | Description |
|---|---|---|
FILE_UPLOAD_MAX_BYTES | 52428800 (50 MB) | Max size per uploaded file |
FILE_UPLOAD_BATCH_MAX_BYTES | 262144000 (250 MB) | Max total size per upload batch |
FILE_UPLOAD_TIMEOUT | 120s | Upload timeout (VM agent) |
FILE_UPLOAD_TIMEOUT_MS | 120000 (120s) | Upload proxy timeout (Worker) |
FILE_DOWNLOAD_TIMEOUT_MS | 60000 (60s) | Download proxy timeout |
FILE_DOWNLOAD_MAX_BYTES | 52428800 (50 MB) | Max download file size |
File Browsing & Raw Proxy
Section titled “File Browsing & Raw Proxy”| Variable | Default | Description |
|---|---|---|
FILE_PROXY_TIMEOUT_MS | 15000 | File proxy request timeout |
FILE_PROXY_MAX_RESPONSE_BYTES | 2097152 (2 MB) | Max file proxy response size |
FILE_RAW_MAX_SIZE | 52428800 (50 MB) | Max raw binary file size (VM agent) |
FILE_RAW_TIMEOUT | 60s | Raw file streaming timeout (VM agent) |
FILE_RAW_PROXY_MAX_BYTES | 52428800 (50 MB) | Max raw file proxy size (Worker) |
Project Files (Remote-Branch Git Browser)
Section titled “Project Files (Remote-Branch Git Browser)”| Variable | Default | Description |
|---|---|---|
REPO_BROWSE_MAX_INLINE_BYTES | 1000000 (1 MB) | Max bytes to inline as text in the file viewer; larger stream raw |
REPO_BROWSE_MAX_COMPARE_FILES | 300 | Max changed files in an Artifacts diff before truncation |
MCP Tool Limits
Section titled “MCP Tool Limits”| Variable | Default | Description |
|---|---|---|
MCP_IDEA_CONTEXT_MAX_LENGTH | 500 | Max characters of idea context shown to agents |
MCP_IDEA_LIST_LIMIT | 20 | Default page size for list_ideas |
MCP_IDEA_LIST_MAX | 100 | Max page size for list_ideas |
MCP_IDEA_SEARCH_MAX | 20 | Max results from search_ideas |
MCP_RELATED_IDEA_SEARCH_LIMIT | 10 | Default results from find_related_ideas |
MCP_TASK_LIST_LIMIT | 10 | Default page size for list_tasks |
MCP_TASK_LIST_MAX | 50 | Max page size for list_tasks |
MCP_TASK_SEARCH_LIMIT | 10 | Default page size for search_tasks |
MCP_TASK_SEARCH_MAX | 20 | Max results from search_tasks |
MCP_TASK_DETAIL_RECENT_MESSAGE_LIMIT | 5 | Recent assistant messages returned by get_task_details |
MCP_TASK_DETAIL_MESSAGE_SNIPPET_LENGTH | 2000 | Max characters per assistant message snippet in get_task_details |
MCP_MESSAGE_SEARCH_LIMIT | 10 | Default page size for search_messages |
MCP_MESSAGE_SEARCH_MAX | 20 | Max results from search_messages |
SEARCH_QUERY_MAX_LENGTH | 4096 | Max total UTF-8 bytes retained by idea, task, knowledge, and message search as a DoS guard; responses disclose truncation |
SEARCH_QUERY_MAX_TERM_LENGTH | 48 | Max LIKE-safe UTF-8 bytes retained per search term; higher values clamp to SQLite’s safe pattern ceiling |
SEARCH_QUERY_MAX_TERMS | 40 | Max whitespace-delimited terms retained by those search surfaces; higher values clamp to the safe D1 parameter ceiling and responses disclose truncation |
MCP_MESSAGE_LIST_LIMIT | 50 | Default page size for get_session_messages |
MCP_MESSAGE_LIST_MAX | 200 | Max messages per get_session_messages request |
MCP_ARCHIVED_TOOL_PAYLOAD_LIST_LIMIT | 10 | Default page size for get_archived_tool_payloads |
MCP_ARCHIVED_TOOL_PAYLOAD_LIST_MAX | 50 | Max archived payloads per get_archived_tool_payloads |
MCP_COMMENT_LIST_LIMIT | 10 | Default page size for list_message_comment_threads |
MCP_COMMENT_LIST_MAX | 25 | Max threads per list_message_comment_threads request |
MCP_COMMENT_BODY_MAX_LENGTH | 4000 | Max comment/reply body characters accepted through MCP |
MCP_COMMENT_QUOTE_MAX_LENGTH | 1000 | Max quoted source-message characters returned to agents |
COMMENT_DIRECTIVE_CONTEXT_MAX_LENGTH | 6000 | Max send-to-agent comment directive prompt length |
MCP_TRIGGER_LIST_LIMIT | 20 | Default page size for list_triggers |
MCP_TRIGGER_LIST_MAX | 100 | Max triggers per list_triggers request |
MCP_INCIDENT_LIST_LIMIT | 10 | Default page size for private list_incident_queue |
MCP_INCIDENT_LIST_MAX | 50 | Max private incidents per list_incident_queue request |
Project event MCP tools use the ProjectData event limits above: PROJECT_EVENT_LIST_LIMIT, PROJECT_EVENT_LIST_MAX, and PROJECT_EVENT_SUBSCRIPTION_EVENT_CURSOR_MAX_LENGTH.
Web UI (Build-Time)
Section titled “Web UI (Build-Time)”| Variable | Default | Description |
|---|---|---|
VITE_FILE_PREVIEW_INLINE_MAX_BYTES | 10485760 (10 MB) | Images below this size render inline automatically |
VITE_FILE_PREVIEW_LOAD_MAX_BYTES | 52428800 (50 MB) | Images below this size show click-to-load; above shows download link |
VITE_ANALYTICS_MAX_QUEUE_SIZE | 100 | Max client-side analytics events retained before oldest events drop |
VITE_ANALYTICS_FLUSH_THRESHOLD | 10 | Client event count that triggers an immediate analytics flush |
VITE_ANALYTICS_FLUSH_INTERVAL_MS | 5000 | Client analytics background flush interval in milliseconds |
VITE_DEBUG_DIAGNOSIS_EVENT_MAX_PAGES | 100 | Max paginated diagnosis-event pages loaded per browser request |
VITE_CHAT_DELTA_MAX_PAGES | 50 | Max newer-message pages one chat refresh drains before failing visibly |
VITE_CHAT_TIMELINE_MAX_PAGES | 200 | Max pages fetched per loop when the chat timeline drawer opens |
VITE_CHAT_LOAD_UNTIL_MAX_PAGES | 400 | Max older-message pages chased while resolving a timeline jump |
VITE_PROJECT_LIST_LIMIT | 50 | Projects loaded into each shared list-cache entry |
VITE_PROJECT_POLL_INTERVAL_MS | 30000 | Project-list page refresh cadence in milliseconds; 0 disables |
VITE_SIDEBAR_PROJECT_POLL_INTERVAL_MS | 60000 | App-shell project-list refresh cadence in milliseconds; 0 disables |
VITE_WORKSPACE_PORTS_POLL_MS | 10000 | Workspace forwarded-port base refresh cadence in milliseconds |
VITE_WORKSPACE_PORTS_BACKOFF_MAX_MS | 120000 | Maximum backoff between forwarded-port readiness polls |
VITE_WORKSPACE_PORTS_FAILURE_BUDGET | 6 | Consecutive unavailable port-list responses before circuit cooldown |
VITE_WORKSPACE_PORTS_BACKOFF_JITTER_RATIO | 0.2 | +/- jitter ratio applied to forwarded-port readiness backoff delays |
VITE_WORKSPACE_PORTS_CIRCUIT_RESET_MS | 300000 | Open-circuit cooldown before probing forwarded-port readiness again |
VITE_SESSION_INFRA_RETRY_DELAYS_MS | 2000,5000,10000 | Comma-separated retry delays for chat Details workspace/node fetches |
VITE_PROJECT_PREFETCH_DELAY_MS | 120 | Mouse dwell before project-detail prefetch; focus/touch are immediate |
VITE_BACKGROUND_FETCH_DELAY_MS | 150 | Delay before background query activity is shown and announced |
VITE_ACP_PERMISSION_POLL_MS | 2000 | Pending ACP permission snapshot refresh cadence |
VITE_ACP_PERMISSION_RECOVERY_POLL_MS | 30000 | Empty-snapshot recovery cadence after a missed realtime attention event |
VITE_ACP_PERMISSION_QUERY_RETRY_COUNT | 3 | Transient ACP permission snapshot retries; 0 disables retries |
VITE_CHUNK_LOAD_RETRY_DELAY_MS | 350 | Wait before retrying a failed lazy route-chunk import |
VITE_CHUNK_RELOAD_COOLDOWN_MS | 15000 | Minimum gap between chunk-recovery reloads; guards against a reload loop |
VITE_ROUTE_FALLBACK_REVEAL_DELAY_MS | 180 | Delay before the route loading spinner fades in, avoiding a flash |
VITE_QUERY_PERSIST_MAX_AGE_MS | 86400000 (24 h) | How long a persisted query-cache record may be restored after writing |
VITE_QUERY_PERSIST_THROTTLE_MS | 1000 | Minimum gap between IndexedDB writes of the query cache |
VITE_QUERY_PERSIST_RESTORE_TIMEOUT_MS | 250 | Budget for the initial cache restore before failing open to no cache |
VITE_CHAT_TRANSCRIPT_CACHE_TTL_MS | 86400000 (24 h) | How long a chat transcript stays cached after it was last used |
VITE_CHAT_TRANSCRIPT_CACHE_MAX_SESSIONS | 20 | Most chat transcripts cached at once; older ones are evicted |
VITE_CHAT_TRANSCRIPT_PERSIST_MAX_ROWS | 500 | Newest rows of each chat transcript written to IndexedDB |
VITE_AGENT_CATALOG_STALE_TIME_MS | 300000 | Freshness window for the installable agent catalog query |
VITE_PROVIDER_CATALOG_STALE_TIME_MS | 300000 | Freshness window for provider catalog size/location/price metadata |
VITE_TRIAL_STATUS_STALE_TIME_MS | 60000 | Freshness window for trial availability status |
VITE_CACHED_COMMANDS_STALE_TIME_MS | 300000 | Freshness window for cached slash-command registries |
VITE_REPORT_ISSUE_CONFIG_STALE_TIME_MS | 300000 | Freshness window for the report-issue availability flag |
VITE_PROJECT_CREATE_CONFIG_STALE_TIME_MS | 300000 | Freshness window for project-creation config flags |
VITE_CREDENTIAL_LIMITS_STALE_TIME_MS | 30000 | Freshness window for credential usage-limit windows shown in the chat header and Settings → Advanced |
VITE_CREDENTIAL_LIMITS_REFETCH_INTERVAL_MS | 60000 | Poll cadence for credential usage-limit windows while the tab is visible (paused in the background) |
Query cache persistence
Section titled “Query cache persistence”The control-plane UI writes an allowlisted slice of its query cache to IndexedDB so a full page reload paints from cache instead of refetching. Persisted slices are limited to allowlisted project summaries, stripped library indexes, and project-chat session messages. Credentials, admin diagnostics, node and workspace runtime details, file contents, signed URLs, and mutation state are never written to disk.
Records are namespaced by authenticated user and by a schema version, and are deleted on sign-out and on account switch, so one account can never be shown another account’s cached data. If IndexedDB is unavailable — private browsing, a storage quota failure, or a disabled store — the app degrades silently to its normal in-memory cache.
Project chat transcripts are kept for VITE_CHAT_TRANSCRIPT_CACHE_TTL_MS after they were last
loaded or updated, capped at the VITE_CHAT_TRANSCRIPT_CACHE_MAX_SESSIONS most recently used. A
chat opened inside that window renders from the cache at once and refreshes in the background; a
chat that is not cached loads its newest page first, and older history loads as you scroll up. On
disk each transcript keeps only its newest VITE_CHAT_TRANSCRIPT_PERSIST_MAX_ROWS rows, however far
back it was read, so a chat restored after a reload pages older history back in the same way. The
cache is only read after the sign-in check completes.
Analytics
Section titled “Analytics”SAM uses first-party analytics ingestion for operational/product aggregates. Browser events are batched to /api/t; request analytics are written by API middleware when enabled. Analytics is best-effort and disabled paths preserve normal application behavior.
Client page/referrer fields follow a privacy normalization contract before enqueue: query strings, fragments, protocol, host/userinfo for page values, credentials, emails, UUIDs/ULIDs, long opaque tokens, common secret prefixes, repository/code file identifiers, and values after sensitive route markers are removed or replaced with [redacted]. Non-sensitive nested path shape, event names, durations, UTM source/medium/campaign, session ID, visitor/authenticated user ID, and explicit safe entity metadata are preserved for aggregate reporting.
| Variable | Default | Description |
|---|---|---|
ANALYTICS_ENABLED | true | Enable API middleware analytics; set false to skip request event writes |
ANALYTICS_SKIP_ROUTES | (built-in skip list) | Comma-separated extra route prefixes/patterns excluded from middleware writes |
ANALYTICS_DATASET | (deployment-generated) | Cloudflare Analytics Engine dataset name |
ANALYTICS_SQL_API_URL | https://api.cloudflare.com/client/v4/accounts | Analytics Engine SQL API base URL override |
ANALYTICS_DEFAULT_PERIOD_DAYS | 30 | Default admin analytics query lookback in days |
ANALYTICS_TOP_EVENTS_LIMIT | 50 | Max rows returned by top-events admin query |
ANALYTICS_GEO_LIMIT | 50 | Max countries in geographic distribution view |
ANALYTICS_RETENTION_WEEKS | 12 | Number of weeks for retention cohort analysis |
ANALYTICS_WEBSITE_TRAFFIC_TOP_PAGES_LIMIT | 20 | Max top pages/referrers/events in website traffic sections |
ANALYTICS_INGEST_ENABLED | true | Enable browser event ingestion at /api/t; false returns success without writes |
RATE_LIMIT_ANALYTICS_INGEST | 500 | Analytics ingest requests allowed per IP per hour |
MAX_ANALYTICS_INGEST_BATCH_SIZE | 25 | Max browser events accepted per ingest request |
MAX_ANALYTICS_INGEST_BODY_BYTES | 65536 | Max ingest request body size in bytes |
MAX_ANALYTICS_DURATION_MS | 3600000 | Max accepted page-duration value; larger values are clamped |
Analytics Forwarding
Section titled “Analytics Forwarding”External analytics forwarding is off by default. When enabled, SAM forwards only analytics rows already accepted by first-party ingestion/middleware; it does not bypass the client-side URL normalization contract.
| Variable | Default | Description |
|---|---|---|
ANALYTICS_FORWARD_ENABLED | false | Enable external analytics event forwarding |
ANALYTICS_FORWARD_EVENTS | key conversion events | Comma-separated list of events to forward |
ANALYTICS_FORWARD_LOOKBACK_HOURS | 25 | Hours to look back for events |
ANALYTICS_FORWARD_CURSOR_KEY | analytics-forward-cursor | KV key used to remember forwarded progress |
ANALYTICS_FORWARD_SQL_LIMIT | 10000 | Max rows fetched per forwarding run |
ANALYTICS_SQL_FETCH_TIMEOUT_MS | 30000 | Timeout for Analytics Engine SQL fetches |
SEGMENT_WRITE_KEY | (unset) | Segment Write Key for event forwarding |
SEGMENT_API_URL | https://api.segment.io/v1/batch | Segment API endpoint |
SEGMENT_MAX_BATCH_SIZE | 100 | Max events per Segment batch request |
GA4_MEASUREMENT_ID | (unset) | Google Analytics 4 Measurement ID |
GA4_API_SECRET | (unset) | Google Analytics 4 API secret |
GA4_API_URL | https://www.google-analytics.com/mp/collect | GA4 Measurement Protocol endpoint |
GA4_MAX_BATCH_SIZE | 25 | Max events per GA4 batch request |
Compact archive shards
Section titled “Compact archive shards”PROJECT_DATA_ARCHIVE_COMPACT_ENABLED=false is the default. Enabling it affects newly journaled terminal-session migrations only. Their format is pinned in D1 and the shard, so retries, reads and copy-back continue with that format after the flag is disabled. Legacy shards are not rewritten by deployment. Keep the existing global-sweep throttle while validating a compact canary.
| Optional Worker variable | Default | Meaning |
|---|---|---|
PROJECT_DATA_ARCHIVE_COMPACT_ENABLED | false | Write new archives as compact SQLite + compressed R2; existing formats remain readable. |
PROJECT_DATA_ARCHIVE_DAILY_WRITE_BUDGET | 250000 | Installation-wide daily allowance of ESTIMATE UNITS (not billed rows) for compact migration attempts; 0 pauses admission. The checked-in wrangler.toml ships 2400000. Divide by 1000 + factor x units-per-session for the migrations/day ceiling: at the shipped factor of 2 and a measured ~10,250-unit session that is ~111/day, above the roughly 72 claim opportunities/day the 18-minute sweep cadence offers. Whichever of the two is smaller is what actually runs, so check both before tuning either. |
PROJECT_DATA_ARCHIVE_WRITE_ESTIMATE_FACTOR | 32 | Safety multiplier applied to the row inventory estimateArchiveWrites already counts (rows, archived tool payloads, grouped rows, and 512-byte units of grouped FTS text, including source deletion) — NOT an independent amplification factor. The checked-in wrangler.toml ships 2. Production measurement 2026-09-14: two project_data_archive_write_budget samples an hour apart give ~65,001 estimated writes per migration, i.e. ~8,000 inventory units at the then-shipped factor of 8; four isolated archive ticks billed 5,950/6,436/7,942/9,823 Durable Object rows. Billed rows per inventory unit therefore spans 0.74-1.23, and 2 exceeds the worst observed ratio by ~63%. Raising it does not make the work safer, only more expensive in budget terms: at 8 the same allowance bought 12 migrations/day while the hourly cadence allowed 24, so half the sweep ticks reclaimed nothing. Sample size is small (2 budget samples, 4 telemetry buckets) — re-measure before changing it. |
PROJECT_DATA_ARCHIVE_R2_TIMEOUT_MS | 10000 | R2 I/O deadline shared across each compact read/export/seal operation, or one chunk write, in milliseconds. |
PROJECT_DATA_ARCHIVE_BUDGET_RECEIPT_RETENTION_MS | 604800000 | Unused reservation receipt retention (7 days); clamped to at least one UTC budget day. |
PROJECT_DATA_ARCHIVE_BUDGET_RECEIPT_CLEANUP_LIMIT | 100 | Maximum expired receipts removed per unused release; zero disables cleanup. |
Compact archives keep session metadata, a dedicated derived search projection, consolidated conversation text, grouped FTS and existing tool-archive pointers in SQLite. Original message rows (IDs, timestamps, order, origins and full tool metadata) live in immutable gzip R2 chunks in the private PROJECT_DATA_ARCHIVE_R2 binding. SQL chunk references contain sizes, SHA-256, role counts, time bounds and message IDs. Exact history and tool expansion fetch and verify those chunks. Missing/corrupt objects fail the request; they do not silently produce a truncated history. The original terminal hash and search-projection count/hash coverage must match before publication and source deletion. Already-published archives without current coverage are repaired through durable raw/grouped cursors: compact raw history advances by configured R2 chunks, while legacy raw rows and grouped fallback advance by the configured hash-page size. The resumable projection commitment and FTS counts are verified before coverage becomes complete. This is one-time backfill work and is reported separately from steady-state query coverage. Version 2 recovery manifests include the compressed-object references; legacy version 1 manifests remain unchanged.
Compact chunk exports use the smaller of PROJECT_DATA_ARCHIVE_CHUNK_BYTES and 2 MiB, with an 8 MiB serialized/decompressed object ceiling. A single valid row above the configured chunk target travels alone up to the absolute Durable Object RPC ceiling; it is never truncated or allowed to livelock the cursor. Rows above that platform ceiling still fail closed and remain on the source. R2 objects have no automatic expiry: deleting them destroys archive history and recovery data.
Everything in this section applies only when PROJECT_DATA_ARCHIVE_COMPACT_ENABLED is true: the write-budget reservation and the derived selection ceiling are both gated on it (selectCandidates / processArchiveMigrationBatch in apps/api/src/scheduled/project-data-archive-sharding.ts). Legacy (non-compact) archiving reserves nothing and is bounded only by PROJECT_DATA_ARCHIVE_SWEEP_MESSAGE_BUDGET.
The write allowance is a durable admission estimate, not a hard invoice cap. Each attempt reserves 1,000 writes plus the estimate factor times the sum of raw rows, grouped rows, tool-pointer rows and grouped UTF-8 text bytes rounded up to 512-byte units. The installation-wide D1 reservation is atomic, resets on the next UTC day and is never refunded after an interrupted attempt. Contenders that definitively lose journal creation or lease acquisition before doing archive writes release their reservation with an idempotent D1 receipt. The pool covers this SAM installation; other installations on the same Cloudflare account need separate headroom. Source inventory is rechecked under the transcript lock before creating its intent. New sessions that cannot reserve remain readable on root; existing interrupted migrations retain their existing recovery fence and can retry when allowance is available. Candidate selection is bounded by the SMALLER of PROJECT_DATA_ARCHIVE_SWEEP_MESSAGE_BUDGET and a ceiling derived from the daily allowance itself. The derivation floors twice, in two steps: archiveAffordableWriteUnits() computes floor((allowance - 1000) / factor) — the largest estimate the allowance can ever admit — and archiveAffordableMessageCeiling() then computes floor(units / (1 + PROJECT_DATA_ARCHIVE_SWEEP_UNIT_OVERHEAD_PERCENT / 100)) (both in apps/api/src/project-data-archive/write-budget.ts). Collapsing the two floors into one expression can differ by one at some factor/overhead combinations. Deriving the second half means a selector can never offer a candidate the write budget must refuse on every attempt. Two independently configured ceilings previously had to agree by hand, and when the deployed allowance was lowered without lowering the message budget, largest-first selection re-picked the same unaffordable session on every pass and reclaimed nothing for four days while still reporting success.
The derived ceiling assumes an overhead; it does not measure one. A session whose real tool-payload, grouped-row and FTS-unit cost exceeds the assumption is still refused at reservation time, and the pass then descends to the next-smaller candidate (PROJECT_DATA_ARCHIVE_SWEEP_FALLTHROUGH_DEPTH controls how many spare candidates it reads for that purpose). A refused candidate opens no migrating fence, stays readable on root, and consumes neither a session slot nor any of the cumulative message budget.
The two refusals are distinguished because they need different responses. exceeds_allowance means the session costs more than the entire daily pool and waiting cannot help; window_exhausted means the day’s pool is spent and the next UTC window refills it. After PROJECT_DATA_ARCHIVE_BUDGET_STALL_ALERT_SWEEPS consecutive passes that migrate nothing and see only the former, the sweep cadence row reports partial with an actionable last_error instead of succeeded. Sessions above every ceiling require an explicitly sized migration plan; increasing the daily budget alone does not lift the message cap.
Whole-session deletion and SQLite/FTS index maintenance still consume writes. project_data_archive_sql_usage reports actual cursor writes. project_data_archive_candidate_migrated.sourceFinalization reports source rows, before/after bytes, duration, and reclaimed bytes; copy checkpoints retain the last operation, operation ID, ordinal, byte totals, lease epoch, and timing for reset attribution. These invocation timings are diagnostic only: billed Durable Object duration comes from Cloudflare durableObjectsPeriodicGroups.sum.duration. Compare aligned daily metrics with these events before raising throughput, and separate one-time repair/backfill from steady-state search. At the default 250,000 estimated writes/day, 30 days admits at most 7.5 million estimated migration writes. Normal application traffic, legacy migrations, operator copy-back and other Workers are outside this pool, so operators must reserve account headroom separately. A zero allowance pauses new compact attempts without breaking reads or completed crash-gap publication. Source deletion is still one whole-session operation, not an interruptible per-row spending limit.
After any compact archive is published, rollback must retain compact readers (disable the writer flag to pause new migrations). Deploying a binary from before compact-reader support would read empty raw SQL tables for those sessions; copy them back and verify root ownership before considering such a downgrade.
Fresh VM boot recovery
Section titled “Fresh VM boot recovery”Certificate issuance retries transport failures, HTTP 429 and HTTP 5xx with capped exponential backoff. Cloud-init also retries its certificate fetch and reports a fixed boot-failure reason when bootstrap cannot continue. Tasks replace a failed fresh VM only after its deletion is confirmed, before any workspace execution. The default is one replacement; a second boot failure ends with its specific reason. Wrong agent versions fail immediately, while a VM without any heartbeat gets six minutes by default.
ORIGIN_CA_RETRY_MAX_ATTEMPTS— Maximum upstream certificate attempts including the first; retries transport, 429 and 5xx only (default: 3).ORIGIN_CA_RETRY_BASE_DELAY_MS— Initial upstream certificate retry delay (default: 500).ORIGIN_CA_RETRY_MAX_DELAY_MS— Cap on upstream certificate exponential backoff (default: 2000).ORIGIN_CA_REQUEST_TIMEOUT_MS— Deadline per upstream certificate request, including reading its body (default: 10000).CLOUD_INIT_AGENT_DOWNLOAD_TIMEOUT_SECONDS— Deadline for the agent binary download, including DNS and connection time (default: 60). Failure triggers the best-effort boot-failure callback.CLOUD_INIT_CERT_MAX_ATTEMPTS— Maximum cloud-init CSR POST attempts including the first (default: 3).CLOUD_INIT_CERT_BASE_DELAY_SECONDS— Initial cloud-init certificate retry delay (default: 2).CLOUD_INIT_CERT_MAX_DELAY_SECONDS— Cap on cloud-init certificate exponential backoff (default: 8).CLOUD_INIT_CERT_REQUEST_TIMEOUT_SECONDS— Deadline per cloud-init certificate request and best-effort boot-failure report; must exceed the upstream API retry budget (default: 45).TASK_RUNNER_FIRST_HEARTBEAT_TIMEOUT_MS— Fresh VM first-heartbeat deadline, capped by TASK_RUNNER_AGENT_READY_TIMEOUT_MS (default: 360000).TASK_RUNNER_BOOT_MAX_REPLACEMENTS— Maximum fresh VM boot replacements per task run; 0 disables replacement; workspace execution is never replayed (default: 1).
Set these optional variables in the GitHub deployment Environment; the deployment pipeline forwards them to the Worker. No new secrets are required. Certificate request and retry budgets should remain below the first-heartbeat deadline.
CLI operation receipt limits
Section titled “CLI operation receipt limits”The optional Worker variables CLI_RECEIPT_REQUEST_MAX_BYTES (default 262144) and
CLI_RECEIPT_RESPONSE_MAX_BYTES (default 65536) accept positive byte counts.
Keyed requests exceeding the request limit are rejected before reservation; replies
exceeding the response limit leave the receipt pending for reconciliation. These
are runtime configuration overrides, not credentials or required deployment secrets.